Tools

Experiment Designer

Most of the changes I’ve shipped couldn’t be A/B tested. The change lived on the product page rather than the person, or pricing had to be identical for every customer, or it reached twelve warehouses on a schedule operations had already committed to.

Standard A/B calculators assume the one thing that isn’t true in those cases. So teams fall back on the month before and the month after, put the difference in a deck, and hope nobody asks what else happened that month. Something usually did. A promotion, a seasonal swing, a competitor outage, a pricing change three teams over. None of it separates out from a before-and-after number, which is why everyone reading one quietly discounts it.

There’s better available. If you can’t randomize who, you can often randomize what: half the catalog gets the new copy. Or when: the feature switches on and off in blocks, or reaches sites in a randomized order. Failing all three, you can still choose a counterfactual you’re willing to defend out loud, which beats not having chosen one.

This tool covers which design survives your constraint, how much data it needs, and how to read the result without claiming more than it supports. It opens by asking why you can’t split traffic, because about half the time the honest answer is “it’s awkward,” not “it’s impossible.” Awkward is worth paying an engineer to fix rather than a statistician to work around.

Start from a worked example

The statistics are the easy part. Almost every experiment that goes wrong goes wrong here — by counting a unit that was never randomized, or by working around a constraint that was never really there.

1Why can't you split traffic?

Be specific. The answer decides everything downstream, and a wrong answer here cannot be fixed by any amount of statistics later.

›Methods, and where this is deliberately simple

Normal CDF via the Hart/West rational approximation; inverse normal via Acklam with a Halley refinement; Student’s t through the regularized incomplete beta (Lentz continued fraction); Beta sampling via Marsaglia–Tsang on a seeded mulberry32 PRNG; stepped-wedge variance from Hussey and Hughes (2007); segmented regression by ordinary least squares with a Gauss–Jordan inverse. Everything runs in your browser. Nothing you type leaves the page.

Known simplifications: cluster sizes are treated as equal, which understates the design effect when they vary a lot; the stepped-wedge power formula assumes a constant treatment effect and a balanced schedule; the difference-in-differences model is the two-period case and cannot test parallel trends for you; the interrupted time series is ordinary least squares with no seasonal or autoregressive terms. There is no sequential-testing correction, because the honest fix is to fix the sample size in advance rather than to patch the math afterward.

Count the unit you randomized

This is the one that ruins most of these tests, and it never looks like a mistake. You split 4,000 products into two halves, collect 300,000 sessions, and read the result as though you had 300,000 independent observations. You had 4,000. Every session on a product carries much of the same information as every other session on it, so the p-value comes back far smaller than the evidence warrants and the test looks decisive when it isn’t.

The fix is arithmetic, not judgment: inflate the variance by the design effect, or analyze at the level you randomized. Both are in the steps above. Worth knowing in advance, because it often multiplies the sample you need by three or more — the difference between a four-week test and one nobody will wait for.

Two ways to read the same data

The read step shows a frequentist and a Bayesian view of identical inputs. That isn’t hedging. They answer different questions, and the second is usually the one that was actually asked. A p-value tells you how surprising this result would be if the change did nothing. Expected regret tells you what shipping costs you if you’re wrong.

The most useful output of either is rarely the verdict. It’s the width of the interval, which tells you whether the test could have detected what you were looking for. A result that “wasn’t significant” on an interval spanning −20% to +25% isn’t a finding. It’s a test that never had a chance, and it should be reported that way.

On the methods

The statistics are standard and the sources are below. This isn’t original work — it’s textbook procedures put behind inputs a product team will actually fill in. Everything runs in your browser; nothing you type is sent anywhere.

The simplifications are listed under Methods in the tool, and they’re real: equal cluster sizes, a constant treatment effect in the wedge, no seasonal or autoregressive terms in the time series. For a decision worth more than an afternoon, have someone check the model against your data rather than trusting these panels.

Sources
  • Hussey, M.A. & Hughes, J.P. (2007). “Design and analysis of stepped wedge cluster randomized trials.” Contemporary Clinical Trials, 28(2), 182–191. — the staggered-rollout variance formula
  • Donner, A. & Klar, N. (2000). Design and Analysis of Cluster Randomization Trials in Health Research. Arnold, London. — the design effect and the ICC
  • Kohavi, R., Tang, D. & Xu, Y. (2020). Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing. Cambridge University Press. — switchbacks, interference, and when to stop looking
  • Wagner, A.K., Soumerai, S.B., Zhang, F. & Ross-Degnan, D. (2002). “Segmented regression analysis of interrupted time series studies in medication use research.” Journal of Clinical Pharmacy and Therapeutics, 27(4), 299–309. — the segmented-regression setup behind the time-series read
  • Bertrand, M., Duflo, E. & Mullainathan, S. (2004). “How much should we trust differences-in-differences estimates?” Quarterly Journal of Economics, 119(1), 249–275. — how much to trust a cohort comparison
  • Gelman, A., Carlin, J.B., Stern, H.S., Dunson, D.B., Vehtari, A. & Rubin, D.B. (2013). Bayesian Data Analysis, 3rd ed. Chapman & Hall/CRC. — the conjugate models behind both Bayesian readings
  • Kruschke, J.K. (2015). Doing Bayesian Data Analysis: A Tutorial with R, JAGS, and Stan, 2nd ed. Academic Press. — the region of practical equivalence and loss-based stopping