Experimentation · Causal Inference

I make A/B testing trustworthy

You ship "wins" that are noise. You kill tests that were real. I'm the fractional senior owner who makes the difference obvious, and tells you which numbers are safe to ship on.

Free · no email required  ·  or book a 20-min call to skip ahead

Guarantee a written verdict on decisions we name together, or you don't pay.
David Arzumanian, fractional experimentation lead
Previously at
Free diagnostic · no email required

Are your tests lying to you?

Five questions. I'll score your pipeline's trustworthiness out of 100, surface the top 3 risks specific to your setup, and tell you exactly which of your "wins" are most likely noise.

~90 seconds
Step 1 of 6
01Which A/B testing tool do you use?
02How many A/B tests do you ship per month?
5per month
03Do you stop tests early when results look "significant"?
Stopping the instant a test crosses 0.05, without a valid sequential rule (always-valid p-values, group-sequential, Bayesian), inflates false positives. If you stop early but use one of those, answer "Never".
04Do you correct for multiple comparisons?
Correcting when you test many metrics, segments, or variants at once (FDR, Bonferroni). Pick "Not sure" if this is new.
05Who reviews experiment results before they ship?
06Did you recently ship something you couldn't A/B test?
A pricing change, redesign, rebrand, or market rollout with no control group. This doesn't affect your score; it tells me which service fits you.
Your trustworthiness score
0/ 100
What you get

The Experimentation Trust Audit

A two-week, fixed-fee review of your A/B testing. I find the numbers you can't trust, show you which "wins" are noise, and hand you a clear way to fix it. Agencies sell this as a 30–60 day engagement starting at $20,000; mine is two weeks from data access, at a third of that.

01
A trustworthiness scorecard
Every recent test scored out of 100, with the false wins flagged and the statistical reason each one is shaky.
02
A prioritized fix roadmap
The handful of changes that move your reads from hopeful to trustworthy, in the order they pay off.
03
A team walkthrough
A working session so the fixes stick after I leave. Not a PDF nobody opens.

$6,500 fixed fee, two weeks from data access, remote. See the pricing →

Why you can trust me

Eight years making numbers tell the truth.

Two kinds of proof: what I've caught, and what I've built.

59×
a revenue metric was inflating results. It looked exactly like real signal. Caught before the team shipped on it.
2 rates ↑
churned and retained both rose in one test, a polarization artifact. I killed a "winner" that would have hurt.
30–40%
earlier that experiments could have stopped, with sequential testing. Faster decisions, same confidence.
Systems I've shipped in production
Sequential testingValid early stopping, so tests answer in days instead of weeks.30–40% faster
CaliperScores which test to run next by the learning it buys.prioritised backlog
Smart AnalyticsCrewAI agents that take the grunt work out of analysis.three crews
Causal measurementInterrupted time series and diff-in-diff for launches you can't A/B test.no-A/B rollouts
Churn modelingXGBoost and SHAP routing, validated with causal rigor.in production
What it costs

Three ways to work together.

Start here
Experimentation Trust Audit
$6,500 fixed
Founding rate. List price $7,500.
Two weeks from data access, remote.
  • Trustworthiness scorecard
  • Fix roadmap, written as your data lead's 90-day plan
  • Live team walkthrough
Guarantee. A written verdict, with evidence you didn't have, on 2–3 shipping decisions we name together, or you don't pay.
Book a 20-minute call
Launch Impact Study
$9,500 fixed
Founding rate. List price $11,000.
You shipped it. Did it work? Three weeks, starting with a $2,000 feasibility check that is fully credited.
  • For pricing changes, redesigns, rollouts: anything you couldn't A/B test
  • Week 1: a board-ready feasibility memo
  • Either side can stop at $2,000, you keep the memo
  • A defensible causal read: interrupted time series, diff-in-diff
The gate is the guarantee. If week one shows your data can't support a credible answer, we stop at $2,000.
Book a call
Fractional Lead
$10,000 /mo
First two client slots. Then $12,000/mo.
2 days a week plus 2 flex days around launches. Month-to-month after a 3-month start. Typically follows an audit or study.
  • I run the roadmap from your audit
  • Own the experiment review and decision rituals
  • Sequential testing policy, rolled out with your team
  • Level up your team
Book a call

Founding pricing holds for the first five project engagements or until June 30, 2027, whichever comes first. The list prices above apply after that.

Common questions

Questions I get from product teams

Isn't "just don't peek" enough?+
Telling people not to peek doesn't survive a launch week. The real fix is sequential testing: valid stopping rules so you CAN look early without inflating false positives.
Do I have to switch tools?+
No. The methodology is tool-agnostic. I audit whatever you run: Amplitude, Statsig, GrowthBook, Eppo, or something built in-house.
We're pre-product-market-fit. Too early?+
Probably, and I'll tell you that on the call rather than take the fee. This is for teams with real traffic and tests they can't fully trust yet.
What do I actually walk away with?+
A trustworthiness scorecard, a prioritized fix roadmap, and a team walkthrough. See the audit →
The next step

Which of your numbers are safe to ship on?

90 sec
Free diagnostic. No email. Score your pipeline and see your top 3 trust risks.
$6,500
Two-week fixed-fee Trust Audit. I find the single riskiest number in your reads.
Guarantee
A written verdict on decisions we name together, or you don't pay.
Taking 2 new audit slots this month.