Free tool

Outcome-weighted split test calculator

Variant B completes 40% better. Before you ship it: how many of those completions closed? This ranks two form variants on completion rate and on Yield rate — closed-won per visitor — and runs a significance test on both. Expect it to tell you the outcome difference is not yet believable. That is usually the true answer.

Runs in your browser. Nothing you type is sent anywhere, stored, or logged — there is no request to send it in. Built by the team behind the dishonest dashboard.

Variant A — the control

Deals from this variant's submissions that your CRM marked won.

Variant B — the challenger

Deals from this variant's submissions that your CRM marked won.

Winner on Yield rate

Variant A

Variant B completes better. Variant A closes better. Completion rate would have shipped the other one.
Two form variants compared on completion rate and Yield rate
MetricVariant AVariant B
Completion rate8.00%11.20%
Yield rate — closed-won per visitor0.238%0.200%
Yield value — closed value per visitor$9.50$8.00
Closed-won deals1916
Completion difference
Significant

p = 0.0000. Clears the 95% bar.

Outcome difference
Not significant

p = 0.6117. Does not clear the 95% bar.

Visitors needed per variant
243,653

To call a difference this size on closed deals at 95% confidence and 80% power.

Visitors you have per variant
8,000

The smaller of the two arms, since that is what the test is limited by.

Watch this

The two metrics disagree, but not yet believably

The completion winner and the yield winner are different variants, which is interesting. The outcome difference does not clear 95% confidence, which means it is also consistent with noise. Interesting is not the same as decided.

How this is calculated

Every figure above comes from the arithmetic below. No weighting, no model, no numbers of ours mixed into yours.

The two rates

completion rate = completions ÷ visitors Yield rate = closed-won ÷ visitors Yield value = closed value ÷ visitors

Same denominator, different numerator. That is the whole idea: Yield rate is a conversion rate whose numerator is money rather than a submit event, so the two are directly comparable.

Significance — two-proportion z-test, two-sided

p̄ = (xA + xB) ÷ (nA + nB) SE = √( p̄ × (1 − p̄) × (1/nA + 1/nB) ) z = (pB − pA) ÷ SE p = 2 × (1 − Φ(|z|))

Run twice: once with x = completions, once with x = closed-won deals. Φ is the standard normal cumulative distribution, computed with the Abramowitz & Stegun 7.1.26 approximation, whose error is under 1.5 × 10⁻⁷ — far below anything we print. No continuity correction.

Sample size, per variant

n = ( z₀.₉₇₅ × √(2p̄(1−p̄)) + z₀.₈ × √(pA(1−pA) + pB(1−pB)) )² ÷ (pB − pA)² with z₀.₉₇₅ = 1.959964 and z₀.₈ = 0.8416212

95% confidence, 80% power, computed on the observed Yield rates. When the two rates are identical the answer is infinite, and we print an em dash rather than a number.

Where the defaults come from

The example is built from a case worth being able to recognise: a challenger that lifts completions by 40% and drops closed deals slightly. It is not data from anyone’s account, ours included.

The statistics are textbook and deliberately so. A two-proportion z-test is the same test every split-testing tool in the market already runs; the only thing this page does differently is point it at closed deals as well as at completions. There is nothing proprietary in the maths and there should not be.

One thing we do not do is compute significance on Yield value. Revenue per visitor is not a proportion — it is a heavily skewed distribution where one large deal moves the mean — and the z-test above would be the wrong instrument. The column is there to be read, not tested.

What this cannot tell you

It cannot tell you the test was run properly. A significance test assumes visitors were randomly assigned, the variants ran over the same period, and you decided to stop before you looked. If you have been checking daily and stopped when it looked good, the p-value here is optimistic and so is every other tool’s.

It cannot see attribution error. If your CRM credits deals to a variant by last touch, some of those closed-won counts belong to the other arm, and no amount of arithmetic downstream fixes that.

And it will not tell you a result is meaningful just because it is significant. With enough traffic a trivial difference clears 95% confidence; the size of the gap in the table matters more than the word next to the p-value.

If the answer here is “not enough closed deals,” the next question is how long it would take to get them: the time-to-outcome checker answers that, and often answers it with a no. The reasoning behind ranking variants this way is in the dishonest dashboard.

Why we built this

Every number on this page is one your form builder could have told you and didn’t.

Endpoint Forms is an open-source form builder for marketers: forms built to convert, data that goes wherever you need it, and every submission carrying what it turned out to be worth. It is not shipped yet. The waitlist is where we tell you when it is.

Waitlist