I instrumented the paywall funnel myself: paywall impression, plan selected, purchase, first cycle cancel or refund. Each event carried the experiment arm and the acquisition source through the MMP and analytics, so every arm's subscribers tied back to paid versus organic cohorts.
FIELD REPORT 05 · SUBSCRIPTION · MATCH GROUP
The onboarding
paywall test
Product passed on an onboarding paywall, so I ran the test myself on a Match Group dating app with well over 50,000 monthly active users. Three arms, a capped exposure, a loss ceiling and a kill line, all written before launch.
net of churn
Illustrative screens, drawn for this page. The three arms match until the screen after the first likes and matches. Net is indexed to control = 100.
There was no ask in onboarding at all.
The app was a Match Group dating app: a free tier plus a paid subscription, with well over 50,000 monthly active users. I owned the paid and organic acquisition cohorts with revenue accountability, roughly $2,000,000 a month in spend.
Onboarding was the most valuable real estate in the product. Every new user went through it, and it never asked for money. New users met the subscription later, from inside the product.
I pitched an onboarding paywall to the product team. They passed. So I ran the experiment myself: the test design, the tracking, the downside controls and the read.
The bet: a paywall in onboarding is the classic way to lose users at the door, so it went after the first value moment, where the ask makes sense. The subscribers it won had to stay past their first billing cycle. That is why the result is read net of churn.
Same plans, same prices. Only the ask moved.
New users only, randomized at sign up, sticky per account and flag driven. Tested on web first. Every lane below is drawn to scale.
- CONTROL 70%
- VARIANT A 15%
- VARIANT B 15%
- FIRST VALUE MOMENT
- ONBOARDING PAYWALL
Worst case was bounded before day one.
Every control had a number and an action, written before launch. None of them needed a meeting to fire.
30% of new users across both test arms, never more. It bounded worst case revenue at risk before day one.
The test arms combined could give up at most 2% of monthly subscription revenue against their control equivalent. Checked weekly. Hitting it turned the arm off.
Day 7 checkpoint: any arm more than 10% relative below control on first time paid subscribers per new user was turned off that day, with no waiting for a recovery.
Six metrics, each with a tolerance band against control. A tripped band put the arm under review at the weekly check.
Weekly review against the ceiling. Flag driven rollback. No peeking based ship calls: the ship decision waited for the read date.
12% below control on first time paid subscribers per new user is past the kill line. The arm goes off at the day 7 checkpoint.
The cap, the ceiling and the kill line are the rules as written. The inputs are illustrative: the real new user share of revenue stayed private. The shortfall applies to the two test arms together. The kill line reads conversion at day 7, so a loss that hides in churn is caught by the guardrails and the weekly ceiling check instead. That is how B went off at week 2.
Four reads, one ship call.
Scrub the six weeks or press play. Each stop shows the allocation, each arm's status and what was decided. A revenue ceiling check ran at the end of every week.
- ENROLMENT, WEEKS 1 TO 4
- MATURATION, WEEKS 5 AND 6
- VARIANT B, OFF AT WEEK 2
- THE FOUR READS
- WEEKLY CEILING CHECK
A ships. B stays off.
113 net of first cycle churn, every guardrail inside its band. A shipped to 100% by flag, then rolled out to iOS and Android.
| METRIC | BAND | VARIANT A | VARIANT B |
|---|---|---|---|
| First cycle churncancel or refund inside the first billing period | inside band vs control | HELD | TRIPPED |
| Paid term mix and average subscription length | inside band vs control | HELD | TRIPPED |
| Refund and chargeback rate | inside band vs control | HELD | OFF AT WK 2 |
| Onboarding completion | inside band vs control | HELD | OFF AT WK 2 |
| Day 1 and day 7 retention | inside band vs control | HELD | OFF AT WK 2 |
| Support contact rate | inside band vs control | HELD | OFF AT WK 2 |
B's plan mix shifted shorter and its early cancels rose. Those two bands tripped at the week 2 checkpoint and B was turned off inside the revenue ceiling. B's other rows were never read at week 6 because the arm was already off.
A's net beat its gross.
Indexed to control = 100 on first time paid subscribers per 1,000 new users. Gross counts every first time paid subscriber in the window, wherever they subscribed. Net is what was left after first cycle churn.
Gross is what a dashboard shows. Net is what stays.
Read each line left to right. A gained 3 points after churn and B lost 13. On gross the two variants were a point apart. On net they were 17 points apart, on either side of control.
| ARM | GROSS | NET OF CHURN | VERDICT |
|---|---|---|---|
| Control | 100 | 100 | baseline |
| Variant A, value first | 110 | 113 | SHIPPED TO 100% |
| Variant B, reordered ladder | 109 | 96 | OFF AT WK 2 |
Graded on gross, B looked like a near tie with A: 109 against 110. Graded on net, A added 13 percent more first time paid subscribers and B finished 4 percent below control. The whole difference sat in the first billing cycle, in plan term and early cancels. That is why the read waited two weeks after enrolment closed.
Fewer of A's subscribers cancelled in the first cycle than control's later, in app subscribers. Asking right after the first value moment attracted better qualified subscribers.
B converted more at the screen. The plan mix shifted shorter and early cancels rose, so net fell below control. B was turned off at the week 2 checkpoint, inside the ceiling. The downside controls did their job.
A shipped to 100% by flag after the read date, then rolled out to iOS and Android. The lift held in the post launch cohort. Retention, refunds and plan mix stayed inside their bands, and subscription revenue from the cohort moved with the subscriber count.
All figures are indexed to control = 100, from my readout at the time with the analyst. Percentages on this page are computed from those indexes. Raw numbers stayed with Match Group.
Five lanes. Mine ran from the pitch to the ship call.
I ran it myself. I made the first designs, in house designers finished them, and developers shipped the code. The tracking, the rules and the read were mine, with an analyst on the readout.
Hypothesis, variants, allocation rule, exposure cap, loss ceiling, kill trigger, guardrail set and the read date, all written before launch. Tested on web first, then the winner went to iOS and Android.
Cohort definitions, the net of churn calculation, the paid and organic split read apart, and the reconciliation of subscriber counts to subscription revenue.
I worked with an analyst on the readout at the time. Two people, one loop.
I run the whole loop myself, from instrumentation to readout, with SQL, Python, Claude Code and a data lake I build, faster than the two of us did then. With AI I would also take on the design and the code.
SAME FAMILY, DIFFERENT APPAt Picniic, a subscription family app, a free trial redesign I ran drove a sustained 25 percent lift in subscription revenue net of churn.
The paywall that converts best is rarely the one that retains best.
When pricing is weekly with a 3 day trial, the first month of renewals decides a paid user's lifetime value, and trial starts are the wrong scoreboard. This is the one I would run.
| VARIANT | COHORT | INSTALLS | TRIAL STARTS | 30 DAY NET REVENUE PER INSTALL | VS CONTROL | GUARDRAILS | DECISION |
|---|---|---|---|---|---|---|---|
| control + each arm | paid and organic, read apart | counted | shown, never graded on | the grade | indexed, control = 100 | held or tripped, per band | WATCH · KILL · SHIP |
Above control on 30 day net revenue per install at the read date, guardrails inside band, in paid and organic alike.
Below the kill trigger at the first checkpoint, or the loss ceiling reached, in either cohort. Both set before launch.
Everything else until the read date. No peeking based ship calls.
Grade on 30 day net revenue per install. Trial starts get shown and never graded.
Set the loss ceiling and the kill trigger before launch.
Read paid and organic cohorts separately.
Test substantially different experiences against each other. Small tweaks can wait.
With enough install volume, the answers arrive in weeks.