FIELD REPORT 05 · SUBSCRIPTION · MATCH GROUP

The onboarding
paywall test

Product passed on an onboarding paywall, so I ran the test myself on a Match Group dating app with well over 50,000 monthly active users. Three arms, a capped exposure, a loss ceiling and a kill line, all written before launch.

+13%first time paid subscribers,
net of churn
ONE ONBOARDING, THREE ARMS
CONTROL · 70%No paywall in onboarding. The offer comes later, in product.NET 100
VARIANT A · 15%A paywall after the first value moment. Same plans, same prices.NET 113 · SHIPPED
VARIANT B · 15%Same placement. Shortest term first, price per week.NET 96 · OFF WK 2

Illustrative screens, drawn for this page. The three arms match until the screen after the first likes and matches. Net is indexed to control = 100.

READ THE REPORT

There was no ask in onboarding at all.

The app was a Match Group dating app: a free tier plus a paid subscription, with well over 50,000 monthly active users. I owned the paid and organic acquisition cohorts with revenue accountability, roughly $2,000,000 a month in spend.

Onboarding was the most valuable real estate in the product. Every new user went through it, and it never asked for money. New users met the subscription later, from inside the product.

I pitched an onboarding paywall to the product team. They passed. So I ran the experiment myself: the test design, the tracking, the downside controls and the read.

ROLEGrowth owner, paid + organic cohorts
SPEND~$2,000,000 a month
APPMatch Group (S&P 500) dating app, 50,000+ MAU
WHENNov 2019 to Mar 2024
The bet: a paywall in onboarding is the classic way to lose users at the door, so it went after the first value moment, where the ask makes sense. The subscribers it won had to stay past their first billing cycle. That is why the result is read net of churn.

Same plans, same prices. Only the ask moved.

New users only, randomized at sign up, sticky per account and flag driven. Tested on web first. Every lane below is drawn to scale.

● THE SPLITNEW USERS · LANE WIDTH = SHARE OF NEW USERS
REGISTER FIRST LIKES + MATCHES NEXT SCREEN NEW USERS RANDOMIZED AT SIGN UP STICKY PER ACCOUNT · FLAG DRIVEN 70% CONTROL 15% VARIANT A 15% VARIANT B A + B = 30%, THE EXPOSURE CAP PAYWALL PAYWALL, LADDER REORDERED NO ASK INTO THE APP OFFER LATER, IN PRODUCT NET 100 IN APP LADDER NET 113 · SHIPPED SHORTEST FIRST, PER WEEK NET 96 · OFF WK 2 NEW USERS 70% 15 15 REGISTER VALUE MOMENT NEXT SCREEN NO ASK INTO THE APP NET 100 PAYWALL NET 113 SHIPPED LADDER NET 96 OFF WK 2 A + B = 30% CAP
  • CONTROL 70%
  • VARIANT A 15%
  • VARIANT B 15%
  • FIRST VALUE MOMENT
  • ONBOARDING PAYWALL
CONTROL · 70%NO ONBOARDING PAYWALLNew users went straight into the app. The subscription was offered later, from inside the product, so onboarding never asked for money.
VARIANT A · 15% · VALUE FIRSTASK AFTER THE VALUE MOMENTA paywall added to onboarding, after the user's first likes and matches are shown. Same plans and prices as the in app offer.
VARIANT B · 15% · VALUE FIRST + LADDERSAME PLACE, NEW LADDERSame placement as A. The plan ladder reordered to lead with the shortest term, price framed per week.
● DESIGNWRITTEN BEFORE LAUNCH
ELIGIBLENEW USERS ONLYExisting users excluded, so renewal cohorts were untouched.
ASSIGNMENTSTICKY PER ACCOUNTRandomized at sign up. Flag driven, so any arm could go back to control in minutes.
POWER5% RELATIVE AT 95%Sized to detect a 5% relative change in first time paid subscription rate. Tens of thousands of new users per arm.
WINDOW4 WEEKS + 2 WEEKSFour weeks of enrolment, two of maturation. The read date was fixed before launch.

Worst case was bounded before day one.

Every control had a number and an action, written before launch. None of them needed a meeting to fire.

01Exposure cap

30% of new users across both test arms, never more. It bounded worst case revenue at risk before day one.

02Revenue at risk ceiling

The test arms combined could give up at most 2% of monthly subscription revenue against their control equivalent. Checked weekly. Hitting it turned the arm off.

03Kill trigger

Day 7 checkpoint: any arm more than 10% relative below control on first time paid subscribers per new user was turned off that day, with no waiting for a recovery.

04Guardrail metrics

Six metrics, each with a tolerance band against control. A tripped band put the arm under review at the weekly check.

05Process

Weekly review against the ceiling. Flag driven rollback. No peeking based ship calls: the ship decision waited for the read date.

● LOSS LIMIT CONSOLERULES AS WRITTEN · ILLUSTRATIVE INPUTS
TRY A CASE
30%
12%
WHERE THE LOSS SHOWS UP
20%
ALLOCATION OF NEW USERS
CONTROL 70%A 15%B 15%
30% CAP
REVENUE AT RISK VS THE 2% CEILING
2% CEILING
0%1%2%3%
DAY 7 KILL CHECK, CONVERSION VS CONTROL
KILL LINE, 10% BELOW
0%10%20%30%40% BELOW
KILL Turned off at day 7.

12% below control on first time paid subscribers per new user is past the kill line. The arm goes off at the day 7 checkpoint.

REVENUE AT RISK0.72%Exposure × shortfall × new user share of revenue.
HEADROOM TO THE CEILING1.28 ptsThe 2% ceiling minus revenue at risk.
SHORTFALL THAT HITS THE CEILING33.3%At this exposure and revenue share.

The cap, the ceiling and the kill line are the rules as written. The inputs are illustrative: the real new user share of revenue stayed private. The shortfall applies to the two test arms together. The kill line reads conversion at day 7, so a loss that hides in churn is caught by the guardrails and the weekly ceiling check instead. That is how B went off at week 2.

Four reads, one ship call.

Scrub the six weeks or press play. Each stop shows the allocation, each arm's status and what was decided. A revenue ceiling check ran at the end of every week.

● EXPERIMENT TIMELINEREAD DATE FIXED BEFORE LAUNCH
USERS ARM A ARM B CHECKS ENROLMENT MATURATION
  • ENROLMENT, WEEKS 1 TO 4
  • MATURATION, WEEKS 5 AND 6
  • VARIANT B, OFF AT WEEK 2
  • THE FOUR READS
  • WEEKLY CEILING CHECK
WEEK 6 · READ DATE

A ships. B stays off.

113 net of first cycle churn, every guardrail inside its band. A shipped to 100% by flag, then rolled out to iOS and Android.

NEW USERS
Enrolment closed.
VARIANT ASHIP113 net, every band held.
VARIANT BOFFOff since the week 2 checkpoint.
● GUARDRAILSTOLERANCE BAND VS CONTROL · WHERE EACH ARM ENDED
The six guardrail metrics and where each test arm ended
METRICBANDVARIANT AVARIANT B
First cycle churncancel or refund inside the first billing periodinside band vs controlHELDTRIPPED
Paid term mix and average subscription lengthinside band vs controlHELDTRIPPED
Refund and chargeback rateinside band vs controlHELDOFF AT WK 2
Onboarding completioninside band vs controlHELDOFF AT WK 2
Day 1 and day 7 retentioninside band vs controlHELDOFF AT WK 2
Support contact rateinside band vs controlHELDOFF AT WK 2

B's plan mix shifted shorter and its early cancels rose. Those two bands tripped at the week 2 checkpoint and B was turned off inside the revenue ceiling. B's other rows were never read at week 6 because the arm was already off.

A's net beat its gross.

Indexed to control = 100 on first time paid subscribers per 1,000 new users. Gross counts every first time paid subscriber in the window, wherever they subscribed. Net is what was left after first cycle churn.

● GROSS TO NETWEEK 6 READ · CONTROL = 100
9095100105110115120 GROSS NET OF CHURN 100 100CONTROL 109 96VARIANT B -13 PTS 110 113VARIANT A +3 PTS

Gross is what a dashboard shows. Net is what stays.

Read each line left to right. A gained 3 points after churn and B lost 13. On gross the two variants were a point apart. On net they were 17 points apart, on either side of control.

A, NET VS CONTROL+13%
B, NET VS CONTROL-4%
A, SHARE OF GROSS KEPT VS CONTROL+2.7%
B, SHARE OF GROSS KEPT VS CONTROL-11.9%
Result by arm, indexed to control = 100 on first time paid subscribers per 1,000 new users
ARMGROSSNET OF CHURNVERDICT
Control100100baseline
Variant A, value first110113SHIPPED TO 100%
Variant B, reordered ladder10996OFF AT WK 2
WHY NET MATTERS

Graded on gross, B looked like a near tie with A: 109 against 110. Graded on net, A added 13 percent more first time paid subscribers and B finished 4 percent below control. The whole difference sat in the first billing cycle, in plan term and early cancels. That is why the read waited two weeks after enrolment closed.

WHY A'S NET BEAT ITS GROSS

Fewer of A's subscribers cancelled in the first cycle than control's later, in app subscribers. Asking right after the first value moment attracted better qualified subscribers.

WHAT HAPPENED TO B

B converted more at the screen. The plan mix shifted shorter and early cancels rose, so net fell below control. B was turned off at the week 2 checkpoint, inside the ceiling. The downside controls did their job.

AFTER THE READ

A shipped to 100% by flag after the read date, then rolled out to iOS and Android. The lift held in the post launch cohort. Retention, refunds and plan mix stayed inside their bands, and subscription revenue from the cohort moved with the subscriber count.

All figures are indexed to control = 100, from my readout at the time with the analyst. Percentages on this page are computed from those indexes. Raw numbers stayed with Match Group.

Five lanes. Mine ran from the pitch to the ship call.

I ran it myself. I made the first designs, in house designers finished them, and developers shipped the code. The tracking, the rules and the read were mine, with an analyst on the readout.

MEgrowth owner PITCHPitched it DESIGNFirst designs TRACKING + RULESEvents, caps, kill line RUNRan the test READOUTNet of churn read ROLLOUTThe ship call
PRODUCT TEAM PITCHPassed
DESIGNERSin house DESIGNFinished the designs
DEVELOPERS CODEShipped the code ROLLOUTiOS + Android
ANALYST READOUTReadout, with me
L1
THE TRACKING

I instrumented the paywall funnel myself: paywall impression, plan selected, purchase, first cycle cancel or refund. Each event carried the experiment arm and the acquisition source through the MMP and analytics, so every arm's subscribers tied back to paid versus organic cohorts.

L2
THE EXPERIMENT DESIGN

Hypothesis, variants, allocation rule, exposure cap, loss ceiling, kill trigger, guardrail set and the read date, all written before launch. Tested on web first, then the winner went to iOS and Android.

L3
THE READOUT

Cohort definitions, the net of churn calculation, the paid and organic split read apart, and the reconciliation of subscriber counts to subscription revenue.

THEN

I worked with an analyst on the readout at the time. Two people, one loop.

NOW

I run the whole loop myself, from instrumentation to readout, with SQL, Python, Claude Code and a data lake I build, faster than the two of us did then. With AI I would also take on the design and the code.

SQLPYTHONCLAUDE CODEDATA LAKEMMP + ANALYTICSFEATURE FLAGS

SAME FAMILY, DIFFERENT APPAt Picniic, a subscription family app, a free trial redesign I ran drove a sustained 25 percent lift in subscription revenue net of churn.

The paywall that converts best is rarely the one that retains best.

When pricing is weekly with a 3 day trial, the first month of renewals decides a paid user's lifetime value, and trial starts are the wrong scoreboard. This is the one I would run.

● SCOREBOARD STRUCTUREEVERY PAYWALL AND ONBOARDING VARIANT
Scoreboard structure: the columns every variant is graded on
VARIANTCOHORTINSTALLSTRIAL STARTS30 DAY NET REVENUE PER INSTALLVS CONTROLGUARDRAILSDECISION
control + each armpaid and organic, read apartcountedshown, never graded onthe gradeindexed, control = 100held or tripped, per bandWATCH · KILL · SHIP
SHIP

Above control on 30 day net revenue per install at the read date, guardrails inside band, in paid and organic alike.

KILL

Below the kill trigger at the first checkpoint, or the loss ceiling reached, in either cohort. Both set before launch.

WATCH

Everything else until the read date. No peeking based ship calls.

01

Grade on 30 day net revenue per install. Trial starts get shown and never graded.

02

Set the loss ceiling and the kill trigger before launch.

03

Read paid and organic cohorts separately.

04

Test substantially different experiences against each other. Small tweaks can wait.

With enough install volume, the answers arrive in weeks.

COMPANION · FIELD REPORT 06 · PRESENTATIONThe paywall test walkthroughThe same test as three boards: the net revenue readout, the build, and the 30 day scoreboard I run now.
NEXT · FIELD REPORT 07 · APP ACQUISITIONThe Female Profile EventApp campaigns offer no gender targeting. I built it out of a conversion event that fired only when a woman created a profile.

Happy to walk through any of these live.