Fundusze Europejskie Województwo Łódzkie Unia Europejska
Retention

Your loyalty members outperform everyone. That proves nothing.

Members buy more in every dataset ever published, because your best customers join first. Loyalz runs the programme and measures whether it actually caused anything.

A loyalty program is a structured incentive to buy again — points, tiers, benefits, rewards. Whether yours works is a harder question than it looks, because the usual evidence does not survive scrutiny: your best customers join first, so members outperform non-members whether or not the programme changed anyone's behaviour. The only way to know is to withhold it from a randomly chosen group and compare.

01the confound

The finding everyone quotes

The chart the category sells on
Revenue per customer
£318 Loyalty members
£94 Non-members
3.4× · the number every deck quotes
Why the gap is uninformative
Already your best customers They join the programme
Programme changes behaviour They spend more
Both diagrams produce the same chart on the left. Nothing in the chart tells you which one you are looking at.

This is the chart the whole category sells on, and it is real. It is also uninformative, because nobody assigned those customers to those groups — they chose.

Your best customers join the loyalty programme because they are your best customers. Their loyalty caused their membership, not the other way round.

Every member-versus-non-member ratio you have ever been shown is consistent with a programme that does nothing at all. We sell a loyalty product and we are telling you this, because the alternative is selling you a number we know is uninformative.

02measurement

The design that does answer it

Six rules, one design
ALL CUSTOMERS randomised once
ARM A · 50% Programme running — points, tiers, quests, benefits
ARM B · 50% Whole programme withheld — no points, no tier benefits the comparison
TIMELINE design frozen cycle 1 — pulled forward? cycle 2 — created?
01
Randomise customers, not orders

An order-level split leaks — the same person lands in both arms and the effect washes out.

02
Assign once, never reassign

A customer who moves arms mid-experiment belongs to neither. Store the arm, the timestamp and the seed.

03
Withhold the whole programme

Withholding points while leaving tier benefits running measures a mechanic, not a programme.

04
Run two purchase cycles

One cycle shows whether you pulled a purchase forward. Two shows whether you created one.

05
Freeze the design first

Metric, window and cut-offs committed before the first assignment. Deciding after you look is how a null becomes a positive.

06
Measure per assigned customer

Dividing by actives quietly drops everyone the programme failed to reach — the group it was supposed to affect.

03the arithmetic

What it costs in customers

Revenue per customer is skewed, so this needs more customers than instinct suggests. Per arm, two-sided, 80% power:

Revenue variability detect +5% detect +10% detect +20%
Narrow catalogue 6,280 1,570 390
Typical store 14,130 3,530 880
Wide catalogue 25,120 6,280 1,570
With a VIP tail 39,240 9,810 2,450
HOLDOUT SIZE VS TOTAL COST
50 / 50 split 1.0×
10 / 90 split 2.8×
Total customers needed. The smaller arm limits the comparison, so a cautious split costs more, not less.
If your store is too small to detect the effect you care about, the honest answer is that you cannot measure it — and running the test anyway produces a number that looks like an answer and is noise.

The instinct is to hold out 5 or 10% because withholding the programme costs revenue. That makes the test more expensive, not less: a 10/90 split needs about 2.8× the total customers of an even split, because the smaller arm limits the comparison.

04the cost

Rewards are a cost, and almost nobody nets them

Worked example
100% of claimed programme revenue
38% 62%
Genuinely incremental38 Margin handed to customers who would have returned anyway62

Worked example of the shape the category never draws. A reward redeemed by someone who would have repurchased anyway is not marketing spend — it is margin handed over for nothing, and the resulting order still gets counted as loyalty-driven revenue.

What to subtract before calling something profit Retention, measured
05the programme

Running it, once you can measure it

Illustrative dashboard
TIERS WITH EXPERIENCE LEVELS LVL 1 → 5
Level 1 Starter 01 STARTER
Level 2 Explorer 02 EXPLORER
Level 3 Elevated 03 ELEVATED
Level 4 Premier 04 PREMIER
Level 5 Icon 05 ICON
EARNED ON SALES SOCIAL ACTIVITY TIME-LIMITED RANKINGS arm A only, during a test
Retention · last 90 days
Repeat rate
31.4%
+2.2pp
Revenue per customer
£186
+£11
Reward cost
£14,280
+9.1%
A COST, NOT A RESULT
Members
8,410
+412
Contribution margin
£96,700
+4.4%
Redemption rate
22.8%
−1.4pp
Second-order rate
18.9%
+0.8pp

Illustrative layout. Points, cashback, tiers with experience levels, quests and benefit collections — with reward cost sitting on the same screen as the revenue it is supposed to have produced, which is the only arrangement that lets you catch a programme paying for its own results.

RFM segmentation tells you who is worth keeping before you spend anything on keeping them. Recency, frequency and monetary value are the three facts about a customer that predict the next order, and they need no model to compute.

Customer retention, measured What a customer is worth
06limits

What this does not do

THE LIKELIEST FINDING
No detectable effect A measured lift A measured loss
We publish a null with the same prominence as a win. We are not putting a probability on which you will get — that would be the same kind of number this page argues against.
SMALLER CLAIM
It does not prove your programme works
It gives you the design that can, and the arithmetic that says whether your store is big enough to run it.
SMALLER CLAIM
CLV here is measured, not projected
Trailing average revenue per identified customer over 24 months. A measurement of the past, labelled as one.
HARD FLOOR
Guest checkouts are excluded
There is nothing to identify them with. Their orders still count toward revenue.
HARD FLOOR
Identity is per store source
A buyer appearing under two identifiers is two customers until those records are reconciled.
YOUR CALL
A holdout has a real cost
It withholds a benefit from real customers for the length of the test. That decision is yours, not ours.
YOUR CALL
Most honest results are null
The likeliest finding is no detectable effect at your sample size. We will publish that with the same prominence as a win.

Find out whether your store is big enough to measure this

Bring your order history. We compute the effect size you could actually detect, before you commit to withholding anything from anyone.

Book the twenty minutes