τ̂ = 1/n Σ [ μ̂₁(Xᵢ) − μ̂₀(Xᵢ) ]
Model the outcome.
Predict what each subject does treated and untreated, average the gap. Correct only if the outcome model is, and silent when it is not.
Method
Every risk score puts them at the top of the call list. Nexyron computes a different quantity, how much your action changes the outcome, and it puts somebody else there.
The subtraction
Retention improved nine points after the save desk was funded. Finance asks what the money bought. The answer is one line of algebra, and it is not nine.
What the report gives you
Δ = E[ Y | W=1 ] − E[ Y | W=0 ]
One subtraction. Nothing in it can be checked, and nothing in it separates the action from the choosing.
What Nexyron gives you
Δ = E[ Y(1) − Y(0) | W=1 ] + ( E[ Y(0) | W=1 ] − E[ Y(0) | W=0 ] )
└── the action ──┘ └──── the targeting ────┘
and the right term is zero only if
W ⫫ Y(0) assignment was random
which your programme is not, so
Δ > τ the reported figure overstates the actionThe same subtraction, carried through to the point where it says which half you paid for.
Everyone reports the sum. Nexyron reports the two parts, so you know which one you paid for.
The right-hand term is zero only when treatment was assigned at random. Save desks call the customers most likely to leave, investigators open the files that look wrong, engineers service the assets that look tired. The selection is the point of the operation, and it is also what puts that term in the equation.
More rows will not fix it. A larger sample makes a biased figure more precise, never more correct.
| reported | the action | the targeting | |
|---|---|---|---|
| retention programme | +9.0 pts | +2.4 | +6.6 |
| fraud delisting | −31% | −9% | −22% |
| maintenance brought forward | −18% | −4% | −14% |
Three ways to compute it
Every estimate rests on a model being correct, and you will not know which of yours is wrong. One estimator survives that, and it is the one that should produce the figure you sign.
What the platform returns
uplift = 2.4 pts
One estimator, unnamed, resting on one model being correct. If that model is wrong nothing in the output says so.
What Nexyron returns
plugin_ate = 2.61 trusts the outcome model
ipw_ate = 2.38 trusts the treatment model
dr_ate = 2.44 trusts either one
spread = 0.23 three assumptions, one answer
where the doubly robust form is
τ̂dr = τ̂plug + 1/n Σ [ Wᵢ(Yᵢ−μ̂₁)/ê − (1−Wᵢ)(Yᵢ−μ̂₀)/(1−ê) ]
└──── the residuals the model left ────┘Three estimators that fail differently, in one row. Their agreement is the evidence; their spread is the warning.
One estimate cannot be checked by anything. Three that fail in different ways check each other.
τ̂ = 1/n Σ [ μ̂₁(Xᵢ) − μ̂₀(Xᵢ) ]
Model the outcome.
Predict what each subject does treated and untreated, average the gap. Correct only if the outcome model is, and silent when it is not.
τ̂ = 1/n Σ [ WᵢYᵢ / ê(Xᵢ) − (1−Wᵢ)Yᵢ / (1−ê(Xᵢ)) ]
Model the treatment.
Reweight until the treated resemble everyone. Correct only if the treatment model is, and the variance runs away as ê approaches zero.
τ̂ = τ̂plug + 1/n Σ [ Wᵢ(Yᵢ−μ̂₁)/ê − (1−Wᵢ)(Yᵢ−μ̂₀)/(1−ê) ]
Either one will do.
The plug-in estimate plus the residuals it left behind. Right outcome model and the correction averages to nothing. Right treatment model and the correction repairs the other. One suffices.
| outcome model right | treatment model right | both wrong | |
|---|---|---|---|
| plug-in | consistent | biased | biased |
| IPW | biased | consistent | biased |
| doubly robust | consistent | consistent | biased |
All three arrive in the same row. Estimators resting on different assumptions landing together is evidence the assumptions hold. When they diverge, one of the models is wrong and the spread says roughly by how much.
Doubly robust estimation: Robins, Rotnitzky and Zhao, Journal of the American Statistical Association.
Whether to believe it
An estimate on its own cannot be audited, and an unauditable number is the one that fails in the meeting where it matters. These come back in the same result, unasked.
What lands on the slide
+2.4 pts
A point estimate. No way to ask whether the groups were comparable, how much sample survived the adjustment, or how wide the answer really is.
What lands in the result
dr_ate = 2.44 bootstrap [ p05, p95 ] = [ 0.31 , 4.52 ] the interval, not the point max_abs_smd = 0.62 before weighting max_abs_weighted_smd = 0.29 after, and still imbalanced treated_ess = 2,910 of 3,400 control_ess = 1,240 of 14,600 the cost of the weighting overlap_count = 15,760 near_one_count = 2,240 no comparison exists here
Nine numbers instead of one, and four of them can veto the fifth.
They hand you a number. Nexyron hands you the number and the four tests that can overrule it.
d = ( x̄₁ − x̄₀ ) / √( (s₁²+s₀²)/2 )
Per feature, before weighting and again after. The first says how different the groups were, the second whether the adjustment fixed it. |d| > 0.1 is imbalance.
ESS = ( Σ wᵢ )² / Σ wᵢ²
Weighting concentrates the answer onto fewer subjects. Fourteen thousand rows can carry the information of twelve hundred, and the row count will never tell you.
#{ i : ε < ê(Xᵢ) < 1−ε }How much of the book sits where a comparison exists at all. A programme that treats everyone at risk has removed its own control group.
[ q₀.₀₅ , q₀.₉₅ ]
An interval across resamples, not a point. An uplift whose interval spans zero has told you something, and it is not what the centre said.
Reporting these can only make an estimate look weaker than reporting the row count alone. That is why they are uncommon, and why a number that comes with them is worth more.
Propensity scores: Rosenbaum and Rubin, Biometrika. Balance diagnostics: Austin, Statistics in Medicine, 2009. Effective sample size: Kish. Bootstrap intervals: Efron, The Annals of Statistics.
What it is worth
What was the action worth, how much of the book can answer, and where does the same budget land instead. Every industry asks them in its own vocabulary.
Illustrative figures showing the shape of the difference, not measurements. The size of each gap is a property of your data, and the engine measures it rather than assuming it.
Did last year's delisting prevent anything?
Which scenario can be turned down, and what would it cost?
When did it start, and who else was in it?
Did the integrity pact work, or did the behaviour move?
Did the control we recommended last year work?
Do reminder notices change anything?
Did pulling the work forward prevent a failure?
Did dual-sourcing achieve anything?
Did the retention spend cause the improvement?
The deciding quantity
A predictive model ranks by how likely the outcome is. A decision needs to know how much acting changes it. These are different functions and they do not order the same list.
The subject most likely to leave is frequently the subject least movable. There is also a group that leaves because you called them, and no risk score has ever separated the two.
How the list is ordered today
rank by E[ Y | X=x ] how likely they are to go
The people at the top are the people most certain to leave. Nothing in that quantity is about whether a call changes anything.
How Nexyron orders it
τ(x) = E[ Y(1) − Y(0) | X=x ] how much the call changes it
at risk, unreachable 0.82 → 0.82 τ = 0.00
persuadable 0.31 → 0.12 τ = 0.19
safe anyway 0.06 → 0.05 τ = 0.01
annoyed by contact 0.14 → 0.22 τ = −0.08
then choose who, under the budget you already have
max Σ τ(Xᵢ)
i∈S
subject to |S| ≤ k
ε < ê(Xᵢ) < 1−ε
and choose k where V(k) stops payingA ranking, a segmentation and a budget, from one quantity the risk score does not contain.
A risk score ranks who is leaving. Nexyron ranks who you can keep, which is a different list and the only one worth calling.
Who to call becomes a constrained maximisation, and how many becomes a question with an answer rather than a number somebody was given.
X-learner for unequal treated and control groups: Kunzel, Sekhon, Bickel and Yu, Proceedings of the National Academy of Sciences, 2019. Uplift gain curves: Radcliffe, Direct Marketing Analytics Journal, 2007.
Before you spend it
You are asked to approve a different policy: call the top two thousand instead of the top one thousand, exit a different band. The evidence is a year of decisions made under the old rule, and that is enough.
How the decision gets made now
run it for a quarter and compare
A quarter of spend to find out, and the comparison at the end has the same selection problem as the one you started with.
What Nexyron computes from the log you already have
wᵢ = π(aᵢ|xᵢ) / p(aᵢ|xᵢ) how much the new rule agrees
IPS = 1/n Σ rᵢ wᵢ unbiased, high variance
SNIPS = Σ wᵢrᵢ / Σ wᵢ self-normalised
DR = SNIPS + reward-model residual
and then the check that matters
ε = 0.01 IPS 1.94 SNIPS 1.31 DR 1.28
ε = 0.05 IPS 1.42 SNIPS 1.29 DR 1.27
ε = 0.10 IPS 1.30 SNIPS 1.28 DR 1.26
IPS moves by a third. SNIPS and DR do not.The proposed rule scored at 1.27 times the current one, and shown to be stable rather than an artefact of a few extreme weights.
The usual answer costs a quarter of spend to find out. Nexyron answers it today, from decisions you have already made.
A proposed rule scored at 1.27 times the one you run today, stable across the range. That stability is the check that separates a real improvement from an artefact of a few extreme weights.
Self-normalised estimator: Swaminathan and Joachims, Neural Information Processing Systems, 2015.
What the usual method cannot see
Community detection is how a network is broken into groups, and the standard objective has a proved blind spot. Its shape is the shape of an organised ring, which is why rings survive the screen bought to find them.
What the standard screen optimises
max Q = 1/2m Σ [ Aᵢⱼ − kᵢkⱼ/2m ] δ(cᵢ,cⱼ)
Modularity. It is the default in every graph tool, and it carries a limit that nobody puts on the box.
What that limit costs you, and what Nexyron does instead
a community with ℓ < √(2m) internal edges cannot be resolved
m = 50,000 floor 316 a ring of 11 is invisible
m = 200,000 floor 632 a ring of 11 is invisible
m = 1,000,000 floor 1,414 a ring of 11 is invisible
the floor grows with your book, so the bigger you are the blinder it is
so the engine also fits, without that floor,
degree-corrected block models
nested block modelsWhich method found a group is part of the finding, because the methods disagree in ways that are themselves informative.
The standard screen is mathematically incapable of seeing a ring of eleven. Nexyron carries methods that are not.
A community with fewer than √(2m) internal edges cannot be resolved. Below that size it is absorbed into a larger one however tightly knit it is. No implementation escapes this, because it follows from what modularity maximises.
At fifty thousand edges the floor is 316. At a million it is 1,414. The larger the book, the larger a group has to be before the method can see it, and a ring of eleven is never large enough.
Degree-corrected and nested block models fit a generative structure rather than maximising modularity, and do not share the blind spot. Which method found a group is part of the finding.
Resolution limit: Fortunato and Barthelemy, PNAS, 2007. Degree-corrected block models: Karrer and Newman, Physical Review E, 2011.
When they ask again
A regulator, a board or an auditor asks you to reproduce a figure from six months ago. Your data has been corrected since. Recompute it and you have quietly answered a different question under the old one's name.
What a rerun gives you
τ̂ recomputed on today's data
The same query, a different answer, and nothing in the output explaining why.
What Nexyron gives you
τ̂ ( at , system_at )
at when the world was in that state
system_at when the database believed it
so the March figure, asked for in September, is
τ̂ ( 2026-03-31 , 2026-03-31 ) what you reported
τ̂ ( 2026-03-31 , 2026-09-30 ) the same period, corrected data
two questions, and the difference between them is visibleWithout both clocks every rerun is a new analysis wearing the old one's name.
Their rerun quietly answers a different question. Nexyron's answers the same one, six months later.
Both clocks are arguments to the estimate. Without them every rerun is a new analysis wearing the old one's name, and the difference is invisible to whoever asked.
One where an action was credited with an improvement. Ask how much was the action and how much was the targeting, which estimator produced it, what the balance was after weighting, and how wide the interval ran.