AI & ML B10 / 09

Holdout Group in iGaming: Definition, How They Measure ML Model Impact and Why Operators Need Them

A Holdout Group is a randomly-selected subset of customers explicitly excluded from an ML-driven intervention, used as a control to measure the actual impact of the intervention. Without holdout groups, operators can only see what happened to customers who received an intervention…

iGaming Glossary · Category: AI & ML · Relevant for: ML, CRM, Marketing, Analytics

iGaming GlossaryMLCRMMarketingAnalytics

TL;DR

A Holdout Group is a randomly-selected subset of customers explicitly excluded from an ML-driven intervention, used as a control to measure the actual impact of the intervention. Without holdout groups, operators can only see what happened to customers who received an intervention, not whether the intervention caused better outcomes than no intervention would have. Holdouts answer the counterfactual question that observational data cannot: would this customer have done the same thing without our action? They are essential for honest ML impact measurement.

Mechanics 02

How it works

Holdout group methodology is conceptually simple but operationally demanding:

  • Define the intervention: a CRM campaign, a recommendation system, a bonus offer, a churn prevention action.
  • Define the eligible population: customers the intervention would target.
  • Randomly assign a subset (typically 5-10%) to the holdout: explicitly excluded from the intervention.
  • Run the intervention against the treatment group; the holdout group receives no treatment (or sometimes a baseline alternative).
  • Measure outcomes for both groups: GGR, retention, conversion, whatever the intervention aims to influence.
  • Compare: the difference between groups is the causal impact of the intervention.

Holdout designs come in several variants:

  • Universal holdouts: small persistent holdout never receiving any CRM intervention, measuring aggregate CRM impact.
  • Campaign-specific holdouts: holdouts per individual campaign measuring that campaign's impact.
  • Model-specific holdouts: holdouts excluded from specific ML-driven recommendations to measure model impact.
  • Staggered rollout: introducing interventions to subsets over time and measuring rollout impact.
Business context 03

Why it matters in iGaming

iGaming operators run many simultaneous interventions: CRM campaigns, ML recommendations, bonus offers, communications, lifecycle programmes. Customers experiencing positive outcomes after an intervention may have produced the same outcomes without it. Without holdouts, operators cannot distinguish actual intervention impact from coincidence, baseline behaviour or other factors. The result is overcredit to interventions that do not actually drive incremental value, underinvestment in interventions that do, and reporting that systematically overstates marketing and CRM ROI.

Different teams use holdouts differently:

  • CRM measures campaign incremental impact through campaign-specific holdouts.
  • ML teams measure model impact through model-specific holdouts.
  • Marketing measures channel and acquisition impact through randomised assignment.
  • VIP management measures intervention impact on high-value customers.
  • Finance benefits from honest ROI measurement supporting investment decisions.

Holdouts also reveal uncomfortable truths. Many CRM campaigns that look successful turn out to have small or zero incremental impact when measured against holdouts. ML recommendation systems sometimes underperform simple alternatives. Bonus offers may produce activity that would have happened anyway. Operators that do not use holdouts continue investing in interventions without knowing whether they work; operators that use holdouts learn what actually drives value and reallocate accordingly.

Failure modes 04

Common mistakes and how operators get holdouts wrong

No holdouts at all. Without holdouts, ROI measurement is fundamentally observational and biased toward overstating intervention impact. Operators relying on "customers who received campaign performed better than customers who did not" comparisons systematically overcredit their CRM and marketing efforts.

Non-random holdout assignment. Holdouts must be randomly assigned, not selected based on customer attributes ("hold out the customers we do not think will respond anyway"). Non-random assignment biases the comparison and produces wrong conclusions about intervention impact.

Holdouts too small. Very small holdouts cannot detect realistic effect sizes. The required holdout size depends on baseline outcome variance and the effect size you want to detect. Statistical power calculations should determine holdout size, not arbitrary defaults.

Holdouts too large. Excessive holdouts forgo intervention value from the held-out customers. The trade-off between measurement precision and forgone intervention value should be explicit. Most operators end up with 5-10% holdouts as a reasonable balance.

Holdouts not maintained. Customers assigned to holdout for one campaign may need to remain in holdout for related campaigns to maintain measurement validity. Holdout discipline matters; ad-hoc reassignment compromises measurement.

Holdouts ignored at decision time. Holdouts produce data that contradicts intuition: campaigns thought to be successful may show no impact. Operators that override holdout findings ("the holdout must be wrong, the campaign clearly worked") get no value from the measurement. Discipline to act on holdout results matters.

No follow-up analysis. Holdouts reveal what interventions work; operators that do not dig into why some work and others do not miss the learning opportunity. Holdout findings should inform iterative improvement, not just yes/no decisions.

What good looks like 05

What good looks like

Holdout practices observed in mature operations:

  • Holdouts on all significant CRM campaigns, ML interventions and customer-facing programmes.
  • Random assignment with documented methodology.
  • Statistical power calculations supporting holdout sizing decisions.
  • Holdout discipline maintained across campaign sequences.
  • Holdout findings reviewed in operational decision-making.
  • Follow-up analysis explaining holdout results and informing iteration.
  • Universal holdouts at the strategic level measuring aggregate CRM impact.
Gamblitude 07

How Gamblitude supports holdout discipline

In Gamblitude, holdout group functionality is built into the dynamic Lists and CRM activation infrastructure. Random assignment to holdout, persistent holdout maintenance across campaigns, statistical power calculations and post-campaign measurement against holdout outcomes are supported natively. CRM teams can run campaigns with confidence in incremental impact measurement; ML teams can measure recommendation engine impact against random control. Universal strategic holdouts at the operator level measure aggregate CRM impact over time. Operators get experimental rigour without building separate experimentation infrastructure.

Explore AI for iGaming ↗
Questions 08

FAQ

Because non-recipients are not randomly selected. Customers who received a campaign were typically selected for reasons (they fit the target profile, they are at the right lifecycle stage). Comparing them to non-recipients compares apples to oranges: the selection itself, not the intervention, may explain outcome differences. Random holdout assignment within the target population is what makes the comparison valid.

Depends on baseline variance and target effect size. Statistical power calculations support specific decisions, but most production CRM holdouts end up in the 5-10% range as a reasonable balance between measurement precision and forgone intervention value. Smaller holdouts (5%) suffice for high-volume campaigns; larger holdouts (10-15%) help for low-volume campaigns or smaller effect detection.

Yes, with care. Multiple overlapping campaigns each with holdouts can produce complex assignment patterns. Customers in holdout for campaign A but receiving campaign B may produce confusing results. Operators with serious experimentation discipline maintain consistent holdout assignment or explicitly account for overlap in analysis.

Real cost, but worth it usually. The forgone value from holdout customers (typically 5-10% of intervention value) is the cost of knowing whether interventions work. Operators that skip holdouts to maximise short-term intervention reach often invest in interventions that do not work, accumulating losses dramatically larger than the holdout cost.

Long enough to measure outcomes. CRM campaigns may produce immediate response measurable within days; LTV-affecting interventions may need months of follow-up for meaningful comparison. The measurement window determines minimum holdout duration. Universal strategic holdouts often persist for years to measure aggregate CRM impact over time.

Explore next 09

Further reading

Keep the glossary useful

Found a mistake or want a term added to the iGaming Glossary? Let us know.

Browse the complete glossary or see how governed definitions work across dashboards, reports, alerts and AI answers.