All posts

Insight

mobile gamingLTV predictionuser acquisitiondata sciencegame analytics

What We Learned Building D7-to-D28 LTV Prediction Systems for Mobile Games

Mai Tran
Mai Tran
Co-Founder & COO
2 June 2026 9 min read
What We Learned Building D7-to-D28 LTV Prediction Systems for Mobile Games
Mobile game analytics dashboard showing LTV prediction curves

In mobile gaming, 28 days is a long time. A user-acquisition team cannot always wait four weeks to know whether a campaign is worth scaling. Creative tests, bids, geographies, and budget allocation decisions often need to happen while the campaign is still alive. That is why many studios want a model that can look at a player's first week and estimate their value by day 28.

On paper, the request sounds simple: use D0–D7 data to predict D28 revenue. In practice, a good LTV system becomes much more than a model. It becomes a way to understand how value is created inside the game economy. It forces the team to answer four questions: Which metric should define success? Which early signals actually matter? How should we treat whales? And can one game teach another?

The observations below come from an anonymized mobile game analytics engagement across casual and puzzle titles. All client identifiers, campaign names, exact data volumes, and commercially sensitive figures have been removed or generalized. The lessons, however, are broadly useful for game founders, UA leaders, product analysts, and data teams building LTV prediction systems.


Lesson 1: MAPE Can Punish the Wrong Mistakes

The most common early confusion in LTV projects is metric selection. Many teams instinctively ask for MAPE — mean absolute percentage error. It sounds intuitive: if the model predicts $0.80 and the player eventually generates $1.00, the error is 20%. Easy.

But mobile game revenue is not normally distributed. Most players generate zero or near-zero revenue, while a small minority generates most of the dollars. In that world, MAPE can become mathematically loud but commercially quiet. A two-cent error on a player worth one cent can look like a catastrophic percentage miss. A multi-dollar error on a high-value player can look less dramatic in percentage terms — even though it matters far more to ROAS.

A useful way to explain this to executives: MAPE screams at pennies and whispers at dollars.

That does not make MAPE useless. It means MAPE should be used carefully, often only above a revenue threshold, and rarely as the only headline metric.

For mobile game LTV, WAPE is often a better primary metric. WAPE asks: out of the total revenue dollars in this cohort, what fraction did the model misallocate? This aligns more closely with UA economics because acquisition decisions are usually made at the cohort, campaign, source, or country level — not from one isolated user prediction.

MetricWhat it tells youBest useMain trap
MAPEAverage percentage error per userDiagnostic only; use with revenue thresholdsCan exaggerate tiny-dollar mistakes
WAPETotal error divided by total actual revenuePrimary metric for UA economicsCan hide poor ranking if used alone
MAEAverage dollar error per userExecutive communication and segment diagnosticsLow-value users can dominate the average
Precision / RecallQuality of high-value user detectionWhale targeting, bidding, CRM, monetizationNeeds a clear threshold and operating point
Cohort ErrorError at campaign, week, country, or source levelBudget and ROAS decisionsCan look good while individual users are misranked

Lesson 2: D7 Revenue Is Useful, but Behavior Explains the Future

Signal heatmap showing D7 behavioral features correlated with D28 LTV)

A player's first-week revenue is obviously valuable. If someone has already generated revenue by day 7, that is a strong signal. But revenue alone is a snapshot. It tells you what has happened; it does not always tell you whether the player is accelerating, decaying, or preparing to become valuable later.

In practice, strong D7-to-D28 models rely on a stack of signals:

- Revenue trajectory — is revenue rising or falling from mid-week to day 7? - Retention — is the player still active on day 6 or day 7? - Progression — are they moving through levels or stuck? - Session depth — are they casually opening the game, or showing the repeated engagement of an invested player? - Ad behavior — are impressions increasing, stable, or collapsing? - Market context — two players with the same engagement can generate different revenue if they are in different ad markets.

One particularly useful category is the game client's own cumulative monetization signal over the first week, provided it is audited carefully for leakage. If it genuinely reflects only observable D0–D7 activity, it can capture revenue dynamics not always fully reflected in server-side tables or raw event counts.

The broader insight: LTV is not predicted by money alone. It is predicted by motion — revenue moving, player activity continuing, progression deepening, ad exposure changing, and the player remaining alive in the economy after the novelty of install day fades.

Signal familyExamplesBusiness question
Revenue motionD7 revenue, revenue velocity, revenue acceleration/decayIs value continuing or fading?
RetentionActive on late D7 days, distinct active days, last active dayIs the player still alive in the economy?
Gameplay depthMax level, sessions, attempts, failures, replay behaviorIs the player invested or just sampling?
Ad behaviorImpressions, rewarded ads, interstitial exposure, ad revenue rhythmIs monetization opportunity expanding?
Market contextCountry, eCPM, source, adgroup/creative contextWhat is each engagement unit worth?
Portfolio contextComparable games, genre-level patterns, campaign familiesCan a mature title teach a newer one?

Lesson 3: Whale Prediction Is a Different Problem from Average Prediction

The biggest mistake in many game analytics projects is treating all users as if they belong to one smooth distribution. They do not. In most free-to-play economies, a small group of high-value users carries a disproportionate share of revenue. These players are not just "slightly higher than average" — they are the business model.

A single global regressor trained on all users has a natural bias: because most users are low-value, the model is rewarded for predicting small values most of the time. That can produce acceptable global error while underestimating the very users who matter most. For UA teams, missing high-value users can be more expensive than being slightly wrong on thousands of low-value users.

A more practical architecture is often two-stage:

    • Use a classifier to rank users by their probability of becoming high-value.
    • Route the top-ranked users into a specialist regressor that estimates how much those likely whales may be worth.
    • Keep a separate full-cohort model for aggregate ROAS reporting.
This gives each model a cleaner job: one model sorts, one model estimates high-value dollars, and one model reports the campaign-level total.

Whale models should be evaluated with precision, recall, and dollars captured — not just average error. Recall tells you whether you caught the whales. Precision tells you whether the flagged users are genuinely valuable or mostly noise. In UA, a false positive that is still a decent payer may be operationally acceptable; a false negative on a future whale can be much more expensive.


Lesson 4: One Game Is a Dataset; a Portfolio Is an Advantage

Cross-game feature sharing diagram across a mobile game portfolio

A single game may not generate enough examples of rare high-value behavior to train a stable model. This is especially true for new titles, new campaigns, new geographies, or games where whales are rare but commercially important. A portfolio of games changes the equation.

Cross-game learning does not mean blindly training one universal model and applying it everywhere. The games still differ — progression systems, ad placements, level difficulty, economy design, payer mix, and market distribution all matter. But a portfolio can share a feature vocabulary: retention, progression depth, revenue velocity, ad engagement, economy interaction, and market context. It can also share validation discipline: temporal splits, leakage audits, cohort-level monitoring, and segment-specific evaluation.

In real projects, two games in the same broad genre can teach different lessons. One may be mature enough that a full-cohort model works well because most payer signal appears early. Another may require a whale-specific stack because the high-value tail behaves differently. The value of a portfolio approach is not that every game becomes identical — it is that every game contributes to a richer understanding of what early value looks like.


A Practical Framework for Game Studios

Before building the next LTV model, start with the decision, not the algorithm. A model used to size bids for high-value users should not be judged the same way as a model used for weekly campaign ROAS reporting. The metric, target, features, validation, and deployment design must follow the decision.

A strong operating checklist:

    • Define the decision — budget allocation, bid sizing, creative testing, cohort ROAS, or high-value user targeting.
    • Define the target and observation window — what exactly is D28 revenue, and which users have a complete post-install window?
    • Choose the metric stack — WAPE for dollars, MAE for communication, thresholded MAPE for diagnostics, precision/recall for whales, cohort error for UA reporting.
    • Build the signal map — revenue trajectory, retention, progression, sessions, ad impressions, in-game economy, market context, and attribution.
    • Prevent leakage — every feature must be available at scoring time. If a cumulative revenue signal is used, audit the time boundary carefully.
    • Segment by value — evaluate all users, payers, whales, and super-whales separately. A single global score can hide the problem.
    • Compare architectures — full-cohort regression, two-stage whale systems, and portfolio or cross-game variants.
    • Validate like production — train on the past, test on the future, and monitor cohort drift.
    • Deploy as a loop — score new D7 users, compare predictions with actual D28 outcomes later, retrain, and feed results into UA dashboards.

The Lesson That Gets Missed

Mobile game LTV prediction is not a Kaggle problem. It is an operating system for faster decision-making.

The best teams will not ask only, "What is the model error?" They will ask: Did we choose the right metric? Did we capture behavior, not just revenue? Did we treat whales as a separate economic segment? Can our games learn from each other? And can this prediction loop improve every month as more data arrives?

When the answer is yes, D7-to-D28 prediction stops being a report. It becomes a strategic advantage for user acquisition, monetization, and portfolio growth.


If your UA team is waiting 28 days to learn campaign quality, your feedback loop may be too slow. A D7 LTV system can turn early behavior into faster, smarter portfolio decisions.

Written by

Mai Tran
Mai Tran
Co-Founder & COO

Want this applied to your data?

In a free 30-minute call we will map the systems involved and suggest the smallest useful first step. No pitch and no obligation.