archive
Withdrawn v1 backtest
A public AppStore2031 research record. The readable view is generated without changing the preserved source.
Open the exact Markdown sourceResearch boundary: Observed evidence, inference, scenario and fictional forecast claims retain the labels used in the source record.
2021→2026 retrospective stress test — withdrawn method v1 record
Historical evidence, not an active validation claim. This stress test belongs to the withdrawn continuity-led method. Its failed data gate and limitations remain public because they informed the rebuild; it does not validate the future-native v2 method or the replacement catalogue.
Status at withdrawal: the historical data gate did not pass and improvement was not demonstrated.
This protocol was intended to test whether method v1 could outperform a naïve “nothing important changes” forecast over a previous five-year period. Its required historical-data gate did not pass, so it cannot support that conclusion.
The test is an external challenge from actual App Store outcomes. It is not a rubric written to reward the forecast, and its result must be published even if the method loses.
1. Question and cutoff
The retrospective question is:
Using only a sealed evidence corpus published by 31 July 2021, what App Store categories and app archetypes would the method have forecast for the Top Free iPhone charts during July 2026?
Constants:
evidenceCutoff: 2021-07-31T23:59:59Z
outcomeWindow:
start: 2026-07-01
end: 2026-07-31
chartFamily: top-free-iphone-apps
listDepth: 10
rboPersistence: 0.9
minimumValidDays: 16
probabilityFloor: 0.01
probabilityCeiling: 0.99
2. Sealed-corpus procedure
- Assemble a 2021 evidence corpus containing only sources with a verified publication date on or before the cutoff.
- Record each URL, publisher, publication date, acquisition date, inclusion reason, language, translation method, licence limitations, and file hash in a manifest.
- Exclude undated sources unless an authoritative archive proves they existed by the cutoff.
- Prevent forecast builders from web browsing or seeing search snippets while generating the hindcast. They receive the sealed corpus only.
- Apply the same category, archetype, probability, regional, scenario, and ranking contract used for 2031. Deviations must be declared before scoring.
- Export the completed 2026 forecast and record its cryptographic hash.
- Only after the forecast hash is frozen may the 2026 outcome packet be opened.
The source manifest, forecast hash, inclusion/exclusion decisions, and tool configuration are public.
3. Coverage and missing-data gate
Preferred coverage is every Apple category for which authoritative July 2021 and July 2026 chart data can be recovered in the frozen regional basket. Categories may not be excluded because their result is inconvenient.
If complete historical charts are unavailable, use this fallback before any 2026 outcome is inspected:
- Always include the All Apps Top Free chart.
- Partition the July 2021 category list into functional strata using only Apple's 2021 descriptions.
- Publish the strata and a deterministic hash seed.
- Select one category per stratum from the hash result.
- Publish every data-availability failure and exclusion.
If authoritative or defensibly archived data cannot resolve a selected chart,
mark it not-resolvable. Do not reconstruct or invent historical ranks.
Full daily data and a one-day snapshot are not equivalent. If only a single
dated snapshot exists, label the result single-day-limited, exclude it from
the primary aggregate, and show it only as a sensitivity check.
4. Naïve persistence baseline
The baseline is frozen without reading 2026 outcomes. It predicts:
- Every July 2021 category still exists in July 2026.
- No new category appears.
- Each July 2021 monthly Top 10 archetype remains in the Top 10 in July 2026.
- Those archetypes retain the same order.
The 2021 apps are converted to the same archetype ontology used by the method. Both candidate and baseline therefore compete at the archetype level. A broad future archetype is not allowed to compete against a narrowly identified incumbent.
The persistence baseline is a ranked-list comparator, not a fabricated probability model. Candidate probabilities are scored for honesty and calibration but are not claimed to beat a probabilistic baseline unless a separate baseline is trained solely on pre-2021 transitions and frozen in advance.
5. Outcome capture
For every selected storefront/category pair:
- Capture the official Top Free iPhone chart once per day at local noon during 1–31 July 2026.
- Store the retrieval time, source URL, raw response, screenshot, and hash.
- Give an app at daily rank
kthe value11-k. - Give an app absent from that day's Top 10 zero.
- Average points across all scheduled days, including zeroes for absence.
- Sort by average points to create the realised monthly Top 10.
- Require at least 16 valid daily captures for a primary outcome.
Tie breakers, in order, are more days at the better rank, lower mean rank on days present, then stable App Store identifier. The raw daily series and computed monthly result must both be downloadable.
Each storefront/category must name one canonical signed-out, human-facing Apple Top Free category page before its outcome packet is opened. The corpus records its exact URL, category ID, locale and extraction rule. Feeds and independent datasets are diagnostics only unless one exact fallback was predeclared, test-captured and frozen before the outcome window because the canonical page had already been permanently withdrawn. A disagreement is resolved in favour of the frozen human-facing page, not whichever surface helps the forecast. Fewer than 16 defensible daily observations from the canonical surface makes the primary outcome ambiguous.
6. Blind archetype mapping
Forecast names, icons, ranks, probabilities, and rationale are hidden from the mapping team.
Three independent coders receive only:
- The locked archetype definitions.
- Required and distinguishing capabilities.
- Disqualifiers.
- Actual 2026 app metadata and observable product capabilities.
A coder records eligible, ineligible, or insufficient-evidence for every
archetype/app pair, plus a checklist result for each required capability and
disqualifier. An edge exists in the matching matrix only when at least two of
three coders independently mark the pair eligible. There is no discretionary
post-hoc adjudication; missing or ambiguous product evidence cannot create an
edge.
Matching is performed independently inside each storefront/resolved-category monthly Top 10. The bipartite left side is the ten locked archetype IDs; the right side is the ten realised App Store IDs. The deterministic optimiser uses this objective tuple, in order:
- Maximise the number of matched pairs.
- Maximise the sum of eligible coder votes across those pairs (
2or3per edge). - Among remaining ties, choose the lexicographically smallest sequence of
archetypeId|appStoreIdpairs after sorting by archetype ID and then App Store ID.
The implementation enumerates or exactly solves the finite 10×10 bipartite assignment; it may not use forecast rank, fictional name, probability, rationale, realised rank, or a semantic similarity score as an edge weight. The complete anonymised 3-vote matrix, selected pairs, unmatched vertices, algorithm version and result hash are published. This makes the one-to-one mapping reproducible while preventing one vague archetype from claiming several realised apps.
Unmatched realised apps are unforeseen entries. Unmatched predictions are misses. Ambiguous product capabilities remain unmatched rather than being resolved in the forecast's favour.
7. Metrics
All metrics and their roles are frozen before outcomes are opened.
Primary comparative metric
Use Rank-Biased Overlap at depth 10 with persistence parameter p = 0.9.
RBO weights the top of the chart most heavily and permits two lists that do not
contain identical members. The reference is Webber, Moffat, and Zobel,
“A Similarity Measure for Indefinite
Rankings”.
Compute RBO separately for every evaluable storefront/category pair. For each category, first average storefront results inside their lens, then average the lens values; finally macro-average categories. This two-stage geography rule prevents a lens with more captured storefronts from dominating.
The method demonstrates improvement over persistence only if:
- Its macro-mean RBO@10 is greater than the baseline's.
- The paired category bootstrap 95% confidence interval for the difference excludes zero.
If either condition fails, publish “improvement not demonstrated.” Do not substitute a more favourable primary metric after seeing the result.
Secondary ranked-list diagnostics
- Top 10 archetype overlap: intersection size divided by ten.
- Kendall's tau-b over archetypes common to prediction and outcome.
- Mean absolute rank error over matched archetypes.
- Count and share of unforeseen realised entries.
- Precision and recall for new, retained, merged, and retired categories.
- Results by storefront, category, and the equal-storefront lens summaries, not only the global mean.
Proper probability scores
For category existence and the joint event that an archetype reaches its
category's Top 10, report logarithmic loss and Brier score. For binary outcome
y and forecast probability p:
logLoss = -(y ln(p) + (1-y) ln(1-p))
brier = (p-y)^2
Lower is better. Published probabilities are constrained to 0.01–0.99 so a wrong certainty cannot produce infinite loss.
For the ordered rank outcomes 1, 2–3, 4–5, 6–10, and
outside-top-10, report:
- Multiclass logarithmic loss as the primary proper score.
- Ranked Probability Score as a distance-sensitive diagnostic.
Also publish calibration tables and a reliability plot. Because one retrospective run may contain too few independent forecasts for stable calibration estimates, show counts in every bin and state that limitation.
These choices follow Metaculus's proper-scoring principle: a score should reward the forecaster for stating their sincere probability rather than gaming the metric.
8. Category and conditional-rank resolution
A category resolves as existing for a storefront only if it appears as a top-level browse category with an independently ranked Top Free chart during at least 16 valid days in the outcome window. Renames and mergers are resolved from the frozen semantic definition, not name similarity alone.
If a category does not exist:
- Its category-existence event resolves No.
- Each app's joint Top 10 event resolves No.
- Conditional rank distributions are not scored separately.
A multi-storefront lens reports the equal-storefront mean of valid outcomes and proper scores; it is not assigned one invented rank or binary event. A global category outcome is the descriptive equal-lens mean of those storefront-coverage shares. It is not assigned a binary forecast probability or scored; category-existence probabilities are evaluated separately by storefront, then macro-averaged within lenses. Market/category pairs with inadequate evidence are ambiguous and excluded, with the exclusion count shown.
9. Contamination disclosure
The result page must reproduce this disclosure prominently:
The backtest restricts cited evidence to information available by July 2021, but an AI model operating in 2026 may retain later facts in its model parameters. Corpus isolation and sealed outcome packets reduce direct leakage; they cannot remove model-memory contamination. This is therefore a retrospective pipeline stress test, not proof of genuine out-of-sample forecasting skill.
Additional limitations to publish:
- Recovered historical charts may be incomplete or use different capture times.
- Apple can change taxonomy and ranking surfaces without documenting the full ranking mechanism.
- Archetype eligibility contains judgement even with blinded independent coders and deterministic matching.
- Regional storefront coverage is not the same as global download share.
- Category-level observations are correlated, so a bootstrap interval is not a guarantee of generalisable skill.
10. Required public artefacts
The backtest is not complete unless the site publishes:
- The cutoff and frozen question.
- Corpus manifest and source inclusion/exclusion decisions.
- Forecast file hash and freeze timestamp.
- Naïve persistence forecast.
- Raw or legally shareable chart captures and hashes.
- Daily-to-monthly outcome calculation.
- Blind mapping packet, anonymised coder matrices, deterministic assignment, and hashes.
- Per-category candidate and baseline metrics.
- Aggregate metrics and paired confidence interval.
- Proper-score calculations and calibration-bin counts.
- Missing data and ambiguous outcomes.
- Contamination disclosure and limitations.
- A plain-language result stating
won,lost, orimprovement not demonstrated.
Future refreshes may improve the method prospectively, but they may not alter this frozen test or rewrite its result.
AppStore2031