archive
Why the first forecast was withdrawn
A public AppStore2031 research record. The readable view is generated without changing the preserved source.
Open the exact Markdown sourceResearch boundary: Observed evidence, inference, scenario and fictional forecast claims retain the labels used in the source record.
Why the first AppStore2031 forecast was withdrawn
Status
The 2031-2026-07-31 inventory is a superseded continuity draft. Its
categories, fictional listings and ranks must not be presented as the active
AppStore2031 forecast. The evidence register and implementation remain useful
inputs, but the forecast inference is being rebuilt.
This decision was made on 1 August 2026 before the edition was sealed, frozen, deployed or published.
What failed
The first method answered a narrower question than the project intended. It found defensible products that could plausibly become popular by 2031. It did not consistently find products and categories that arise from a materially changed 2031 world.
Several design choices pushed the result towards continuity:
- Candidate strength rewarded delivered recurring behaviour, established reference classes and credible present-day distribution.
- High-impact weak signals received more uncertainty and therefore tended to lose rank to products resembling things already delivered.
- The taxonomy process began by resolving the fate of Apple's current categories, which anchored later work to today's marketplace ontology.
- The four scenarios varied delegation depth and ecosystem openness but did not vary the underlying level of machine capability, the price of cognitive labour or the resulting institutional transformation.
- Candidate authors were not required to pass a dated current-market and published-concept collision audit test before ranking.
- The quality loop tested reproducibility, source integrity, calibration, geographic evidence, accessibility and disclosure without testing temporal distance from the evidence cutoff.
- A later product-explanation pass improved the descriptions while deliberately leaving the invalidated model inputs and ranks unchanged.
The result often forecast a present-day product plus better AI, consent, provenance, interoperability or regulation. Those may be valuable continuity bets, but they cannot dominate an exercise intended to help readers experience how different 2031 could be.
Diagnostic evidence
- The draft contained 23 categories and 230 fictional listings.
- Its scenario axes were autonomy and ecosystem openness; capability level was not an independent uncertainty.
- Its category process was explicitly connected to a frozen set of 25 current Apple categories.
- The product-story layer classified digital currencies and blockchain as
not-neededfor 226 listings andoptionalfor four; none classified them as required. The problem is not that blockchain must be required by a quota. It is that the process asked whether a present-shaped product technically needed it instead of exploring how alternative ownership and settlement systems could change products, actors and markets. - Representative listing jobs could be matched to products or infrastructure already delivered by the evidence cutoff. Their forecast additions were usually incremental rather than category-creating.
These observations invalidate the inventory as the main forecast. They do not prove that every individual premise or source is false.
What remains useful
- The dated source register, including contradictory, rejected and superseded evidence.
- Regional evidence and translation provenance after claim-level review.
- The 2026 marketplace capture as a later dated collision-audit boundary.
- Fictional disclosure, public provenance, immutable-edition and route-proof infrastructure.
- The static-site implementation, accessible detail panel and research reader.
- The failed draft itself as an inspectable record of a modelling error.
No category, listing, probability or rank carries into the replacement by default.
Replacement separation rule
The replacement run has three workflow-separated stages. This limits supplied and actively retrievable run material; it does not erase model pretraining or unavoidable developer/workspace context.
- Future construction. Researchers scan disruptions, weak signals, counterforces and cross-impacts. Scenario, category and product builders see no previous inventory and no current-app comparison material.
- Dated collision audit. After candidate generation, separate researchers compare each product with current markets and published concepts at the evidence cutoff. A candidate that is substantially delivered must be rejected or demonstrate a material difference created by a named future condition. A pass says no collision was found in those dated surfaces; it does not claim global novelty.
- Ranking. Only candidates that survive the future-distance review may be ranked. Consequential uncertainty remains visible rather than being treated automatically as low importance.
The comparison stage is dynamic. No named company, current product or favoured technology is embedded in the method. Examples that exposed this failure are historical evidence, not permanent rules.
Refresh consequence
The original refresh skill is disabled during the replacement Gauntlet. Its delta-only research window, previous-edition cloning, unchanged-forecast rule and fixed-model constraint would reproduce the same anchoring.
The replacement skill must regenerate the world model, categories and products without being supplied the previous forecast, run a new cutoff-date collision audit, and compare editions only after independent generation is complete.
Second correction: the Stage D threshold error
The replacement process then made a narrower but important mistake. It treated “Top 10” as evidence that a category needed ten different recurring jobs before products could be designed. That confused two questions:
- Stage D asks whether a category is a useful, separately browsable part of the future marketplace.
- Stage F asks whether ten genuinely different products survive dated current-market and published-concept collision research.
The threshold rewarded broad catch-all categories and removed precise categories with fewer job labels, even though several very different products can solve one need. The one-category result was a modelling error, not a surprising forecast about 2031.
We preserved Stages A–C byte-for-byte, archived every threshold-tainted Stage D proposal, repair, packet, receipt, output and tool with hashes, and reset Stage D. The corrected method uses purpose, independent distribution, changed-need traceability, boundary rules and a July 2031 resolution test. Ten is now a post-audit publication gate only.
A passing Stage F result is now called
future-dependent-no-collision-found: no collision was found in the named
search surfaces by the evidence cutoff. This is deliberately narrower than
claiming nobody has built the idea anywhere.
Third correction: the lived-experience input gap
Removing the Stage D numeric threshold fixed one bias but exposed another. The 114 Stage C needs were valid accounts of institutional, infrastructure, regulated-service, care-pathway and recovery work. They were not a complete account of ordinary life.
Two independent comparative critics found that all three resulting ontology proposals therefore behaved more like procurement or commissioned-service catalogues than a full successor app store. Important everyday purposes were too thin or absent, including relationships, belonging, creativity, play, culture, ordinary communication and self-directed activity.
We archived that Stage D attempt byte-for-byte and marked it provenance-only. We did not invent missing consumer categories at Stage D, and we did not discard the institutional needs. Stage C2 now researches lived-experience needs independently. Only a later hash-bound union packet may restart Stage D.
AppStore2031