# Why the first AppStore2031 forecast was withdrawn

## Status

The `2031-2026-07-31` inventory is a superseded continuity draft. Its
categories, fictional listings and ranks must not be presented as the active
AppStore2031 forecast. The evidence register and implementation remain useful
inputs, but the forecast inference is being rebuilt.

This decision was made on 1 August 2026 before the edition was sealed, frozen,
deployed or published.

## What failed

The first method answered a narrower question than the project intended. It
found defensible products that could plausibly become popular by 2031. It did
not consistently find products and categories that arise from a materially
changed 2031 world.

Several design choices pushed the result towards continuity:

1. Candidate strength rewarded delivered recurring behaviour, established
   reference classes and credible present-day distribution.
2. High-impact weak signals received more uncertainty and therefore tended to
   lose rank to products resembling things already delivered.
3. The taxonomy process began by resolving the fate of Apple's current
   categories, which anchored later work to today's marketplace ontology.
4. The four scenarios varied delegation depth and ecosystem openness but did
   not vary the underlying level of machine capability, the price of cognitive
   labour or the resulting institutional transformation.
5. Candidate authors were not required to pass a dated current-market and
   published-concept collision audit
   test before ranking.
6. The quality loop tested reproducibility, source integrity, calibration,
   geographic evidence, accessibility and disclosure without testing temporal
   distance from the evidence cutoff.
7. A later product-explanation pass improved the descriptions while
   deliberately leaving the invalidated model inputs and ranks unchanged.

The result often forecast a present-day product plus better AI, consent,
provenance, interoperability or regulation. Those may be valuable continuity
bets, but they cannot dominate an exercise intended to help readers experience
how different 2031 could be.

## Diagnostic evidence

- The draft contained 23 categories and 230 fictional listings.
- Its scenario axes were autonomy and ecosystem openness; capability level was
  not an independent uncertainty.
- Its category process was explicitly connected to a frozen set of 25 current
  Apple categories.
- The product-story layer classified digital currencies and blockchain as
  `not-needed` for 226 listings and `optional` for four; none classified them as
  required. The problem is not that blockchain must be required by a quota. It
  is that the process asked whether a present-shaped product technically needed
  it instead of exploring how alternative ownership and settlement systems
  could change products, actors and markets.
- Representative listing jobs could be matched to products or infrastructure
  already delivered by the evidence cutoff. Their forecast additions were
  usually incremental rather than category-creating.

These observations invalidate the inventory as the main forecast. They do not
prove that every individual premise or source is false.

## What remains useful

- The dated source register, including contradictory, rejected and superseded
  evidence.
- Regional evidence and translation provenance after claim-level review.
- The 2026 marketplace capture as a later dated collision-audit boundary.
- Fictional disclosure, public provenance, immutable-edition and route-proof
  infrastructure.
- The static-site implementation, accessible detail panel and research reader.
- The failed draft itself as an inspectable record of a modelling error.

No category, listing, probability or rank carries into the replacement by
default.

## Replacement separation rule

The replacement run has three workflow-separated stages. This limits supplied
and actively retrievable run material; it does not erase model pretraining or
unavoidable developer/workspace context.

1. **Future construction.** Researchers scan disruptions, weak signals,
   counterforces and cross-impacts. Scenario, category and product builders see
   no previous inventory and no current-app comparison material.
2. **Dated collision audit.** After candidate generation, separate researchers
   compare each product with current markets and published concepts at the
   evidence cutoff. A candidate that is substantially delivered must be
   rejected or demonstrate a material difference created by a named future
   condition. A pass says no collision was found in those dated surfaces; it
   does not claim global novelty.
3. **Ranking.** Only candidates that survive the future-distance review may be
   ranked. Consequential uncertainty remains visible rather than being treated
   automatically as low importance.

The comparison stage is dynamic. No named company, current product or favoured
technology is embedded in the method. Examples that exposed this failure are
historical evidence, not permanent rules.

## Refresh consequence

The original refresh skill is disabled during the replacement Gauntlet. Its
delta-only research window, previous-edition cloning, unchanged-forecast rule
and fixed-model constraint would reproduce the same anchoring.

The replacement skill must regenerate the world model, categories and products
without being supplied the previous forecast, run a new cutoff-date collision
audit, and compare editions only after independent generation is complete.

## Second correction: the Stage D threshold error

The replacement process then made a narrower but important mistake. It treated
“Top 10” as evidence that a category needed ten different recurring jobs before
products could be designed. That confused two questions:

- Stage D asks whether a category is a useful, separately browsable part of the
  future marketplace.
- Stage F asks whether ten genuinely different products survive dated
  current-market and published-concept collision research.

The threshold rewarded broad catch-all categories and removed precise
categories with fewer job labels, even though several very different products
can solve one need. The one-category result was a modelling error, not a
surprising forecast about 2031.

We preserved Stages A–C byte-for-byte, archived every threshold-tainted Stage D
proposal, repair, packet, receipt, output and tool with hashes, and reset Stage
D. The corrected method uses purpose, independent distribution, changed-need
traceability, boundary rules and a July 2031 resolution test. Ten is now a
post-audit publication gate only.

A passing Stage F result is now called
`future-dependent-no-collision-found`: no collision was found in the named
search surfaces by the evidence cutoff. This is deliberately narrower than
claiming nobody has built the idea anywhere.

## Third correction: the lived-experience input gap

Removing the Stage D numeric threshold fixed one bias but exposed another. The
114 Stage C needs were valid accounts of institutional, infrastructure,
regulated-service, care-pathway and recovery work. They were not a complete
account of ordinary life.

Two independent comparative critics found that all three resulting ontology
proposals therefore behaved more like procurement or commissioned-service
catalogues than a full successor app store. Important everyday purposes were
too thin or absent, including relationships, belonging, creativity, play,
culture, ordinary communication and self-directed activity.

We archived that Stage D attempt byte-for-byte and marked it provenance-only.
We did not invent missing consumer categories at Stage D, and we did not
discard the institutional needs. Stage C2 now researches lived-experience
needs independently. Only a later hash-bound union packet may restart Stage D.
