# How we are building AppStore2031

**Status:** public process record for a forecast still in progress  
**Evidence cutoff:** 2 August 2026
**Forecast date:** 31 July 2031  
**Method:** future worlds first, products later

AppStore2031 is an educational forecast. It asks what people, machines and
organisations might discover in an app store—or whatever replaces an app
store—in July 2031.

The product names, developers, icons, ratings, reviews and chart positions on
the finished storefront will be invented. The evidence and the route from that
evidence to each invention will be inspectable. A source can show that a trend,
constraint or policy exists. It cannot prove that a fictional product will be
popular five years later.

This document records the method while the replacement forecast is being
built. It does not imply that the future worlds, categories, products or ranks
already exist. The live progress and critic verdicts are in the
[Gauntlet record](../GAUNTLET.md).

The final bounded ranking and listing synthesis, including the stopped audit
transport and its limitations, is documented in
[From sealed research to the AppStore2031 Top 10s](FINAL_SYNTHESIS_AND_LIMITS.md).

## Why we withdrew the first attempt

Our first attempt was careful, sourced and plausible—but it answered the wrong
question. It often imagined a product that could be built now, then made it
faster or added better AI, permissions, provenance or payments. Its categories
also began with today's app-store categories. That made the result feel like
2026–2028 rather than a world transformed by 2031.

The ranking method strengthened this bias. Products with current users,
familiar business models and working reference points looked safer, while
uncertain but important changes lost points. Quality checks caught source and
accessibility problems, but did not ask the crucial question: **is this product
materially from 2031, or is it mainly a better version of something that
already exists?**

We therefore withdrew the whole inventory before publication. We preserved it
as a labelled audit record so readers can see the mistake, but no category,
product or rank carries into the replacement automatically. The full diagnosis
is in [Why the first forecast was withdrawn](FORECAST_POSTMORTEM_V1.md).

This matters beyond one website. A forecast can be well cited and still be
anchored to the present. Transparency should show not only the sources, but
also how the research process tries to stop that anchoring.

## The central rule: build the world before the product

The replacement reverses the order of work:

```text
global evidence
    → interacting disruptions
    → several coherent 2031 worlds
    → changed actors and changed needs
    → a new marketplace shape and new categories
    → sealed product candidates
    → current-market rejection test
    → scenario and regional rankings
    → fictional storefront
```

Researchers generating future worlds are not supplied the old inventory, its
categories or comparisons with current products. Product authors are not
supplied today's app catalogue, and unbound current-market retrieval is blocked
before sealing. Only after a candidate is sealed does a different researcher
search named current-market and published-concept surfaces. This controls the
project material used as the imagination prompt; it cannot remove model
pretraining or unavoidable platform and workspace context.

A ranked item may not even be a conventional phone app. Depending on the world,
the marketplace could distribute an agent, capability pack, persistent
companion, synthetic environment, autonomous-organisation interface, robot
service or another kind of software unit. The research, rather than a fixed
quota, decides how many categories exist.

The candidate method is specified in
[Forecast method v2](FORECAST_METHOD_V2.md). It is still a Gauntlet candidate,
not a claim that every stage has passed.

## What happens at each stage

### A. Scan the horizon without inventing apps

Eight workflow-separated research streams examine the forces that could
reshape daily life by 2031:

1. [AI, agents, compute and human agency](../research/rebuild/ai-agents-compute.md)
2. [Economy, work and finance](../research/rebuild/economy-work-finance.md)
3. [Health, biotechnology and demography](../research/rebuild/health-biotech-demography.md)
4. [Education, social connection and culture](../research/rebuild/education-social-culture.md)
5. [Energy, climate, food and water](../research/rebuild/energy-climate-food-water.md)
6. [Robotics, spatial systems and mobility](../research/rebuild/robotics-spatial-mobility.md)
7. [Geopolitics, governance and security](../research/rebuild/geopolitics-governance-security.md)
8. [Institutional scan of scans](../research/rebuild/institutional-scan-of-scans.md)

The institutional scan tests the breadth and blind spots of major public
forecasts. The seven domain scans then look at observations, projections, weak
signals, counterforces, regional differences and open questions. They are
research inputs—not predictions, categories or product wish lists.

The scans use the Three Horizons model:

- **Horizon 1 (H1):** the systems operating now, including their strengths and
  lock-ins.
- **Horizon 2 (H2):** the contested transition—investment, standards, law,
  behaviour, backlash, failure and repair.
- **Horizon 3 (H3):** materially different conditions that could be normal by
  2031.

H3 is not automatically the most likely outcome. H2 explains what would have
to happen for the world to move from H1 towards an H3 condition, and what could
block it.

### B. Combine changes into several futures

No disruption happens alone. Cheap machine reasoning could collide with energy
limits, ageing, new identity rules, climate shocks, changed money, robotics or
geopolitical fragmentation. Scenario builders will receive only the horizon
evidence and examine those cross-impacts.

They will build several internally consistent worlds, including constrained
progress, transformative machine capability, and material institutional or
geopolitical discontinuity. They will trace direct effects, then second- and
third-order consequences. For example, a technology can change the price of a
task; that can change an organisation; that can then change a right, duty or
scarcity. The worlds are not “good”, “average” and “bad” versions of one story,
and they will not be assigned fake precision just to create a consensus.

Every world must state what would make it inconsistent and which measurable
signs would strengthen or weaken it.

### C. Find the changed actors and needs

The next researchers will ask who can act in each world, what they control or
owe, and what has become scarce, cheap, compulsory or dangerous. Actors may be
people, households, communities, human–machine teams, autonomous agents,
public bodies, robot fleets or infrastructure.

They will identify repeated problems that could support a discovery market.
They will not start with “what app should we build?” A need must exist because
the world, an actor or a relationship has changed.

### D. Derive the marketplace and its categories

Taxonomy researchers will receive the changed-actor and changed-need records,
not today's category list. They will ask:

- What does this marketplace actually distribute?
- Who—or what—chooses and trusts a listing?
- Which jobs deserve separate listings?
- Which recurring purposes form useful browse categories?
- Which categories exist only in certain worlds or regions?

A provisional category must have a clear 2031 purpose, a kind of product that
could be listed on its own, a trail back to changed needs, and rules that let a
later assessor decide what belongs. It does not need ten different jobs. That
old rule accidentally rewarded vague catch-all categories and erased useful,
specific ones.

The number of categories is not decided in advance. A category reaches the
storefront only if later research leaves ten genuinely different products that
all survive an independent current-product audit. Ten is a publication gate,
not a Stage D invention target.

### E. Generate and seal future-native candidates

Product authors will work from a future-world and category packet. For each
candidate they must explain, in plain language:

- who or what uses it;
- the new job it performs;
- what changed by 2031 to create that job;
- what the complete experience feels like;
- what it must be able to do—and what would not count;
- how it could be distributed and paid for;
- what technical, legal, physical or social systems it depends on;
- how it could help, harm or be abused; and
- what would stop it from existing.

The candidate is then sealed. Its author cannot browse current products and
quietly reshape it around whatever is already on sale.

### F. Try to reject every candidate with a dated collision audit

Only now does a separate auditor search current app stores, official product
sites, open-source projects, research prototypes and delivered infrastructure.
The aim is to defeat the candidate, not to decorate it with competitors.

A candidate fails if its central experience already exists and its 2031 claim
is mainly better quality, speed, price, scale, AI assistance, permissions,
provenance, interoperability or a different payment rail. It may survive only
if a named future change materially alters the user, job, autonomy, ownership,
physical or social experience, delivery institution, or what is legally,
technically or economically possible.

The auditor returns one of three results:

- **Reject:** substantially delivered by the evidence cutoff.
- **Revise and re-audit:** the difference may be real but is not yet clear.
- **Future-dependent; no collision found:** the candidate depends on a named
  2031 change and no substantial collision was found in the named, dated search
  surfaces.

The final wording maps to `future-dependent-no-collision-found`. It is not a
claim that the candidate is novel, unique, first, or absent from every market.

The rules never name a favoured company, current product or technology. The
market search is repeated at each refresh because the comparison boundary
moves.

### G. Rank by world and region, then show uncertainty

Only surviving candidates can be ranked. Each is considered separately in
each world and geographic lens. The judgement includes need, reach,
conditional feasibility by 2031, repeat use, distribution, infrastructure,
regulation, substitutes, network effects, possible harm and evidence quality.

One fresh author context handles one category and one metric under immutable
protocol `causal-judgement-v8`. It sees no other metric judgement or ranking.
The author explains the causal path, assumptions, falsifiers and evidence
limits, then chooses ordinal lower, central and upper bands. It does not assign
a decimal score or probability. This stops longer, more complete-looking prose
from becoming a ranking advantage.

The output must point to content-hash anchors for that exact candidate
mechanism, world change and geographic constraint. It explains each link in a
separate field and separates candidate, world and geographic evidence. Each
link must independently retain a distinctive multi-word part from every source
it claims to connect: the mechanism and delivered outcome, the general and
candidate-specific world change, and the general and candidate-specific local
constraint. It must then add at least five residual concepts spanning at least
two causal roles. Generic words such as “stated mechanism changes the assessed
result” do not count. Before detecting cloned prose, the validator removes
identity labels, exact anchor vocabulary, metric wording and changing bands.
It then compares both the exact residual sequence and an order-insensitive
semantic form that stems inflections and groups generic synonyms. Words count
only when they occur in explicit semantic evidence-field allowlists for that
row's packet-bound candidate, world, lens, source and claim material. URL,
identifier, provenance and other metadata never ground reasoning. All unbound
words become one `unbound-concept`, so adding numbered placeholders or
unrelated nouns cannot manufacture distinct reasoning. Each authored causal
relation must connect locally grounded arguments on both sides, while every
authored consequence needs a locally grounded argument; grounded words in a
separate aside do not satisfy the rule. Only a role word inside an exact fully
copied anchor span is exempt: repeating that same word later creates a new
authored claim that must be grounded. Matching happens separately inside each
causal clause, so punctuation and `->` cannot reconnect copied fragments; case
and commas inside one clause may still be normalised. Compact, one-sided,
space-padded and tab-padded arrows are all boundaries. A single template with different anchor
text, bands, grammar, synonyms or filler is therefore still a clone and is
rejected alongside generic ID insertion or repeated band matrices.

The deterministic simulator samples those authored bands. It does not add
synthetic noise or turn citation counts into confidence. Public results use
broad simulation-share bands and integer rank ranges, not a probability that a
world or fictional product will occur.

There is no editable judgement layer between authors and the simulator. A
canonical projector derives it from the immutable author packets, outputs and
session receipts. The public build re-reads those sources and requires the
compiled file to match exactly, so changing one band, explanation, anchor or
author pointer cannot silently change a rank.

The same rule now applies to the rankings themselves. One canonical projector
is shared by compilation and production validation. After rebuilding the
judgements, the validator replays every conditional, world, main and
sensitivity chart from the sealed candidates, active audits, worlds and lenses
and requires identical bytes. A receipt chain exposes hashes from the causal
author source tree through the judgement file to each chart, adjacent-rank
explanation and complete ranking record. Recomputing a local hash after an edit
does not make that edit part of the forecast.

The finished analysis will provide:

- a Top 10 for each category inside each world;
- regional variations where evidence supports them;
- a scenario-balanced main chart, clearly labelled as a judgement rather than
  an event probability;
- rank sensitivity when reasonable world weights change;
- confidence and the strongest counter-case; and
- why each product sits above the next candidate.

A high-impact candidate must not disappear simply because its enabling world
is uncertain. The site should let a reader see where it dominates, where it
fails and how sensitive its rank is.

## Evidence is labelled, not blended

Institutional reports often put facts, models, expert views, ambitions and
scenarios beside one another. They are not the same kind of evidence. The
research register separates:

| Label | Meaning | What it can tell us |
| --- | --- | --- |
| Observed or estimated data | A measured baseline or documented event | What is happening or has happened by the cutoff |
| Measured trend | A dated direction across observations | What has been changing, not whether it must continue |
| Modelled projection | An output conditional on model inputs | One possible trajectory under stated assumptions |
| Expert elicitation | What a sampled group expects or fears | A view worth testing, not an event probability |
| Policy intent | A law, target, strategy or announced project | Direction and possible support, not delivery |
| Weak signal | Small or early evidence with larger possible consequences | A branch to investigate, not proof of scale |
| Scenario assumption | A condition deliberately used to build a coherent world | What follows if it holds, not a factual claim |
| Design inference | Our explanation joining evidence to a possible product | A contestable research judgement |
| Explicit unknown | Something the available evidence cannot settle | A limitation and a future research task |

Every factual premise should retain its publisher, URL, publication and access
date, geography, language or translation route, limitations and status.
Contradictory, rejected and superseded evidence stays visible. Popularity does
not equal truth: a subject covered by many institutions is not necessarily the
most important combination of changes.

## Global does not mean one global average

The scan deliberately includes the United States, China and other leading
technology economies, but does not treat them as the whole future. It also
examines regions with different demographics, infrastructure, climate
exposure, informal economies, institutional capacity and technology paths.

Each research stream records regional divergence and missing coverage. Later
worlds and rankings must preserve meaningful differences rather than smoothing
them into an imaginary global user. Translated evidence must identify the
original language and method. Uneven evidence is itself a limitation: public
institutions and English-language sources describe formal systems better than
informal work, unpaid care, minority languages, low-resource settings and
local trust networks.

## What the fictional storefront will show

The interface will resemble a familiar marketplace so that an unfamiliar
future is easy to explore. Familiar presentation must not imply that the
marketplace itself is unchanged.

Every listing will separate six layers:

1. **Observed evidence:** the dated facts and trends beneath the forecast.
2. **Inference:** how we connect those facts.
3. **Scenario condition:** what must be true in that future world.
4. **Product design:** the imagined response to a changed need.
5. **Rank judgement:** why it occupies this position and how sensitive it is.
6. **Imagined marketplace copy:** the fictional name, sales line, rating and
   comments used to make the idea understandable.

The listing's future-distance panel will answer four simple questions:

1. Why could this not be the same product in 2026?
2. What changed in the world?
3. What does the 2031 product let someone—or something—do?
4. What could prevent it from appearing?

Imagined reviews can illustrate benefits and failures, but they are never
evidence. Fiction will be labelled at the point where readers encounter it.

## How the Gauntlet challenges the work

The Gauntlet is an adversarial quality loop. A builder produces a real
artifact; a fresh critic that is not supplied the builder's conversation compares it with
the external bar and identifies the single biggest gap. The builder repairs
that gap, then a different fresh critic judges the revised artifact. Builders
do not grade themselves, and we do not decide the number of rounds in advance.
This separation controls project inputs; it does not erase model pretraining or
unavoidable platform and workspace context.

The external bar combines the UK Government Office for Science
[Futures Toolkit](https://www.gov.uk/government/publications/futures-toolkit-for-policy-makers-and-analysts/the-futures-toolkit-html),
the OECD [Strategic Foresight Toolkit for Resilient Public Policy](https://www.oecd.org/en/publications/foresight-toolkit-for-resilient-public-policy_bcdd9304-en.html),
and a live, dated current-market and published-concept collision audit.
Criticism checks stage separation, global
coverage, counterforces, harms, temporal distance, evidence links,
understandability, data quality, accessibility and real local routes.

The [Gauntlet record](../GAUNTLET.md) publishes the pieces, rounds, verdicts,
open gaps and regression checks. Passing a critic means the artifact beat the
stated bar in that comparison; it does not turn a scenario into a fact.

## How refreshes will work

A refresh will be a new forecast from the new evidence cutoff—not an edit of
the previous chart. It will update the horizon scan, rebuild the interactions
and worlds, rederive actors, needs, marketplace and categories, generate fresh
workflow-separated candidates, run a dated current-market and
published-concept collision audit, and rerank the survivors. This separation
controls supplied and actively retrievable run material; it cannot remove
model pretraining or unavoidable developer and workspace context.

Only after the new edition is complete will it be compared with the previous
one. The public change log can then explain which categories or products
appeared, vanished, merged or moved. Those comparisons never become prompts
for the next generation step.

Old frozen editions will remain unchanged, so readers can inspect what the
project believed at each cutoff. A method change must be versioned and pass its
own Gauntlet.

## What we publish—and what we do not

Credible transparency means publishing enough for someone else to challenge
the result:

- source registers and evidence labels;
- dates, geographies, limitations and contradictory evidence;
- stage inputs and isolation rules;
- scenario conditions, causal links and failure conditions;
- candidate inclusion and rejection reasons;
- model fields, rank sensitivity and counter-cases;
- critic verdicts, repairs, tests and known limitations; and
- superseded editions and changes between completed editions.

We do **not** publish private chain-of-thought: hidden token-by-token reasoning,
private scratch work or internal deliberation. That material is neither a
reliable explanation nor necessary for audit. The public rationale is the
structured record above—evidence, assumptions, transformations, decisions,
alternatives, outcomes and tests. It is designed to be inspected without
pretending that raw internal monologue is scientific evidence.

## How a reader can inspect or falsify the forecast

You do not have to agree with a ranking. Try to break it:

1. Open the source behind a factual claim. Check its date, geography, evidence
   type and limitation.
2. Find a missing counterexample, non-English source or regional path.
3. Challenge an H2 transition: does it really connect today's H1 condition to
   the proposed H3 condition?
4. Break a world by finding two assumptions that cannot both hold.
5. Show that a supposedly new actor or need is unchanged from 2026.
6. Find a dated current-market product or published concept that already
   delivers the candidate's central
   experience.
7. Remove one enabling condition and test whether the product is still merely
   a present-day app.
8. Change the world or regional weights and see whether the chart is stable.
9. Track the published signposts over time and compare later evidence with the
   forecast's explicit conditions.
10. At the target date, map real marketplace units to the published archetype
    definitions before looking at the predicted ranks.

A falsified candidate, broken scenario or missing world is useful evidence.
The project should show the correction and preserve the earlier edition rather
than silently rewriting history.

## Current progress and limitations

The first forecast remains withdrawn. Because this process record can outlive
individual research rounds, it does not claim that a particular replacement
stage is complete. See the [live Gauntlet record](../GAUNTLET.md) for the
current stage, critic verdicts, repairs and remaining work.

Important limitations already known:

- no scan can guarantee discovery of an unknown unknown;
- public, formal and English-language evidence is uneven;
- a five-year target sits awkwardly between 2030 forecasts and longer
  2035–2050 scenarios;
- scenario coherence does not make a scenario likely;
- dated collision searches can miss private, local or newly launched products;
- exact chart positions are authored judgements, not measured future facts;
- the marketplace itself may change in ways our ontology misses; and
- imagined products can make a causal story feel more certain than its
  evidence supports.

The forecast therefore includes explicit unknowns and an “unforeseen” field.
Its purpose is disciplined imagination—not certainty.

## Concise glossary

| Term | Plain-language meaning |
| --- | --- |
| **Actor** | A person, group, machine, organisation or system able to make or carry out a choice. |
| **Archetype** | The underlying kind of product or job, separate from its fictional brand. |
| **Cross-impact** | How one change strengthens, weakens or redirects another change. |
| **Evidence cutoff** | The latest date from which factual evidence may enter an edition. |
| **Forecast world** | A coherent set of conditions that could exist together in 2031. |
| **Future-distance** | The material difference between what exists at the cutoff and the proposed 2031 experience. |
| **H1 / H2 / H3** | Today's system / the contested transition / a materially different future condition. |
| **Ontology** | A definition of what kinds of things the marketplace contains and how they relate. |
| **Collision** | A current product, service, prototype or infrastructure found in a named, dated search surface that already performs part or all of a candidate's job. |
| **Resolution** | The rule used at the target date to decide whether an observed result matches the forecast. |
| **Scenario sensitivity** | How much a rank changes when the assumed future world or weighting changes. |
| **Signpost** | A measurable development that makes a future condition more or less plausible. |
| **Weak signal** | Early, limited evidence that could become important but has not yet scaled. |

## Inspect the working record

- [Replacement method](FORECAST_METHOD_V2.md)
- [Withdrawal postmortem](FORECAST_POSTMORTEM_V1.md)
- [Gauntlet rounds and regressions](../GAUNTLET.md)
- [Horizon research index](../research/rebuild/README.md)
