AppStore2031

method

How we built the forecast

A public AppStore2031 research record. The readable view is generated without changing the preserved source.

Open the exact Markdown source

Research boundary: Observed evidence, inference, scenario and fictional forecast claims retain the labels used in the source record.

How we are building AppStore2031

Status: public process record for a forecast still in progress
Evidence cutoff: 2 August 2026 Forecast date: 31 July 2031
Method: future worlds first, products later

AppStore2031 is an educational forecast. It asks what people, machines and organisations might discover in an app store—or whatever replaces an app store—in July 2031.

The product names, developers, icons, ratings, reviews and chart positions on the finished storefront will be invented. The evidence and the route from that evidence to each invention will be inspectable. A source can show that a trend, constraint or policy exists. It cannot prove that a fictional product will be popular five years later.

This document records the method while the replacement forecast is being built. It does not imply that the future worlds, categories, products or ranks already exist. The live progress and critic verdicts are in the Gauntlet record.

The final bounded ranking and listing synthesis, including the stopped audit transport and its limitations, is documented in From sealed research to the AppStore2031 Top 10s.

Why we withdrew the first attempt

Our first attempt was careful, sourced and plausible—but it answered the wrong question. It often imagined a product that could be built now, then made it faster or added better AI, permissions, provenance or payments. Its categories also began with today's app-store categories. That made the result feel like 2026–2028 rather than a world transformed by 2031.

The ranking method strengthened this bias. Products with current users, familiar business models and working reference points looked safer, while uncertain but important changes lost points. Quality checks caught source and accessibility problems, but did not ask the crucial question: is this product materially from 2031, or is it mainly a better version of something that already exists?

We therefore withdrew the whole inventory before publication. We preserved it as a labelled audit record so readers can see the mistake, but no category, product or rank carries into the replacement automatically. The full diagnosis is in Why the first forecast was withdrawn.

This matters beyond one website. A forecast can be well cited and still be anchored to the present. Transparency should show not only the sources, but also how the research process tries to stop that anchoring.

The central rule: build the world before the product

The replacement reverses the order of work:

global evidence
    → interacting disruptions
    → several coherent 2031 worlds
    → changed actors and changed needs
    → a new marketplace shape and new categories
    → sealed product candidates
    → current-market rejection test
    → scenario and regional rankings
    → fictional storefront

Researchers generating future worlds are not supplied the old inventory, its categories or comparisons with current products. Product authors are not supplied today's app catalogue, and unbound current-market retrieval is blocked before sealing. Only after a candidate is sealed does a different researcher search named current-market and published-concept surfaces. This controls the project material used as the imagination prompt; it cannot remove model pretraining or unavoidable platform and workspace context.

A ranked item may not even be a conventional phone app. Depending on the world, the marketplace could distribute an agent, capability pack, persistent companion, synthetic environment, autonomous-organisation interface, robot service or another kind of software unit. The research, rather than a fixed quota, decides how many categories exist.

The candidate method is specified in Forecast method v2. It is still a Gauntlet candidate, not a claim that every stage has passed.

What happens at each stage

A. Scan the horizon without inventing apps

Eight workflow-separated research streams examine the forces that could reshape daily life by 2031:

  1. AI, agents, compute and human agency
  2. Economy, work and finance
  3. Health, biotechnology and demography
  4. Education, social connection and culture
  5. Energy, climate, food and water
  6. Robotics, spatial systems and mobility
  7. Geopolitics, governance and security
  8. Institutional scan of scans

The institutional scan tests the breadth and blind spots of major public forecasts. The seven domain scans then look at observations, projections, weak signals, counterforces, regional differences and open questions. They are research inputs—not predictions, categories or product wish lists.

The scans use the Three Horizons model:

  • Horizon 1 (H1): the systems operating now, including their strengths and lock-ins.
  • Horizon 2 (H2): the contested transition—investment, standards, law, behaviour, backlash, failure and repair.
  • Horizon 3 (H3): materially different conditions that could be normal by 2031.

H3 is not automatically the most likely outcome. H2 explains what would have to happen for the world to move from H1 towards an H3 condition, and what could block it.

B. Combine changes into several futures

No disruption happens alone. Cheap machine reasoning could collide with energy limits, ageing, new identity rules, climate shocks, changed money, robotics or geopolitical fragmentation. Scenario builders will receive only the horizon evidence and examine those cross-impacts.

They will build several internally consistent worlds, including constrained progress, transformative machine capability, and material institutional or geopolitical discontinuity. They will trace direct effects, then second- and third-order consequences. For example, a technology can change the price of a task; that can change an organisation; that can then change a right, duty or scarcity. The worlds are not “good”, “average” and “bad” versions of one story, and they will not be assigned fake precision just to create a consensus.

Every world must state what would make it inconsistent and which measurable signs would strengthen or weaken it.

C. Find the changed actors and needs

The next researchers will ask who can act in each world, what they control or owe, and what has become scarce, cheap, compulsory or dangerous. Actors may be people, households, communities, human–machine teams, autonomous agents, public bodies, robot fleets or infrastructure.

They will identify repeated problems that could support a discovery market. They will not start with “what app should we build?” A need must exist because the world, an actor or a relationship has changed.

D. Derive the marketplace and its categories

Taxonomy researchers will receive the changed-actor and changed-need records, not today's category list. They will ask:

  • What does this marketplace actually distribute?
  • Who—or what—chooses and trusts a listing?
  • Which jobs deserve separate listings?
  • Which recurring purposes form useful browse categories?
  • Which categories exist only in certain worlds or regions?

A provisional category must have a clear 2031 purpose, a kind of product that could be listed on its own, a trail back to changed needs, and rules that let a later assessor decide what belongs. It does not need ten different jobs. That old rule accidentally rewarded vague catch-all categories and erased useful, specific ones.

The number of categories is not decided in advance. A category reaches the storefront only if later research leaves ten genuinely different products that all survive an independent current-product audit. Ten is a publication gate, not a Stage D invention target.

E. Generate and seal future-native candidates

Product authors will work from a future-world and category packet. For each candidate they must explain, in plain language:

  • who or what uses it;
  • the new job it performs;
  • what changed by 2031 to create that job;
  • what the complete experience feels like;
  • what it must be able to do—and what would not count;
  • how it could be distributed and paid for;
  • what technical, legal, physical or social systems it depends on;
  • how it could help, harm or be abused; and
  • what would stop it from existing.

The candidate is then sealed. Its author cannot browse current products and quietly reshape it around whatever is already on sale.

F. Try to reject every candidate with a dated collision audit

Only now does a separate auditor search current app stores, official product sites, open-source projects, research prototypes and delivered infrastructure. The aim is to defeat the candidate, not to decorate it with competitors.

A candidate fails if its central experience already exists and its 2031 claim is mainly better quality, speed, price, scale, AI assistance, permissions, provenance, interoperability or a different payment rail. It may survive only if a named future change materially alters the user, job, autonomy, ownership, physical or social experience, delivery institution, or what is legally, technically or economically possible.

The auditor returns one of three results:

  • Reject: substantially delivered by the evidence cutoff.
  • Revise and re-audit: the difference may be real but is not yet clear.
  • Future-dependent; no collision found: the candidate depends on a named 2031 change and no substantial collision was found in the named, dated search surfaces.

The final wording maps to future-dependent-no-collision-found. It is not a claim that the candidate is novel, unique, first, or absent from every market.

The rules never name a favoured company, current product or technology. The market search is repeated at each refresh because the comparison boundary moves.

G. Rank by world and region, then show uncertainty

Only surviving candidates can be ranked. Each is considered separately in each world and geographic lens. The judgement includes need, reach, conditional feasibility by 2031, repeat use, distribution, infrastructure, regulation, substitutes, network effects, possible harm and evidence quality.

One fresh author context handles one category and one metric under immutable protocol causal-judgement-v8. It sees no other metric judgement or ranking. The author explains the causal path, assumptions, falsifiers and evidence limits, then chooses ordinal lower, central and upper bands. It does not assign a decimal score or probability. This stops longer, more complete-looking prose from becoming a ranking advantage.

The output must point to content-hash anchors for that exact candidate mechanism, world change and geographic constraint. It explains each link in a separate field and separates candidate, world and geographic evidence. Each link must independently retain a distinctive multi-word part from every source it claims to connect: the mechanism and delivered outcome, the general and candidate-specific world change, and the general and candidate-specific local constraint. It must then add at least five residual concepts spanning at least two causal roles. Generic words such as “stated mechanism changes the assessed result” do not count. Before detecting cloned prose, the validator removes identity labels, exact anchor vocabulary, metric wording and changing bands. It then compares both the exact residual sequence and an order-insensitive semantic form that stems inflections and groups generic synonyms. Words count only when they occur in explicit semantic evidence-field allowlists for that row's packet-bound candidate, world, lens, source and claim material. URL, identifier, provenance and other metadata never ground reasoning. All unbound words become one unbound-concept, so adding numbered placeholders or unrelated nouns cannot manufacture distinct reasoning. Each authored causal relation must connect locally grounded arguments on both sides, while every authored consequence needs a locally grounded argument; grounded words in a separate aside do not satisfy the rule. Only a role word inside an exact fully copied anchor span is exempt: repeating that same word later creates a new authored claim that must be grounded. Matching happens separately inside each causal clause, so punctuation and -> cannot reconnect copied fragments; case and commas inside one clause may still be normalised. Compact, one-sided, space-padded and tab-padded arrows are all boundaries. A single template with different anchor text, bands, grammar, synonyms or filler is therefore still a clone and is rejected alongside generic ID insertion or repeated band matrices.

The deterministic simulator samples those authored bands. It does not add synthetic noise or turn citation counts into confidence. Public results use broad simulation-share bands and integer rank ranges, not a probability that a world or fictional product will occur.

There is no editable judgement layer between authors and the simulator. A canonical projector derives it from the immutable author packets, outputs and session receipts. The public build re-reads those sources and requires the compiled file to match exactly, so changing one band, explanation, anchor or author pointer cannot silently change a rank.

The same rule now applies to the rankings themselves. One canonical projector is shared by compilation and production validation. After rebuilding the judgements, the validator replays every conditional, world, main and sensitivity chart from the sealed candidates, active audits, worlds and lenses and requires identical bytes. A receipt chain exposes hashes from the causal author source tree through the judgement file to each chart, adjacent-rank explanation and complete ranking record. Recomputing a local hash after an edit does not make that edit part of the forecast.

The finished analysis will provide:

  • a Top 10 for each category inside each world;
  • regional variations where evidence supports them;
  • a scenario-balanced main chart, clearly labelled as a judgement rather than an event probability;
  • rank sensitivity when reasonable world weights change;
  • confidence and the strongest counter-case; and
  • why each product sits above the next candidate.

A high-impact candidate must not disappear simply because its enabling world is uncertain. The site should let a reader see where it dominates, where it fails and how sensitive its rank is.

Evidence is labelled, not blended

Institutional reports often put facts, models, expert views, ambitions and scenarios beside one another. They are not the same kind of evidence. The research register separates:

LabelMeaningWhat it can tell us
Observed or estimated dataA measured baseline or documented eventWhat is happening or has happened by the cutoff
Measured trendA dated direction across observationsWhat has been changing, not whether it must continue
Modelled projectionAn output conditional on model inputsOne possible trajectory under stated assumptions
Expert elicitationWhat a sampled group expects or fearsA view worth testing, not an event probability
Policy intentA law, target, strategy or announced projectDirection and possible support, not delivery
Weak signalSmall or early evidence with larger possible consequencesA branch to investigate, not proof of scale
Scenario assumptionA condition deliberately used to build a coherent worldWhat follows if it holds, not a factual claim
Design inferenceOur explanation joining evidence to a possible productA contestable research judgement
Explicit unknownSomething the available evidence cannot settleA limitation and a future research task

Every factual premise should retain its publisher, URL, publication and access date, geography, language or translation route, limitations and status. Contradictory, rejected and superseded evidence stays visible. Popularity does not equal truth: a subject covered by many institutions is not necessarily the most important combination of changes.

Global does not mean one global average

The scan deliberately includes the United States, China and other leading technology economies, but does not treat them as the whole future. It also examines regions with different demographics, infrastructure, climate exposure, informal economies, institutional capacity and technology paths.

Each research stream records regional divergence and missing coverage. Later worlds and rankings must preserve meaningful differences rather than smoothing them into an imaginary global user. Translated evidence must identify the original language and method. Uneven evidence is itself a limitation: public institutions and English-language sources describe formal systems better than informal work, unpaid care, minority languages, low-resource settings and local trust networks.

What the fictional storefront will show

The interface will resemble a familiar marketplace so that an unfamiliar future is easy to explore. Familiar presentation must not imply that the marketplace itself is unchanged.

Every listing will separate six layers:

  1. Observed evidence: the dated facts and trends beneath the forecast.
  2. Inference: how we connect those facts.
  3. Scenario condition: what must be true in that future world.
  4. Product design: the imagined response to a changed need.
  5. Rank judgement: why it occupies this position and how sensitive it is.
  6. Imagined marketplace copy: the fictional name, sales line, rating and comments used to make the idea understandable.

The listing's future-distance panel will answer four simple questions:

  1. Why could this not be the same product in 2026?
  2. What changed in the world?
  3. What does the 2031 product let someone—or something—do?
  4. What could prevent it from appearing?

Imagined reviews can illustrate benefits and failures, but they are never evidence. Fiction will be labelled at the point where readers encounter it.

How the Gauntlet challenges the work

The Gauntlet is an adversarial quality loop. A builder produces a real artifact; a fresh critic that is not supplied the builder's conversation compares it with the external bar and identifies the single biggest gap. The builder repairs that gap, then a different fresh critic judges the revised artifact. Builders do not grade themselves, and we do not decide the number of rounds in advance. This separation controls project inputs; it does not erase model pretraining or unavoidable platform and workspace context.

The external bar combines the UK Government Office for Science Futures Toolkit, the OECD Strategic Foresight Toolkit for Resilient Public Policy, and a live, dated current-market and published-concept collision audit. Criticism checks stage separation, global coverage, counterforces, harms, temporal distance, evidence links, understandability, data quality, accessibility and real local routes.

The Gauntlet record publishes the pieces, rounds, verdicts, open gaps and regression checks. Passing a critic means the artifact beat the stated bar in that comparison; it does not turn a scenario into a fact.

How refreshes will work

A refresh will be a new forecast from the new evidence cutoff—not an edit of the previous chart. It will update the horizon scan, rebuild the interactions and worlds, rederive actors, needs, marketplace and categories, generate fresh workflow-separated candidates, run a dated current-market and published-concept collision audit, and rerank the survivors. This separation controls supplied and actively retrievable run material; it cannot remove model pretraining or unavoidable developer and workspace context.

Only after the new edition is complete will it be compared with the previous one. The public change log can then explain which categories or products appeared, vanished, merged or moved. Those comparisons never become prompts for the next generation step.

Old frozen editions will remain unchanged, so readers can inspect what the project believed at each cutoff. A method change must be versioned and pass its own Gauntlet.

What we publish—and what we do not

Credible transparency means publishing enough for someone else to challenge the result:

  • source registers and evidence labels;
  • dates, geographies, limitations and contradictory evidence;
  • stage inputs and isolation rules;
  • scenario conditions, causal links and failure conditions;
  • candidate inclusion and rejection reasons;
  • model fields, rank sensitivity and counter-cases;
  • critic verdicts, repairs, tests and known limitations; and
  • superseded editions and changes between completed editions.

We do not publish private chain-of-thought: hidden token-by-token reasoning, private scratch work or internal deliberation. That material is neither a reliable explanation nor necessary for audit. The public rationale is the structured record above—evidence, assumptions, transformations, decisions, alternatives, outcomes and tests. It is designed to be inspected without pretending that raw internal monologue is scientific evidence.

How a reader can inspect or falsify the forecast

You do not have to agree with a ranking. Try to break it:

  1. Open the source behind a factual claim. Check its date, geography, evidence type and limitation.
  2. Find a missing counterexample, non-English source or regional path.
  3. Challenge an H2 transition: does it really connect today's H1 condition to the proposed H3 condition?
  4. Break a world by finding two assumptions that cannot both hold.
  5. Show that a supposedly new actor or need is unchanged from 2026.
  6. Find a dated current-market product or published concept that already delivers the candidate's central experience.
  7. Remove one enabling condition and test whether the product is still merely a present-day app.
  8. Change the world or regional weights and see whether the chart is stable.
  9. Track the published signposts over time and compare later evidence with the forecast's explicit conditions.
  10. At the target date, map real marketplace units to the published archetype definitions before looking at the predicted ranks.

A falsified candidate, broken scenario or missing world is useful evidence. The project should show the correction and preserve the earlier edition rather than silently rewriting history.

Current progress and limitations

The first forecast remains withdrawn. Because this process record can outlive individual research rounds, it does not claim that a particular replacement stage is complete. See the live Gauntlet record for the current stage, critic verdicts, repairs and remaining work.

Important limitations already known:

  • no scan can guarantee discovery of an unknown unknown;
  • public, formal and English-language evidence is uneven;
  • a five-year target sits awkwardly between 2030 forecasts and longer 2035–2050 scenarios;
  • scenario coherence does not make a scenario likely;
  • dated collision searches can miss private, local or newly launched products;
  • exact chart positions are authored judgements, not measured future facts;
  • the marketplace itself may change in ways our ontology misses; and
  • imagined products can make a causal story feel more certain than its evidence supports.

The forecast therefore includes explicit unknowns and an “unforeseen” field. Its purpose is disciplined imagination—not certainty.

Concise glossary

TermPlain-language meaning
ActorA person, group, machine, organisation or system able to make or carry out a choice.
ArchetypeThe underlying kind of product or job, separate from its fictional brand.
Cross-impactHow one change strengthens, weakens or redirects another change.
Evidence cutoffThe latest date from which factual evidence may enter an edition.
Forecast worldA coherent set of conditions that could exist together in 2031.
Future-distanceThe material difference between what exists at the cutoff and the proposed 2031 experience.
H1 / H2 / H3Today's system / the contested transition / a materially different future condition.
OntologyA definition of what kinds of things the marketplace contains and how they relate.
CollisionA current product, service, prototype or infrastructure found in a named, dated search surface that already performs part or all of a candidate's job.
ResolutionThe rule used at the target date to decide whether an observed result matches the forecast.
Scenario sensitivityHow much a rank changes when the assumed future world or weighting changes.
SignpostA measurable development that makes a future condition more or less plausible.
Weak signalEarly, limited evidence that could become important but has not yet scaled.

Inspect the working record