method
AI process
A public AppStore2031 research record. The readable view is generated without changing the preserved source.
Open the exact Markdown sourceResearch boundary: Observed evidence, inference, scenario and fictional forecast claims retain the labels used in the source record.
How AI was used
Living process record. Sections describing the first catalogue are historical and do not make that catalogue current again. The replacement forecast uses the separated future-native workflow in
FORECAST_METHOD_V2.md; the reasons for withdrawing v1 are recorded inFORECAST_POSTMORTEM_V1.md.
AppStore2031 is an AI-assisted research and design exercise. AI did not discover a hidden dataset of 2031 apps. It helped organise public evidence, construct falsifiable app archetypes, expose disagreements and implement the site. The forecast remains a dated set of judgements.
The initiating brief
The project began with a human brief to:
- forecast the Top 10 apps in every plausible 2031 category;
- avoid assuming that today's categories all survive or that no new ones are created;
- consider China, the United States, the European Union and other leading technology economies alongside AI, energy, renewables, money, demography and social isolation;
- emulate the information rhythm of an iPhone marketplace without downloads, Apple assets or a claim of affiliation;
- invent names, descriptions and realistic reviews while labelling all of them as fictional;
- publish the evidence, method, prompts, decisions, critic verdicts, limitations and update history; and
- create a reusable skill for future evidence-led refreshes.
The human owner first chose the evergreen working identity App Store
Forecast, then deliberately renamed the public project AppStore2031 and
supplied its faceted 31 logo. The 2031 name now makes the forecast horizon
immediately clear; future editions will preserve this dated edition rather
than pretending its original prediction moved. The owner also reduced the
requested inventory from Top 20 to Top 10, expanded the original UK focus to a
global one, required public source citations, approved the full Gauntlet cost
and excluded deployment/domain registration from this build.
Withdrawn v1 role separation — historical record
The build used bounded AI roles so one synthesis did not silently become its own reviewer:
| Role | Scope | Public output |
|---|---|---|
| Current-store researcher | 2026 category, chart, product-page and legal baseline | research/baseline/ and chart snapshot JSON |
| US/EU/UK researcher | Primary-source regional evidence and contradictions | research/regions/us-eu-uk.md |
| Mainland China researcher | Policy, infrastructure, distribution, ageing, payments, energy and embodied AI | research/regions/china.md |
| Asia/Gulf researcher | Japan, Korea, India, ASEAN, Taiwan, UAE and Saudi Arabia | research/regions/asia-gulf.md |
| Multilateral researcher | Global demographic, labour, health, connectivity, energy, climate and finance drivers | research/global-drivers.md |
| Method designer | Resolution, probabilities, scoring, backtest and update contract | docs/FORECAST_CONTRACT.md and docs/BACKTEST.md |
| Inventory authors | Disjoint groups of categories using one shared schema | Edition app JSON files |
| Category calibrators | Seven fresh, disjoint assignments using probability-blind evidence packets | 621 storefront calibration records |
| App/storefront calibrators | Disjoint assignments using rank-blind packets | 460 country records, 1,620 named reasons and 2,980 explicit gap-forced zeros |
| Fresh blind critics | One new context-free reviewer per Gauntlet round | Concise verdicts in GAUNTLET.md |
| Coordinator | Scope, integration, corrections, UI, validation and public record | Repository and site |
The critic was replaced after every verdict. A critic never judged its own repair. Each losing round fixed only its single most important gap before the next blind review.
Withdrawn v1 evidence protocol — historical record
Research prompts required primary or authoritative sources where practical,
an evidence class, dates, geography, limitations, contradictions, measurable
signals and app implications. A policy target counted as intent or investment,
not delivery. Availability did not count as adoption. Every reviewed source
used in the synthesis is present in the edition-local sources.json snapshot
(with data/sources.json as the mutable working register); regional packs retain
translation caveats and explicitly excluded research.
The source register is a record of material pages reviewed, not every search result returned by a search engine. Copyrighted reports are linked and summarised; they are not republished wholesale.
Withdrawn v1 category calibration — historical record
An early draft copied regional category probabilities and then generated plausible-sounding explanations around those numbers. A blind critic correctly rejected that as post-hoc justification. The replacement process first created one packet per primary category containing its definition, boundaries, structural drivers, strongest counter-case, local evidence, evidence roles and source limitations. The packets explicitly omitted every previous category probability, lens/global coverage score, app rank and joint-model output.
Seven fresh calibrators received disjoint category groups. For each of 27 storefronts they chose a probability, compared it with the two nearest public anchors, cited only evidence present in the packet, and stated what would have moved the judgement up or down. Brazil and Mexico were judged separately; South Africa, Nigeria and Kenya were also judged separately. Machine checks require all 621 rationales to be unique, source-linked, substantive and exactly bound to the authored record before the probabilities can enter the model. The public calibration table is rendered from those records after integration; it is not the authoring surface.
Withdrawn v1 app/storefront adjustment — historical record
A late global-research critic found that separate country evidence changed category existence but not conditional app order inside the EU, Southeast Asia, Gulf, Latin America or Africa. The repair kept the twelve broad app-fit priors, then built a second packet family containing app archetypes, direct evidence, each local category judgement and its sources—but no generated rank, PMF, expected points, adjacent comparison or joint-model output.
Seven disjoint authors judged a bounded ten-app adjustment vector for each of
the 460 category/storefront cells inside those multi-country lenses. Every
evidenceful record is zero-sum, uses a 0.025 grid from -0.15 to +0.15,
cites only packet-local evidence whose registered geography names the exact
storefront country, and explains every app movement by name. Regional, global
and neighbouring-country sources remain visible context but cannot move a
value. This changes relative conditional rank without treating a policy source
as absolute adoption precision. If a packet contains no qualifying local
evidence, the record is forced to ten zeros and explicitly labels the gap.
Equal values are retained when independently authored local evidence supports
no finer distinction. Integration and edition validation independently enforce
the source, authorship, gap and no-rank-leakage rules before the v3 model
consumes the inputs.
Withdrawn v1 source-to-rank process — historical record
The process was:
- Observe the 2026 marketplace and freeze the forecast question.
- Gather regional and multilateral evidence through 31 July 2026.
- Classify 16 drivers and their counterforces.
- Build four non-probabilistic scenarios from delegation depth and ecosystem interoperability.
- Account for all 25 current categories, the Kids surface and plausible new discovery purposes. This produced 23 primary categories and five watchlist candidates; the count was not preselected.
- Define exactly ten non-overlapping, resolvable archetypes per primary category, then wrap each in a fictional product listing.
- Author 621 storefront/category probabilities from rank- and probability-blind packets.
- Author 460 multi-country storefront adjustment records from rank-blind app packets, then jointly simulate each category with ten explicit unknown- field competitors. Ranks are computed, not hand-edited.
- Attach evidence, regional variation, dependencies, counter-cases, entrepreneur openings and observable update signals.
- Run blind critic and implementation checks before freezing the edition.
The precise numerical equations, random generator, aggregation rules and
resolution protocol are in FORECAST_CONTRACT.md and data/SCHEMA.md.
Prompt record
The run-level actor, model/tool disclosure, prompt-retention status, bounded
task and output are published in AI_RUN_LOG.md. This
section explains the recurring contracts; it is not a substitute for that
ledger.
Public prompts are recorded as task briefs and constraints, not fabricated private chain-of-thought. The recurring research prompt pattern was:
Find dated primary evidence through the cutoff; separate observed deployment, projection, enacted rule and policy intent; state limitations and regional differences; identify implications, counter-cases and measurable signals; do not invent adoption or edit another researcher's files.
The inventory prompt pattern was:
Create exactly ten distinct fictional listings for the assigned category. Map each to a bounded archetype with observable capabilities and disqualifiers. Explain its adjacent rank, strongest counter-case, regional fit, evidence, dependencies, update signals and entrepreneurial opening. Include exactly two visibly imagined reviews and validate every source ID.
The blind critic prompt named the goal, external standard and regression
constraints, then requested only PASS or FAIL plus the single most
important gap. Full task-level decisions and the exact gaps found are preserved
in RESEARCH_LOG.md, DECISION_LEDGER.md and GAUNTLET.md.
What AI was not allowed to do
- Present fictional brands, ratings or reviews as observations.
- Claim that Apple publishes a global chart.
- Convert policy ambition directly into adoption.
- Hide contradictory evidence or silently remove a source.
- Rewrite a frozen edition during a refresh.
- Deploy the site, register the domain or act externally without separate human authority.
- Publish private chain-of-thought as if it were evidence.
Withdrawn v1 product opportunity pass — historical record
Browser review showed that the first panel answered “why this exact rank” more clearly than three earlier questions: what the product does, why the need could become mass-market, and how a team might build it. The human owner also set a reader standard: a bright 16-year-old should understand the new explanations and bullets without specialist knowledge.
The repair deliberately left model inputs, model hashes and ranks unchanged. Four bounded audit agents first inspected the schema, editorial quality, information order and evidence boundary. Eight separate category authors then wrote companion product-story files, with non-overlapping file ownership. The coordinator integrated and validated the results.
Each of the 230 withdrawn v1 listings received:
- a concrete benefit-led sales promise;
- a plain explanation of what the user does and who needs it;
- two or three named world shifts, linked to sources already in that app case or explicitly marked as a forecast assumption when no source directly supports them;
- a separate explanation of why demand could reach Top 10 scale;
- a first-version build route, core components, external dependencies, readiness and failure consequences; and
- an explicit
required,optionalornot-neededassessment for digital currencies and blockchain.
Evidence links in these sections support present-day premises only. They do not prove the fictional product, the 2031 demand forecast or its rank. A build dependency may instead be labelled as our design judgement. A CBDC, stablecoin, tokenised bank deposit and public blockchain are not treated as synonyms, and words such as “ledger” do not by themselves justify blockchain.
Mechanical rules cap reader-facing sentences at 28 words, bullets at 24 words and sales promises at 8–20 words. The stronger safeguard is editorial: one idea per sentence, concrete verbs, app-specific consequences and a plain translation wherever a technical term is necessary.
Withdrawn v1 limitations — historical record
- Five-year consumer rankings are highly uncertain and sensitive to short-term events, distribution policy and incumbent behaviour.
- The coverage-balanced composite values geographic breadth, not population, revenue, downloads or iPhone installed base.
- Several country-level questions inside Latin America and Africa still have thinner primary-source coverage than the US, China, Europe and leading Asian economies. The calibration therefore widens those judgements toward 0.50 where evidence is thin, while preserving separate country records.
- Model inputs remain inspectable analyst judgements; equations make their consequences reproducible but do not make the judgements objective.
- The 2021→2026 backtest could not meet its data gate. Its weaker sensitivity is published, but the proper verdict is improvement not demonstrated.
- An important 2031 capability may live inside an operating system, wallet, superapp, browser or device rather than resolving as a standalone app.
Current future-native v2 workflow
The active edition does not evolve today's apps or reuse the withdrawn v1 catalogue. It starts with six materially different 2031 worlds, identifies the people and institutions acting inside them, and derives unmet needs before any category or product is named. That produced 60 actors, 74 needs, eight proposed categories and 72 sealed candidate concepts. Six categories advanced to the storefront; two remain watch categories.
Six bounded category authors then compared all twelve sealed concepts in their assigned category. They selected exactly ten, excluded two and scored each on future distance, need scale, geographic breadth, marketplace clarity, 2031 delivery readiness, and trust and safety. Scores are authored judgements, not probabilities. Every selected listing explains what it is, why it could matter in 2031, why it holds that rank, what would change the judgement, how it might be built and which present-day evidence supports its premises. A separate bounded exact-name search found no material collision for the 60 final fictional names; this is explicitly not legal, trademark, domain or cultural clearance.
An attempted present-day semantic-overlap audit was stopped because the review
packets were too large for reliable complete inspection. Its partial output was
not accepted. This limitation is published rather than disguised as a passed
gate. The complete current method, ranking formula and limits are in
FINAL_SYNTHESIS_AND_LIMITS.md.
Future refreshes
The installed $app-store-forecast-refresh skill creates a new immutable
edition and repeats the future-world-first process. It reviews the full current
evidence base, rechecks previous signals, permits genuinely new worlds,
categories and product types, and does not begin from today's app catalogue or
copy the previous Top 10 forward. It records what changed, screens final
fictional names, runs one bounded final critic with at most one targeted repair,
and publishes an edition comparison. It never deploys.
AppStore2031