# AppStore2031 replacement Gauntlet

## Run

- Goal: Replace the continuity-biased draft with a future-native July 2031 marketplace forecast derived from global disruptions, cross-impacts, changed actors and new needs before categories or products are proposed.
- External bar: UK Government Office for Science Futures Toolkit (horizon scanning, Three Horizons, driver mapping, scenarios, futures wheels and backcasting); OECD Strategic Foresight Toolkit for Resilient Public Policy (cross-domain disruptions, assumption challenge, cross-impact scenarios and delayed solution generation); live, dated current-market and published-concept evidence as a collision boundary only.
- Approved scope: Preserve the failed draft as a superseded audit record; rebuild the method, world model, marketplace ontology, categories, Top 10 inventories, rankings, explanations, public process record, site data and both matching copies of the explicit refresh skill. Remove or replace claims and content invalidated by the new method. No deployment, domain work, freezing, publication, commit or push.
- Approved time/compute: 20-35 unattended elapsed hours with substantial web research, generation and fresh-agent criticism.
- Started: 2026-08-01 Europe/London.
- Overall status: converged as a local release candidate. No deployment,
  publication, domain work, freeze, commit or push was performed.

## Pieces

| Piece | Round | Latest verdict | Confidence | Open gap | Status |
| --- | ---: | --- | --- | --- | --- |
| Superseded-draft boundary and method reset | 1 | A | not close | none | converged |
| Global horizon scan | 2 | A across all eight streams | not close or high | none | converged |
| Cross-impact worlds and Three Horizons | final synthesis | Six conditional worlds became the sole creative input to actors, needs and candidates | high | none | converged |
| Changed actors and recurring needs | final synthesis | 60 changed actors and 74 recurring needs carried into the sealed concepts | high | none | converged |
| Marketplace ontology and new categories | final synthesis | Six categories published, two retained as watch categories | high | none | converged |
| Future-native candidate inventory | final synthesis | 72 sealed concepts; 60 selected and 12 excluded with explicit reasons | high | none | converged |
| Dated collision and future-distance audit | bounded stop | Oversized semantic-overlap packets truncated after candidate 01 | - | Disclosed as a public limitation; no partial result accepted | stopped |
| Scenario-dependent ranking and uncertainty | final synthesis | Six transparent Top 10 judgements with worlds, lenses, counter-cases and move conditions | high | Scores remain authored judgements, not probabilities | converged |
| Storefront and public process record | final release | Live desktop/mobile proof passed with six categories, ten rows, six worlds and zero browser errors | high | none | converged |
| Refresh skill v2 | final release | Repository and installed copies are byte-identical; governance and release tests pass | high | none | converged |
| Integrated local release | final critic | PASS after restart of the crash-stalled local preview | high | Not deployed or published | converged |

## Non-negotiable regressions

- [x] Candidate generation begins from future worlds, changed actors and changed
  needs, not today's categories or incumbents.
- [x] Current products and published concepts appear only after generation as a
  dated collision boundary; the incomplete broad semantic audit is disclosed.
- [x] No company, product or fashionable technology example is hard-coded into
  the reusable method or refresh skill.
- [x] A refresh rederives worlds, categories, candidates and ranks; no listing
  or quota carries forward automatically.
- [x] Every selected concept explains what changes by 2031, what it does, how it
  could exist, its harms, authority limits, counter-case and rank conditions.
- [x] The taxonomy permits new categories and marketplace units that are not
  ordinary phone apps.
- [x] Public rights, legal permission, professional judgement and final remedy
  remain with accountable external authorities where required.
- [x] Exact-name results are labelled bounded screens, never novelty, trademark
  or legal clearance.
- [x] High-impact uncertain futures remain visible as conditional worlds rather
  than being treated as inevitable or suppressed by present-day baselines.
- [x] The superseded catalogue remains provenance only and is not presented as
  the active forecast.
- [x] Repository and installed refresh-skill copies are byte-identical.
- [x] The final verdict came from a fresh critic inspecting the real live output.
- [x] No deployment, domain configuration, freeze, publication, commit or push
  occurred.

## Round log

The entries below are a chronological audit trail and intentionally preserve
the status language that was true at each round. Current status is the Pieces
table and Final evidence above and below, not an earlier “pending” sentence.

- Round 0 — the July 2026 continuity-led edition was marked `superseded`, removed from the active chart UI and retained only as an audit record. `FORECAST_POSTMORTEM_V1.md` records the failure mechanism. The old refresh skill was placed behind a read-only safety gate in both copies.
- Round 0 — `FORECAST_METHOD_V2.md` established a workflow-separated sequence: disruptions, cross-impact worlds, changed actors and needs, ontology, candidates, then a dated collision audit and scenario ranking. Supplied/retrievable run material is controlled; pretraining and ambient developer/workspace context are disclosed.
- Round 1 — eight workflow-separated horizon researchers were launched without supplied old inventory, taxonomy or current-product comparisons. AI/agents/compute and education/social/culture completed; fresh context-separated criticism followed. Pretraining and ambient developer/workspace context were not removed.
- Round 1 — the v2 method beat its UK/OECD external bar with `Winner: A`, `Confidence: not close` and no gap. The withdrawal/separation boundary received the same verdict.
- Round 1 — AI/agents/compute, economy/work/finance, education/social/culture after one repair, and the institutional scan-of-scans passed fresh criticism. Education's first critic exposed an implicit H2 hand-off; the author added a transition map and Round 2 passed `not close`.
- Round 1 — climate/energy, health/biotech, governance/security and robotics/spatial each lost on the same single actionable gap: strong H1, weak-signal and H3 material without an explicit H2 transition map. Each owner is repairing only that gap before a fresh-context Round 2.
- Round 2 — climate/energy, health/biotech, governance/security and robotics/spatial each added the missing H2 mechanisms, actors, blockers, regional branch points and bounded Stage B hand-off. Four new workflow-separated critics returned `Winner: A` with no remaining gap. All eight Stage A streams now converge.
- Round 1 — the Stage A atlas lost because 38 sampled links did not account for all 496 variable pairs. The repair published a four-class pair ledger, 80 evidenced direct interactions, 24 evidence-gap pairs and reconciled all 496 pairs; a fresh critic passed it at 0.96 confidence.
- Rounds 1–8 — the machine-readable source registry repeatedly lost on extraction fidelity: decimal splitting, translation/limitation separation, slash-separated time horizons, publisher contamination, time ranges and domain labels misclassified as geography, robotics title/publisher/date boundaries and finally robotics evidence-family token mapping. Each repair addressed only the surfaced gap and added a regression. Round 8 now reconciles all 49 robotics rows and class totals A=17, B=38, C=7; a ninth fresh critic is inspecting it.
- Rounds 9–10 — the ninth critic found two strategy periods misread as publication dates. S24 and S25 now preserve those periods only as `temporalScope`, qualify publication date as unknown and reconcile the public totals. A tenth fresh critic independently reproduced 378 source rows, 440 retained links, 363 unique primary URLs, all eight register totals, hashes, numeric tokens, geography, evidence families and datedness; it returned `Winner A`, no gap, 0.99 confidence.
- Round 1 — the v2 schema and public process record each passed fresh criticism with no gap. The schema removes incumbent/current-category strength anchors and models worlds, needs, ontology, sealed candidates, later prior art and conditional ranks.
- Stage B generation — three workflow-separated builders produced six-world sets from the hash-bound Stage A packet. One critic selected Set B for its complete audit closure; another selected Set C for its discontinuities, regional plurality and physical realism. The canonical synthesis contains six distinct causal worlds, the complete 32-variable by six-world and 16-transition by six-world audits, all nine regional lenses per world, no probabilities and no product/category leakage. A fresh critic returned `Winner: A`, no gap, `not close`. Its two artifacts are sealed in `STAGE_B_PACKET.json` as the Stage C creative inputs.
- Refresh-skill v2 tests — the helper and governance tests now prove an empty edition scaffold, no inherited inventory/evidence/model output, sealed candidates before dynamic prior-art search, no named present-day product rules and separate research, seal and publication authority. Focused tests pass 4/4 and the full suite passes 76/76; independent quality judgement remains.
- Store experience specification — Round 1 passed narrowly but found that the build explanation would sit too far down the detail view. The repair added a compact first-viewport `What it does / Why 2031 needs it / How it could exist` brief linked to the full validated sections. A fresh Round 2 critic returned `Winner A`, no gap, 0.96 confidence.
- Refresh-skill v2 runnable path — Round 1 criticism found an incompatible scaffold and v1-only validation path. The repair aligned the v2 manifest and directory layout, added honest prepared-stage validation and method-aware edition dispatch, and raised the full suite to 77/77. Round 2 then found the helper still validated its source edition through the legacy validator, preventing a future v2-to-v2 refresh; that single gap is being repaired and will receive a new critic.
- Stage C changed actors and needs — three workflow-separated builders covered all six worlds with 103 world-specific actors and 114 recurring jobs. First critics found three bounded omissions: equipment makers and grid operators in the physical worlds, caregiver continuity when relationship records disappear, and the digital-to-physical handoff after back-office control chains shrink. Each source packet was repaired and a new critic returned `Winner A`, no gap, `not close`. The canonical synthesis preserved all 114 needs, 103 actors, 77 institutions, 142 resource shifts and 54 world-region lenses. Independent structural and semantic critics both returned `Winner A`, no gap, `not close`.
- Refresh-skill v2 workflow-separation controls — Round 3 showed packet separation was still self-declared. Later repairs created hash-bound packets and task receipts, deterministic candidate-tree seals, fail-closed session reconstruction and exactly six externally reconstructed fresh critic receipts. Round 13 passed 127/127 tests; Round 14 criticism is in progress. The method wording now avoids claiming that workflow controls erase pretrained or ambient context.
- Stage D reset — after seven threshold-tainted ontology rounds, a fresh protocol critic chose the corrected rule with 0.99 confidence. The number of different jobs was rewarding catch-all categories and erasing precise categories before product research. All proposals, repairs, round packets/receipts, outputs and five active repair tools were moved to a recoverable 36-file hash manifest. Stages A–C stayed byte-for-byte unchanged.
- Corrected Stage D hand-off — the first corrected packet removed the numeric threshold and directly hash-referenced the canonical actor-needs and world artifacts. Its archive mechanics and hashes passed, but later comparative criticism showed that the upstream need set was not broad enough to support a complete successor app store.
- Stage D incomplete-input verdict — two independent critics rejected all three ontology proposals as a complete taxonomy. Both found a procurement and commissioned-service catalogue shaped by strong institutional needs but missing ordinary lived experience such as relationships, creativity, play, culture, everyday communication and self-directed activity. This is an upstream input gap, not a category merge problem.
- Stage D incomplete-run archive — the incomplete packet, reset note and six proposal files were moved byte-for-byte to a provenance-only eight-file archive with a 310,817-byte hash manifest. Stage C2 then added 30 recurring lived needs: 23 eligible, four provisional and three context-only. The 114 institutional needs remained valid and were combined with lived experience in a hash-bound input packet.
- Stage D repaired ontology — R3 reconciles all 144 needs exactly once into 27 marketplace-unit types and 23 categories. Nineteen categories are provisional and four remain on the watch list because their boundaries are not yet strong enough for product authoring. The canonical package passed independent structural and market criticism only after the R1 failures were retained and the R2 critics rechecked the repaired bytes through a non-circular verification envelope.
- Stage E packet preparation — the initial edition prepared 12 alternatives for each of the 19 provisional categories: 228 immutable candidate briefs and zero watch-category briefs. Every brief binds its category, future worlds, institutional and lived need context, marketplace unit, material-difference axis, authority boundary and source receipts. The canonical prepared-edition validator passed before any candidate was written.
- Stage E execution correction — the first proposed plan placed three or four categories in one author context. The rendered delivery for the first group measured 3,464,126 bytes and required one session to produce 36 rich candidates, so it was rejected before dispatch. The replacement uses one fresh context per category, deduplicates repeated packet evidence without changing the 228 immutable packet receipts, and size-checks the delivery. After exact authorship, evidence and dependency allowlists were added, the largest category delivery is 183,472 bytes. Nineteen one-category groups validate.
- Stage E author-wave failure — the first five category contexts exposed a path-semantics defect: edition-relative output roots were applied from the project root. Forty-six partial files were written outside the edition, none entered canonical data, and the deterministic validator rejected the completed attempt. All five contexts were retired and the invalid root was moved to `research/rebuild/archive/rejected-stage-e-wrong-output-root-2026-08-02/`. Author dispatch is paused until every delivery contains an exact absolute `candidate.json` path and a project-root application/readback regression passes.
- Stage E receipt correction — the first corrected wave then exposed three transcript and metadata boundaries before sealing. One otherwise schema-valid context was retired for using an ad-hoc parser outside the exact call allowlist. Three integrated groups were later retired because their author-supplied `authoredAt` placeholders predated the externally verified tasks. All content and receipt bytes remain in hash-manifested failed-run archives. The new contract supplies exact authorship/evidence/dependency values, hashes substantive content independently of integration metadata, and mechanically replaces `authoredAt` with the verified task completion time while adding the shared receipt. Focused tests passed before a fresh six-category wave.
- Stage E protocol-version correction — hardening the author instructions while older sessions were still running changed the reconstructed delivery bytes and made compliant historical transcripts look invalid. The receipt system now records and rebuilds an immutable named protocol version (`legacy-v1`, `readiness-v1` or `hardened-v2`) for every session instead of comparing all sessions with the latest prompt. Old protocol bytes are frozen; future changes require a new version. This recovered honest verification without grandfathering bad behaviour: a legacy category whose author independently added a JSON parsing command still failed the allowlist and must be rerun from a fresh context under `hardened-v2`.
- Stage E-to-F integrity convergence — the lifecycle gate now reopens the exact live Stage E packet, all eight ordered allowed inputs and exactly the candidates, collisions and seal outputs before Stage F. Process-crash and POSIX fsync recovery are exercised. A fresh critic then found that identical external bytes could enter through final or parent symlinks; segment-walking real-file reads and transaction/recovery checks closed that gap. The final fresh verdict returned `Winner A` at 0.99 confidence with 30/30 focused tests.
- Stage F audit-v3 convergence — five adversarial rounds hardened non-English evidence and search attribution before any replacement candidate research. An authoritative transcript-bound local-to-English relation is required; false juxtaposition, capability-ID relabelling and cross-anchor relabelling of the same normalized relation fail group-globally. Each query has one search and one direct open, and exact result ordinal plus canonical record hash binds the opened URL. The final fresh verdict returned `Winner A` at 0.99 confidence. This validates the audit mechanism only; the invalidated candidate set was never audited.
- Stage H publication-chain repair — a fresh critic found that internally consistent edited drafts and later name-audit receipts could bypass the original packet-bound author work. The draft now carries a complete source-set tree receipt, compilation rederives every packet from sealed Stage E, Stage F and Stage G inputs, and canonical compile/final validation re-create the exact normalised draft from author outputs plus reconstructed session receipts before applying only the audited fictional identity choice. Seven negative mutation classes are regression-tested: prose, anchor, review, dependency, fingerprint, session pointer and rank.
- Stage H name-audit convergence — three adversarial rounds removed output-order selection, unreachable normalized near-match detection and ambiguous search-to-open attribution. The final repair derives a bounded transcript-fragment span from the searched identity, so a five-word provider with a one-edit collision cannot clear a four-token scan. Exact result ordinals and canonical record hashes bind each one-query/one-open check. A final fresh critic returned `Winner A` at 0.98 confidence; its focused suite passed 8/8.
- Stage G semantic-template repair — a fresh adversarial critic proved that `causal-judgement-v2` accepted 864 grammatical and synonym paraphrases of one generic template after exact row anchors and bands were swapped. Protocol `causal-judgement-v3` added source-by-source anchors, five residual concepts across two causal roles, and exact plus order-insensitive semantic fingerprints. Its fresh recritic then proved that arbitrary numbered placeholders and unrelated nouns could still manufacture residual novelty. Protocol `causal-judgement-v4` limited distinct credit to row-bound stems, but a third critic found two remaining scope leaks: URL metadata could enter the grounding corpus, and a grounded aside elsewhere in a clause could excuse ungrounded causal arguments. Protocol `causal-judgement-v5` added semantic-field allowlists and local argument binding; a fourth critic then proved that a later ungrounded role occurrence inherited exemption merely by sharing a stem with copied anchor text. Protocol `causal-judgement-v6` bound exemptions to copied-anchor token positions, but a fifth critic proved that token normalisation reconnected anchor fragments across full stops, including when inserted two-letter words vanished. Protocol `causal-judgement-v7` split causal clauses before copied-span matching, but a sixth critic showed its arrow splitter recognised only arrows padded on both sides. Protocol `causal-judgement-v8` treats compact, both one-sided, space-padded and tab-padded `->` forms as identical clause boundaries while retaining case/comma normalisation inside a clause. The grammatical-paraphrase, numbered-filler, real-noun-filler, grounded-adjunct, `relatedUrls` sentinel, repeated-anchor-stem, punctuation-splice and five arrow-spacing attacks plus genuinely diverse positive rows are permanent regressions; canonical judgement and ranking reprojection remains unchanged.
- Stage A–D discontinuity recritic — after all 228 draft candidates were authored but before collision sealing, a fresh critic applied the user's original objection as a strict gate. It passed the global scope, jobs-to-be-done ontology, sovereignty, care, relationships, climate and non-carry-over boundaries, but failed the world set at 0.91 confidence because no scenario explicitly stress-tests possible AGI, near-zero cognitive-labour cost, large employment discontinuity, ownership and aggregate-demand change, or contribution and belonging after paid work. Stage E stopped immediately. The current candidates remain failure-history evidence and cannot proceed to prior-art research, ranking or publication.
- Failed-run verdict record — a new critic inspected only the old Stage A-D artifacts and wrote a machine-readable `FAIL` at `B-worlds`, with 27 exact path/byte/SHA-256 receipts. It preserves the old run's scoped passes while prohibiting default carryover of its worlds, needs, ontology, categories or candidates.
- Possible-AGI evidence repair — 34 primary or official sources now support a bounded, non-canonical W07 handoff: `Abundant Cognition, Scarce Human Standing`. Four capability gates plus two of three diffusion gates classify the scenario without assigning a probability. The handoff covers economic substitution, AI-research acceleration, work and ownership shocks, contribution after paid work, physical scarcity, nine regional expressions, countercases and redirects. A provenance critic first rejected loose evidence pointers; exact path/byte/SHA-256 receipts repaired the gap, and the final fresh verdict returned `Winner A` at 1.00 confidence.
- Social-connection evidence repair — 35 primary or official records keep loneliness, isolation, living alone, mental illness, trust, belonging, AI attachment and work meaning separate. Relational-AI augmentation and substitution remain open branches. Five person-invoked need hypotheses define observable completion and authority boundaries without promising friendship, treatment or belonging. The first critic required exact source IDs in the Stage C handoff; after repair, the fresh verdict returned `Winner A` at 0.99 confidence.
- Physical-constraint evidence repair — 15 primary or official sources distinguish measured deployment, projection, policy intent and company plans across compute, grids, energy, chips, robotics and digital-to-physical delivery. Three needs enter the strongest Stage C slice: scarce machine-capacity access, bounded physical proxy work and trusted design-to-real-object handoff. A critic rejected an unsupported claim about robot place-data reuse; that hypothesis moved to Watch with explicit missing-evidence requirements. The final fresh verdict returned `Winner A` at 0.98 confidence.
- Clean-room Stage B repair packet — the first replacement packet was rejected because allowlisted repair files still exposed failed-run world and downstream need cues. Eight evidence-only supplements now retain all 84 cognition, social-connection and physical-constraint source identities, limitations and counterevidence while excluding old world IDs, world names, counts, needs, categories, products, ranks and current-market cues. All 19 input receipts matched; a 403-check leakage scan found zero matches. A fresh critic returned `Winner A` at 0.98 confidence. World count, names and IDs remain for the new Stage B authors to justify.
- Failed-edition invalidation authority — the first invalidation writer was rejected because any synthetic `FAIL/B-worlds` text could have withdrawn the edition. The repaired boundary requires an edition-bound authorisation over the exact manifest and refresh-run bytes, reached stage, R1 JSON and Markdown bytes, canonical ordered 27-artifact verdict receipt set, integration receipt, and reconciled live Stage B-D outputs. It rejects wrong-edition, wrong-prestate, changed-verdict, altered-receipt and unrelated-run attacks before transaction staging. A fresh critic returned `Winner A`, high confidence, not close. The first real transition attempt then stopped before writing because the pre-Stage-F edition legitimately lacked a Stage F directory; the corrected guard treats absent or empty as not started while continuing to reject files, symlinks and non-empty state. Its fresh recritic returned `Winner A`, high confidence, not close, with 19/19 focused and 5/5 independent adversarial checks. The exact 14,922-byte authorisation then invalidated edition `2031-2026-08-01` at its reached `candidate-packets-prepared` stage, preserving 806 protected files and publishing zero categories or listings from it.
- Clean replacement edition — `2031-2026-08-02` was created only after the invalidated source revalidated. It is a `prepared` research scaffold with `basedOnEdition: null`, no valid forecast predecessor, no comparison eligibility and explicit `catalogueCarryoverPermitted: false`. No categories, candidates, ranks, listings, current-product data or fake freeze receipts were copied. Its only source relation is the hash-bound invalidation disposition of the failed research edition.
- Dynamic world-and-need contract — the new protocol derives contiguous lower-case internal world IDs and upper-case author-facing IDs from the selected world array rather than a global count. It accepts independently justified 4-, 7- and 11-world fixtures, binds exact artifact paths/bytes/hashes/order, treats need counts as observations and rejects gaps, swaps, duplicates, forged paths, layer changes, unknown worlds, dropped needs and cross-layer overlap. The first critic found receipt-derived path trust, missing per-record layer checks and ambiguous ID case; after repair a fresh recritic returned `Winner A` at 0.99 confidence with 14/14 focused tests. The exact six-world, 114-institutional and 30-lived constants remain only in the frozen legacy protocol.
- Edition-local research boundary — schema-2 integration now snapshots every allowlisted input and validates it by exact source path, edition path, bytes, SHA-256 and ordered set. Validator and Stage E read only those local bytes. Three critic rounds closed missing-receipt and schema-version downgrade fallbacks plus an output-set self-attestation gap; expected outputs now come from the independent edition lifecycle, accept dynamic 4/7-output probes and cannot be removed or substituted by recomputing receipt/manifest hashes. The final fresh critic returned `Winner A` at 0.97 confidence with 46/46 focused checks, typecheck and all-edition validation. The invalidated schema-1 history stayed byte-identical and validator-clean.
- Stage G causal-grounding convergence — eight protocol rounds progressively closed generic paraphrase, numbered-filler, unrelated-noun, metadata-vocabulary, grounded-adjunct, copied-role-stem, punctuation-splice and compact-arrow bypasses. `causal-judgement-v8` builds grounding only from semantic evidence fields, requires locally grounded relation/consequence arguments, exempts only exact copied anchor spans inside one clause, and splits ASCII `->` regardless of surrounding whitespace. The final fresh critic returned `Winner A` at 0.99 confidence; genuine 864-row output, session reconstruction, canonical replay and deterministic rehash still pass. This validates the mechanism only; it did not rank the invalidated candidates.
- Fresh Stage B alternatives — three closed-allowlist authors independently selected four, five and six worlds from the clean 25,625-byte packet. Their first critics surfaced false every-world social-connection passes in all three sets. Repairs now keep loneliness, social isolation, living arrangements and mental illness distinct; test relational-AI help and harm; and trace contribution, status, time structure and belonging in every world. Alpha also corrected four evidence-class labels and added plain summaries/glossary. Fresh recritics returned `A` for Alpha, `A` at 0.97 for Beta and `A` high-confidence for Gamma. All remain non-canonical while a separate synthesizer compares causal configurations and justifies the minimum sufficient count.
- Fresh Stage B synthesis — the synthesizer mapped all 15 proposed worlds rather than voting. Three repeated clusters cover supervised work, broad-ownership post-labour settlement and concentrated-ownership demand fracture; Alpha's bundled constraint world splits into maintained physical renewal and hardened service corridors; Gamma's relational-mediation branch survives only behind a four-conjunct placement discriminator and an auditable five-edge self-reinforcing loop. Two critic rounds required that W06 prove an observable dominant closure, replace 14 ambiguous short source IDs with unique namespaces, state its fifth invalidator accurately and expose every loop link as structured data. Final count criticism returned `Winner A` at 0.95 confidence and final coverage criticism returned `Winner A`, high confidence. The synthesis has six minimum-sufficient worlds, not an inherited six-world target; it remains unweighted and contains no category, product, rank or current-market material.
- Reviewed Stage B binding — a deterministic compiler projects the reviewed synthesis to lower-case internal `w01`–`w06` records while preserving upper-case source snapshots and a dynamic world-set receipt. Its first critic proved that rebound A verdicts could hide numeric probability/downstream fields and that hardlinked inputs were accepted. The repaired compiler uses fail-closed semantic-key boundaries, exact single-link inputs, atomic no-overwrite commit and single-link final readback. A fresh recritic returned `Winner A` at 0.995 confidence; independent four-world fixtures still pass, proving that six is not compiled as a global target.
- Fresh Stage C authoring — three isolated authors covered W01–W06 without category, product, current-market, probability or rank material. Their final outputs contain 60 changed actors, 41 institutional needs and 33 lived-experience needs. Those counts are observations, not quotas. Critics required direct-versus-scenario evidence boundaries, harm-specific recovery, real dependency explanations, geographic humility and natural bright-16 prose. Final author verdicts were A at 0.999, 0.99 and 0.98 confidence.
- Fresh Stage C synthesis — a production approval anchor now binds the exact reviewed Stage B inputs and cannot be replaced by callers. Stage C preserves every actor and need one-to-one; 60-to-1, 41-to-1 and 33-to-1 “picnic” collapses fail. Grouping belongs only in Stage D. The fixed-path generator produced a byte-identical proposal and synthesis as distinct single-link files plus 41-need and 33-need layer documents and a receipt bundle. A blind critic preferred the implementation to the external contract bar `not close`, with no gap; 249 focused checks passed.
- Fresh Stage D input — the invalidated 114/30 need set, old ontology, old categories and their critics are mechanically excluded. A new code-owned Stage C approval binds seven exact canonical/verdict receipts. The input contract dynamically derives six worlds, nine lenses, 129 evidence IDs, 60 actors and 74 needs, then gives the same full universe to three isolated authors. The internal proposal contract passed fresh criticism only after closing forged evidence, non-prefix world, self-owned authority, negated capability, duplicate-category and collision-mapping attacks. No proposal may become publishable without a later independent A semantic/readability verdict.
- Stage D assignment-handshake failure — all three isolated builders independently proved that the first prepared assignments were impossible to satisfy: their trusted registry exposed only `trusted-*` owner IDs while the proposal validator required those same IDs to begin `authority-*`, `appeal-*` and `remedy-*`. A separate hostile review graded the prepared set B and also reproduced self-authorised forged evidence through an incompletely revalidated Stage C chain, a late reserved-proposal-path collision that did not abort the packet commit, and semantic author verdicts that were not bound exactly to the replayed author outputs/check set. The builders were stopped before any proposal was accepted. The four prepared input files remain byte-for-byte failure evidence; a corrected set must use new versioned paths, close all four attacks and pass fresh criticism before authoring restarts.
- Stage D version decision — `internal-dynamic-stage-d-proposal-v3` is the explicit corrected successor to the unsatisfiable proposal v2 contract. The preserved root-v1 packet, assignments and failure verdict are mechanically denied as inputs. The corrected input pipeline uses its own versioned `/v2` directory and must bind every builder assignment to proposal v3; the different numbers describe different contracts, not an attempt to relabel failed bytes.
- Corrected Stage D handoff — input-v2 now replays the exact Stage B and Stage C approval chains, author outputs, semantic verdict rows and check keys; derives its evidence allowlist only from reviewed Stage B; rejects the failed root paths; keeps all proposal paths reserved through final commit; and exposes satisfiable external authority, appeal and remedy role placeholders without pretending to know the names of 2031 institutions. The exact 71,955-byte packet and three 2,846-byte assignments passed public byte-for-byte replay and fresh independent criticism at A/0.995. Focused checks passed 63/63, the full cross-contract run passed 319/319, and no proposal existed at acceptance.
- Dynamic lens storefront boundary — the inactive v2 storefront no longer requires or labels a fixed 12-lens set. It accepts a non-empty unique validated lens array and renders the authored count. This prevents the failed method's geography count from becoming a hidden future schema target; focused storefront checks and typecheck pass.
- Stage D proposal A R1 — the 735,748-byte proposal passed structural v3 validation with six retained categories, 74 one-to-one primary need dispositions and 15 pairwise records. A fresh semantic critic then graded it C at 0.99 confidence because all 370 directional need comparisons used the same match flags, mechanically forcing every pair to `distinct`. The exact bytes and verdict are retained; the proposal is excluded from synthesis unless a new-path repair reconstructs real substitution logic.
- Stage D proposal C R1 — the 744,605-byte proposal passed structural v3 validation with seven retained categories, but its semantic critic returned C at 0.99 confidence. All 74 reconciliation rationales repeated the source job plus one of seven category taglines, hiding incompatible observable finishes inside broad workflow buckets. The exact bytes and verdict are retained and excluded; a new-path repair must first recut human-handoff and physical-restoration families around need-specific mechanism, completion, liability and recovery.
- Stage D proposal B R1 — the 813,274-byte proposal passed structural v3 validation with eight retained categories, then received C at high confidence. All 306 regional cells claimed direct evidence, reused world/lens bundles across unrelated categories and sometimes admitted comparable evidence was absent in the same cell. The exact bytes and verdict are retained and excluded; a new-path repair must rebuild the matrix from `gap` upward and reserve `direct` for category-and-lens-specific evidence.
- Stage D proposal B R2 — the regional repair produced 0 direct, 114 adjacent, 84 inference and 108 gap cells and passed exact structural validation, improving the fresh semantic grade from C to B at 0.98. The remaining blocker is claim-level qualification: all 34 causal paths still say direct despite inference/projection/policy-intent/scenario-condition source chains, while cell status still follows citation count. R2 remains excluded; R3 is limited to causal/regional evidence blocks and their claim-specific limitations/unknowns.
- Stage D proposal C R2 — recutting two collapsed categories expanded seven broad units into fourteen genuinely distinct callable finishes; the fresh critic confirmed unique need sets, bounded work and observable finishes. The grade improved from C to a close B. The remaining blocker is limited to 148 reconciliation prose fields: 67/74 boundary rationales overlap their main rationale by more than 80 percent and at least fourteen pairs contain mangled template sentences. Categories and units are frozen for a bright-16 R3 rewrite.
- Stage D proposal A R2 — rebuilding 370 collision mappings with varied source-set sizes and fourteen flag signatures improved the proposal from C to B at 0.99, but a fresh critic found true flags unsupported by the named source needs. One mapping used climate-service outcomes plus income-record correction to claim completion of supervised learning and demonstrated skill. R2 remains excluded; R3 must re-author collision truth from explicit observable finishes and accept any resulting merge or reassignment.
- Stage D proposal B R3 — claim-level evidence regrading passed exact structural validation with no direct claims, but the fresh semantic critic returned C at 0.99 on deeper no-loss grounds. Fifteen multi-need causal paths covering 55 need references used only the first need's observable completion. A W06 path cited six needs while preserving child safety but losing safe disagreement, cross-group help, trust checking and dispute repair. R3 remains excluded; R4 must split/group paths by compatible finishes and carry every cited need's specific harm and recovery.
- Stage D proposal A R3 — the full-text audit removed all 437 unsupported true flags, but replaced them with a blanket false strategy. A fresh critic returned B at 0.99 and demonstrated a mapping where completion differs while authority, liability and recovery fully match. R3 remains excluded; R4 must score each of 370 mappings dimension by dimension and name the precise residual behind every false flag.
- Stage D proposal C R3 — the 148-field reconciliation rewrite passed its target: a fresh critic human-read all 74 pairs and found them coherent. The proposal remains B at 0.99 because the next largest readability defect sits in previously frozen material: 17/27 causal paths and 37 regional uncertainty cells contain mangled template splices such as `requires The` and `unknown how whether`. R4 is limited to those sentences; accepted categories, fourteen units, reconciliations and collisions remain frozen.
- Stage D proposal B R4 structural stop — the author fixed semantic loss by splitting 34 combined paths into 74 single-need paths, but exact root replay rejected the 1,247,824-byte file because proposal v3 requires exactly one causal-path object per applicable world. This is a contract defect as well as an artifact failure: the one-path shape encouraged incompatible finishes to collapse into a single string. R4 is preserved and barred from semantic criticism; the contract repair waits for in-flight A/C work to reach a safe boundary.
- Stage D proposal C R4 — a blind critic independently revalidated the exact 731,212-byte proposal and returned A at 0.93 with no blocker. The accepted synthesis input contains seven future-native categories, fourteen callable units, 74 reconciliations, 27 causal paths, 243 regional cells and all 21 pairwise boundaries. It remains non-canonical and non-public until cross-proposal synthesis and the repaired causal-path contract complete.
- Stage D proposal-v4 replay — all three alternatives validate against the new
  input-v3 assignments, but none inherited an old semantic verdict. Fresh blind
  v4 critics returned B at 0.995 for A and 0.99 for B and C. A shortened 52
  multi-sentence completion tests to their first step; B's broad public category
  boundaries did not entail several need-specific finishes; C lost two worker-
  agency remedies and used mechanically uniform regional/collision reasoning.
  Each exact file remains excluded. Separate R5 assignments repair only the
  critic-bound gap and require another fresh exact-byte verdict.

- Builder B R5 repaired three broad category boundaries and 27 broken
  rationales without changing the 74 need assignments. A fresh critic read all
  74 entailments, 306 regional cells, eight callable units and 518 collision
  mappings, then returned A at 0.97. Eight categories are now an accepted
  synthesis alternative, not a canonical count.
- Builder A required an adjudication after two critics contradicted one another
  on one collision dimension. The adjudicator ruled completion true but
  authority, liability and recovery false; neither critic's full interpretation
  survived. A later blind audit found four needs still exceeded category-001's
  finish and six other mappings overclaimed partial dimensions. R8 moved those
  needs into provider-exit and safe-adoption boundaries, removed 99 of 100 prior
  true collision flags and preserved only the adjudicated completion overlap.
  Its fresh critic returned B at 0.995 after finding nine shortened public
  finishes, a disjunctive human-service/certification category and twelve stale
  pair summaries. R9 split staffed human service from safe-system adoption,
  restored every shortened finish and rebuilt all 21 pairs. A new blind critic
  reviewed 74 need entailments, 27 causal paths, 234 regional cells and 444
  directional mappings and returned A at 0.97. Seven categories are therefore
  a second accepted synthesis alternative, not a canonical count.
- Builder C R6 removed 80 collision false positives but still closed 19 needs
  before the person's full result, left one social route buildable with present
  tools and published fragmented audit prose. R7 added executable finishes,
  reframed the social watch category around cross-service reciprocity that must
  continue with AI switched off, rebuilt all 444 mappings and reduced collision
  explanations from a median 89 words to 30. Its fresh critic returned B at
  0.99 because eight paths and ten units still used shorthand, two classes were
  credible before 2031 and the W02/W06 branch test was conjunctive. R8 merged or
  rejected those classes, preserved all 74 exact paths and separated the world
  branches, but a fresh hostile critic returned B at 0.998. Forty-four of its
  74 units hid an unrelated seed outcome merely to satisfy the structural rule,
  and one closest-substitute collision was not actually tested. R9 reduced the
  inventory to 20 units, repaired the named collision and capped completion text
  at 59 words. Its full audit still found 16 incomplete bindings across 11 units
  and 293 of 296 collision mappings using one generic template, with 73 priority
  plausible-overlap cases. The B at 0.999 verdict excludes C R9. We do not force
  a third alternative or a five-category result; the exact failures remain as
  process evidence.
- The first synthesis-contract implementation failed hostile review at C/0.999
  despite passing its own 11 tests. Its positive fixture was not a valid v4
  proposal. Attacks could self-author an A, hide present-product/quota material,
  drop an accepted watch category, leave a world with no category or typed gap,
  duplicate collision records and turn the observed 74 needs into a refresh
  quota. R2 now passes fresh hostile review at A/0.995 inside its explicitly
  injected internal-authority boundary. A code-owned production wrapper now
  replays the reviewed sources and exact A-grade A R9/B R5 alternatives, owns
  the canonical assignment and accepts only `{projectRoot}`. Its first hostile
  critic found alternate-path relabelling and lens-universe substitution; R2
  path-locked candidate receipts and derives the exact ordered lenses from the
  approved input packet. Fresh re-review returned A at 0.995, scoped only to the
  wrapper boundary; the wrapper still cannot approve a synthesis report itself.
- Canonical synthesis R1 preserved all 74 needs, six worlds, eight retained
  boundaries and 28 pairs, then failed both fresh gates. The structural critic
  returned C at 0.999 because newly receipt-bound report prose could still carry
  a named current app or invalidated inherited ontology. The contract now rejects
  those payloads recursively. The semantic critic returned B at 0.999 because one
  A-source disposition omitted institutional-need-011, the canonical collision
  matrix dropped the adjudicated institutional-need-020 to institutional-need-009
  completion overlap, and 162 stale pre-rename category titles remained. R1 is
  immutable failure evidence; R2 is a new path and must receive two fresh A grades.
- Canonical synthesis R2 repaired all three semantic defects and the structural
  report gate. Exact replay preserves 74 needs, six worlds, nine lenses, eight
  categories, eight units and 28 pairs. Fresh structural criticism returned A at
  0.997 after 13 hostile attack groups; fresh semantic criticism returned A at
  0.995 after 15/15 source dispositions, 518 mappings, the adjudicated overlap
  and all eight stale-title digests were replayed. A separate code-owned canonical
  approval is now the only remaining Stage D gate.
- The separate canonical approval replays R2 and both independent A verdicts,
  binds 54 nested critic receipts, keeps failed R1 non-authoritative and exposes
  only `canonical-approved-for-stage-e`. Its hostile critic returned A at 0.995
  with zero blockers. App content, inventory, ranking, storefront publication,
  deployment and domain permissions are all explicitly false. Stage E must call
  only `validateReviewedStageDCanonicalApproval({projectRoot})`.
- Edition-local Stage E hand-off — the direct-v4 research integration and
  recoverable lifecycle transaction advanced the clean edition to
  `ontology-complete`. A real preparation run then found a schema-boundary bug:
  shared worlds were read from the approval identity instead of the hash-bound
  need record. The bounded repair and regression passed, and Stage E prepared
  72 immutable clean-room briefs across six provisional categories while
  keeping both watch categories out of publication dispatch.
- Refresh-governance repair — the project skill and installed
  `$app-store-forecast-refresh` copy now describe a full future-native rebuild,
  contain no permanent present-product example, and derive category, candidate,
  world and lens counts from the new edition. Governance and helper suites pass;
  both skill copies are byte-identical.

## Smoothing pass

- Needed: yes, after independently judged research, scenario, taxonomy, inventory, model, skill and storefront pieces converge.
- Changes: The final synthesis replaced implementation-led taglines with
  result-led sales promises; put “what it does,” “why 2031” and “how it could
  exist” before the ranking machinery; standardised bright-16 summaries;
  preserved harms and counter-cases; removed 28 colliding or uncertain
  fictional names; and compiled the six category records into the actual
  storefront experience.
- Final fresh-critic verdict: `PASS` — after the crash-stalled local preview was
  restarted, it returned HTTP 200 and the critic confirmed the mobile and
  desktop route proof: six categories, ten rows, six worlds and zero browser
  errors.

## Final evidence

- Final data: 463 public sources, 264 public claims, six conditional 2031 worlds,
  60 actors, 74 recurring needs, six published categories, 60 ranked listings,
  12 explicit exclusions and nine geographic lenses.
- Final projection: six category indexes, 60 rich listing records, 60 accessible
  icons, nine lens files per category and a hash-bound integrity manifest.
- Release checks: final synthesis, names and projection validated; typecheck,
  production build and 30/30 active release tests passed; signed-out mobile and
  desktop route proof passed with zero browser errors; the Cloudflare dry run
  inspected the built Worker assets without publishing.
- Final independent judgement: `PASS` after the only final repair, restarting
  the local preview process left stalled by the machine crash.
- Remaining limitations: The attempted category-wide present-day overlap audit
  was stopped because its 360–410 KB packets were truncated by the live agent
  interface after candidate 01. No result from that stage was accepted. The
  final exact-name search is bounded and not legal clearance; the Top 10 score
  is a transparent authored judgement rather than a calibrated probability;
  scenarios remain conditional worlds rather than predictions. These limits
  are published in `docs/FINAL_SYNTHESIS_AND_LIMITS.md`.
