# Tokyo delivery — September 6, 2026

Tokyo is implemented on the Hamburg/Tulsa game baseline, and the research, reproduction tools and review captures are packaged. **The full-city 100,000-byte goal has not been achieved.** The current runtime uses **2,861,052 bytes of city data** before ordinary compression. No boundary, road selection or fidelity requirement was relaxed to claim a win.

The newest lossless experiments improve the measured frontier. Their smallest road and feature packets, plus the unchanged profile, total **2,534,567 bytes**. This is an offline accounting comparison, not an integrated game package: the current readers and profile pins cannot consume those replacements unchanged. Tokyo-specific source would add further bytes. The playable build therefore retains the verified, bounded runtime format.

## What to open

- `tokyo-playable-web.zip`: Tokyo, Tulsa, Hamburg and the interactive Tokyo video study. Extract and serve its `web/` directory over HTTP; `python3 -m http.server 8000 --directory web` works for local review at `http://localhost:8000`. The browser must support WebGL2. Do not open `index.html` through a `file:` URL.
- `tokyo-research-reproduction.zip`: pinned runtime source snapshot, exact city packets, frozen road input, supplemental source evidence, old and new experiments, formats, ledgers, verification receipts, plan and execution prompt. Its README gives reproduction commands.
- `tokyo-gameplay-captures.zip`: 14 original MP4s, 14 posters and 14 GIFs, with provenance and validation. Six GIFs are original; eight are clearly labeled replacements derived from preserved MP4s. They are not new gameplay captures.
- `tokyo-astra-ultra-plan.md`: the recovered complete plan, including resolution sweeps, sparse names, shared-code accounting, merge strategy and Free Space.
- `tokyo-astra-ultra-prompt.md`: a concise reusable execution prompt with the current evidence and stopping conditions.

The code is in [PR #39](https://github.com/ThatGuySam/planetwide/pull/39), revision `c504972ad0d7f0cf793fccf1d31cd58a891ad67c`, based on `b9568341d8c1a7d2349701a3c66f9e0fa82cc908`. The PR remains unmerged. This recovery did not change game source or wait for another branch to merge.

## Scope and recognition

The frozen OSM JP-13 selection contains **158,235 road ways** and **738,822 retained vertices**, including remote Tokyo islands. It is not all Greater Tokyo. Its inherited selection excludes service roads and footpaths. Repeated-node pinning repaired 105 ways and retained 96 additional vertices. Source identities, road classes, bridge/tunnel flags and layers are preserved; the evidence concerns an undirected road atlas, not legal turn-by-turn routing.

Supplemental geometry is concentrated around Tokyo Station and selected reference areas: 3,770 building/relation/part records, 491 rail/water/park constraints and 100 landmark descriptors. These are not 100 bespoke, fully surveyed or enterable landmark models. Unknown heights, facades, interiors and bridge dimensions remain inferred.

The map stores **25 street names and 25 landmark names**, permits at most **six visible names**, and supports names-off discovery. Captured named routes showed at most two names. Shape and district recognition remain important; removing names alone does not solve the dense survey geometry. Human recognition testing and acceptance of visual parity with Hamburg/Tulsa remain open.

## Bytes and reusable code

| Item | Raw bytes | Treatment |
| --- | ---: | --- |
| Runtime roads | 2,735,678 | City data |
| Runtime supplemental geography | 124,649 | City data |
| Runtime profile | 725 | City data |
| **Runtime city-data total** | **2,861,052** | **28.61 times the target** |
| Identified Tokyo content embedded in source | 5,135 | Additional city charge |
| Data plus identified embedded content | 2,866,187 | Not an exhaustive city-code attribution |
| Data plus all categorized new mixed source | 2,941,318 | Conservative for these files, not a proven global upper bound |
| New generic runtime source, five modules | 57,783 | Tracked separately; future cross-city reuse is not yet established |
| New mixed runtime source, four modules | 80,266 | Contains generic and Tokyo-specific behavior; not silently exempted |
| Existing runtime changes, seven files | +10,477 | Net source growth; not proven entirely generic |
| Existing shared interior source, eight files | 23,465 | Zero Tokyo delta in the measured files |
| Committed offline tools, 22 files | 151,502 | Outside runtime city budget |

The 5,135-byte subset is already inside the mixed-source figure; do not add it twice. The three city files total 2,464,659 bytes when individually gzipped, a transport diagnostic that does not satisfy the raw-byte contract. The recovered Web PCK is 8,969,068 bytes; the generic engine is 37,685,705 bytes raw / 9,377,158 bytes gzipped. Those are separate export/runtime measurements.

## New Free Space results

Three independent lanes tested ideas suggested by the earlier geography and byte work. All retained the reviewed quantized geometry and metadata exactly.

| Experiment | Best complete raw component | Finding |
| --- | ---: | --- |
| Shared coordinates, Hilbert order, patched bit-packed columns | 2,436,242 road bytes | 10.95% smaller; includes a 14,856-byte bbox index, but decoding remains global |
| Local prediction recipes with exact residuals | 2,613,938 road bytes | 4.45% smaller; recipe mode packed into the point-count varint |
| Exact footprint recipes plus patched columns | 97,600 feature bytes | 21.70% smaller; reconstructs the original feature packet byte for byte |

The column lane compared three packets. The recipe lane measured ten complete serialized variants; only its winner is retained, with code to reproduce the others. The feature lane retains four packets. Independent review re-decoded the recipe winner and restored the feature winner to the exact original bytes, and found no uncounted city dictionary dependency.

The combined 2,534,567-byte comparison still exceeds the goal by 2,434,567 bytes before additional embedded source. In the Hilbert candidate, original way IDs alone require 329,540 bytes, unique coordinates 1,115,529 bytes, and coordinate references 829,087 bytes. These are failures of the tested representations, not a universal impossibility proof. Moving IDs offline would require stable replacement runtime identity and a new measured representation; it cannot be assumed free.

The current byte-varint road format has a conditional floor of 2,004,193 bytes with its current way/vertex/block counts, before names. Resolution tuning within that representation cannot reach 100 KB. Different representations are still allowed; no credible measured candidate currently approaches the required reduction.

## Resolution and runtime tradeoffs already measured

- Road coordinate grids from 0.125 to 4 m were tested independently of physical world scale. Moving from 0.5 to 4 m saved only about 17.2%, while same-grade rounded-coordinate collision groups grew from 2 to 2,663. The runtime keeps 0.5 m.
- Supplemental polygons retain 0.25 m coordinates: maximum vertex movement 0.176 m and no collapsed polygons or added proper self-intersections in that sweep. Coarser variants introduced failures. Blanket rectangle replacement had P95 boundary error about 8.95 m and was rejected; the new exact selective recipes avoid that approximation.
- Spatial blocks capped at 256 ways added 26,975 bytes while measured native cold-query P95 fell from 25.351 to 14.666 ms.
- Shared interior batching reduced median draw calls from 4,908 to 1,317 on the matched 80-frame stair probe. The maximum observed pixel difference was 42 of 230,400 pixels; stair rise remained 3.160226 m.
- Seven routes were captured at full and half 3D render scale. Half scale softened the world and did not show a consistent timing benefit. These capture intervals include recording overhead and concurrent work; they are not controlled device benchmarks. The HUD stayed at full output resolution.

The planned procedural surface-detail-density sweep and human recognition study remain unrun. They are recorded as open work, not presented as completed experiments.

## Integration and verification

The runtime carries the merged city travel, original weapons/HUD, survey/orbit, display/input and shared interior behavior available at the pinned base. Hamburg presentation stays in Hamburg; Tokyo uses its own profile. Volatile codec details stay behind descriptor-producing readers. New names are data changes. Shared orchestration remains owned by one integrator. Later refactors and city updates can target those existing call sites without making them prerequisites for this delivery.

The preserved full verifier passed at the reviewed revision, including **664 Tokyo assertions**, with its original receipt and 181,439-byte log. Recovery rechecked all **327 pinned runtime source files**, rebuilt the road packet from the frozen source to the exact original SHA256, and recovered all three final city packets with matching hashes.

The rebuilt PCK has the same 308 resource paths. Compiled scripts, imported assets, geographic data and project settings match the original export; generated scene/UID resources differ after reimport. This is a newly built export from the same source, not the byte-identical old PCK. Fresh headless starts reached the expected Tokyo, Tulsa and Hamburg readiness markers. Tulsa retains an exit-time ObjectDB warning. The export contains no recovered coordination/task data.

All 14 recovered MP4s and six original GIFs fully decode at 640×360, 120 frames and 20 fps; all eight derivative GIFs also pass complete decoding. Original frame sequences and the prior all-game legacy capture archive were not recovered. The original aggregate receipts remain unchanged and the new recovery manifest identifies these limits.

Browser gameplay and physical phone/MacBook performance remain unverified; the prior browser environment lacked WebGL2. Human visual acceptance is pending. The documentation content checks passed previously, but the complete documentation build was blocked by 54 missing historical media files. None of these limits is turned into a passing claim by this recovery.

## Research and Astra/Ultra execution

The original plan records the prior-thread constraints and practitioner research, including Fabian Giesen/Farbrausch's construction recipes, TopoJSON shared arcs, MapLibre columns, topology-preserving simplification, Godot batching and the small-interface guidance of Matt Pocock, Martin Fowler/Kent Beck and Sandi Metz. The application here is measured local changes and narrow interfaces, not a new orchestration framework.

The new column experiment draws on [MapLibre's encoding specification](https://maplibre.org/maplibre-tile-spec/encodings/) and [its vector-tile paper](https://arxiv.org/abs/2508.10791). Published gains against MVT cannot be multiplied onto this already compact binary baseline. The older research lead [Compression of Digital Road Networks](https://doi.org/10.1007/978-3-540-73540-3_24) motivates shared graph structure and predicted bends; our exact residual measurements determine whether those ideas help Tokyo. [CGAL's simplification manual](https://doc.cgal.org/latest/Polyline_simplification_2/index.html) supports topology checks, with separate grade/clearance checks still required here.

Current [official Astra guidance](https://developers.openai.com/api/docs/guides/latest-model?model=gpt-6-astra) supports explicit follow-through, scoped delegation, concise evidence handoffs and proportional testing. [ChatGPT's subagent guidance](https://learn.chatgpt.com/docs/agent-configuration/subagents) describes Ultra's maximum reasoning and proactive delegation. This run used one integrator, bounded independent experiments, frozen inputs, compact artifact checkpoints and an independent result audit. These are execution choices, not a benchmark or a claim that the assistant changed the UI model setting.

The next substantive 100 KB attempt needs a representation that removes most remaining geometry/association cost while reconstructing the agreed geography. The delivered measurements make that requirement concrete. The current result is a playable baseline and a reproducible negative result for the tested 100 KB approaches.
