# Tokyo at 100 KB: Astra Ultra execution plan

Date: September 5, 2026. Status: researched implementation plan, not an implemented Tokyo build.

Tease: Spend bytes on the real city; generate its appearance.

Lede: Build Tokyo on the combined Hamburg/Tulsa runtime, and test whether 100,000 bytes can preserve the frozen Tokyo street network, a curated set of names, 100 landmark descriptors, and city-specific rules. The measured road-only baseline is 2,785,712 bytes before gzip. A roughly 28-fold reduction is an experimental target, not an established capability.

Why it matters: A small, attractive invented city would miss the geography requirement. A faithful packet that freezes the browser would miss the game requirement.

Go deeper: The contract, evidence, compression experiments, integration sequence, acceptance gates, and copyable launch prompt follow.

## Recommendation

Use a compact constraint atlas plus a shared deterministic renderer. First establish one functioning combined runtime and an honest Tokyo reference. Then compare a small set of independent encoders against that reference. Preserve topology and recognizable geometry while fitting reusable rules only where the exceptions actually cost less than literal data.

My assessment: preserving the current visual style through procedural rendering is plausible. Fitting all frozen JP-13 streets plus the other required data into 100 KB is substantially less certain. None of the researched techniques demonstrates that result. Measure before committing to a codec or reducing geographic scope.

## 1. Recovered constraints and evidence

Previous conversation retrieval was incomplete. This plan combines the visible Snow/Snow 2 and Planetwide requests, retrieved prior decisions, local source and measurements, and current GitHub evidence. Local handoffs sometimes lag or lead GitHub; those distinctions matter.

| Evidence | Observed result | Consequence |
| --- | --- | --- |
| Frozen Tokyo audit, September 5 | 158,235 retained street ways; 738,726 simplified vertices; 3,959 names; 2,785,712 raw compact bytes; 2,174,940 gzip bytes | About 27.86× raw or 21.75× gzip reduction to 100,000 bytes, before added landmarks/buildings |
| Tokyo scope | OSM administrative JP-13, including remote islands; full touching ways retained | Reproduce this scope first. Greater Tokyo and a contiguous urban subset are separate benchmarks |
| Tokyo name table | 69,339 bytes in the experimental representation | Historical all-name baseline; compare curated runtime names separately and report omitted name bytes |
| Tulsa playable-area packet | 78,219 vector bytes; 95,144 total with building envelopes and labels | This is a downtown slice, not a whole-city precedent |
| Tulsa exact encoding | 135,124 logical vector bytes become 78,219 bytes; zero additional codec error | Existing PWVECT2 recipes are a useful baseline |
| Tulsa approximation experiment | An extra 8 m simplification saved only 3.2% | Arbitrarily deleting bends is a poor first strategy |
| Hamburg | Fictional local district at a real Earth anchor | Use Hamburg as an appearance/interaction reference, not a surveyed street reference |
| Prior Earth target | 10,000,000 bytes | Track Tokyo and shared-generator growth against the wider atlas budget |

The frozen Tokyo source hash is `cb7e75c4df23ec2dec7dc9fb24d76aa37a18f828d22dbce307f056ce1e1a0833`. Its OSM base timestamp is `2026-09-05T04:48:40Z`. The experiment used 0.5 m quantization and class-dependent 0.6–1.5 m simplification with shared source nodes pinned. Its round-trip assertions do not establish production decoding or grade-aware routing correctness.

Retain the previously selected road classes and explicitly audit useful service connections. The old Tokyo query excluded service roads and footpaths at acquisition time. “All highways” in one experiment's filename does not mean all Tokyo highway features. Do not broaden or shrink that definition invisibly.

The frozen experiment contains roads, not a complete Tokyo art/landmark atlas. During baseline preparation, acquire and freeze compatible source records for the 100 selected landmarks and required rail, water, park and building constraints. Record query/boundary policy, acquisition date, attribution, content hashes and measured section costs. Keep these additions separate in comparisons with the old road-only result. Select landmarks across the declared scope and review their source positions, forms and heights; missing heights remain explicit inferences.

### Budget contract

The primary gate is **100,000 decimal bytes of city-specific encoded data before ordinary gzip/Brotli/Zstd compression**. Report transport bytes separately. This follows the prior request for domain compression rather than a transport-only win.

Count geometry, the selected runtime names, landmark descriptors, indices, headers, checksums, fitted parameters, dictionaries, residuals, city-specific shader constants or code, and every shipped geographic LOD. Count Tokyo-specific content even if embedded in a GDScript file, generated source, or a supposedly shared lookup table. Count all required shards, not just the first fetched one.

Report shared engine/generator bytes and their incremental growth separately. A generic operator reused across cities belongs there; hard-coded Tokyo coordinates do not. Also report decoded RAM, generated GPU resources, and generation time. Generated textures can have zero download bytes and substantial memory/runtime cost.

Keep source snapshots and rebuild provenance outside the runtime budget when they are not required at play time. Preserve their licenses and availability. Source IDs may move to an offline provenance table only after confirming that runtime identity does not depend on them. Do not confuse this with removing provenance.

Use this provisional allocation to diagnose pressure, not to justify cutting data:

| Component | Target bytes |
| --- | ---: |
| Road topology, geometry and residuals | 55,000 |
| Water, rail, parks and other geographic constraints | 8,000 |
| Names and 100 landmark descriptors | 15,000 |
| Building/district recipes and exceptional envelopes | 12,000 |
| Headers, directory, integrity, versioning | 3,000 |
| Unallocated margin | 7,000 |
| Total | 100,000 |

These are allocations, not predicted sizes. The target permits only about 0.632 bytes per retained source way if roads consumed the entire package. That arithmetic illustrates the challenge; it is not an information-theoretic impossibility proof.

## 2. Start from merged main and keep later integration local

GitHub discovery on September 5 showed these integration commits:

- `f18614c2c93a1e794cde7b7fdff6efcdda9af049`: Hamburg in the production world; aligned glazed openings, lifts, exterior placement and touch fixes.
- `5c6d5027c11cad87f65ef53fa05139fd3cfb62a4`: Tulsa geometry codec, original weapons/audio, phone work, geographic survey scale, surface/cloud/skyline and performance changes.
- `e1b942da64ccaeae56b1c9bdc64f753a6ed4c7bd`: desktop controls and CRT modes.

These are verified discovery anchors, not a promise that main will remain there. Refresh live main and check ancestry at execution time. Existing materialized worktrees can have synthetic or partial history; their HEAD is not authoritative upstream history.

| Layer | Bring forward | State at this planning pass |
| --- | --- | --- |
| Shared world | Absolute position, floating origin, map travel, survey/orbit, stable chunk identity | In observed integration history |
| Tulsa | Exact packet decoder, intersection/sidewalk repairs, shared visible/collision footprints, original 2D weapons/HUD, rain, audio warmup, clouds and skyline | In observed integration history |
| Hamburg | Open floors, matched windows/glass, lifts/ladders, swimming, world-space snow, sheltered interiors, footsteps/ambience | In observed integration history; retain documented limitations |
| Display/input | Sharp/VGA/RGB-TV/composite approximations; aspect/resolution choices; keyboard, pointer lock, trackpad drag, touch | CRT/desktop commit observed; physical-device acceptance still open |
| New Tulsa loading/input | Separate engine download and scene progress, validated bytes, stall timeout, nonfatal pointer-lock rejection, combat bridge, disconnect travel callbacks | Present in local handoff and source; do not assume landed until live verification |
| Snow 2 additions | 60 room traits, deterministic room history, procedural graphics, elevator panels/buttons, varied floor connections, requested screenshot defaults | User requirements recovered; completed implementation and exact defaults not verified |

Do not copy Hamburg wholesale into Tokyo or attach Tulsa weather globally. Share mechanisms and keep city presentation choices in city profiles. Treat the CSS/div comparison as a request for computed graphics; Godot should still use generated meshes, materials and shaders rather than thousands of DOM elements.

### Merge procedure

1. Read live `AGENTS.md`, scoped instructions and `docs/agent-workflow.md`. Inventory main, relevant active branches, worktrees and handoffs. Record source revision, owner, changed paths and verification for each required layer.
2. Start one Tokyo branch from current main and pin its SHA. Proceed on merged behavior. A stable unmerged change may be reused selectively when its exact revision, narrow surface, tests and lack of volatile dependencies are established; record that judgment and keep it independently revertible. Never make an unfinished branch a prerequisite or blindly cherry-pick already-landed changes.
3. Run the existing cross-city travel, input and rendering gates against that available baseline. Capture what is actually present. Do not wait for pending Snow 2, Tulsa or technical-debt refactor merges to begin Tokyo. A clean Git merge alone does not prove behavioral compatibility.
4. Let one integrator own `src/game.gd`, `src/world_stream.gd`, `tests/run_tests.gd`, shell/input wiring and exports. Experimental workers use separate paths/worktrees. No simultaneous shared-file writers.
5. Add Tokyo behind the existing city-selection/profile mechanism. Introduce only the extension points needed by a second real consumer. Keep PWVECT1/PWVECT2 decoding compatible; version any full-city format separately.
6. Refresh main at deliberate integration checkpoints and before final validation. Incorporate relevant landed changes, rerun affected gates and the required full verifier, then prepare one coherent review branch/PR. Record the final tested base/head. Pending work remains a follow-up integration item, not a reason to postpone the Tokyo deliverable. Preserve existing approval boundaries for merge and publication.

Sam clarified that waiting for other merges is unnecessary. Build a working Tokyo against current main and stable existing behavior. Preserve the goal of reusing later Snow 2 and Tulsa improvements through localized integration, but do not require those additions for this delivery. Describe exactly which revisions and capabilities are included, and retain a short follow-up map for the remainder. Never claim pending work has landed.

### Engineering approach: make the next change local

Research checked September 5, 2026. These are design principles applied to Planetwide, not instructions to install a new framework or execute every referenced skill.

- **Matt Pocock's codebase-design guidance:** put substantial behavior behind a small interface; keep implementation choices internal; introduce an adapter boundary when there is real variation. Application: codec experiments should produce the same city descriptors so the renderer does not learn each encoding. Keep private parsing, integer packing and validation behind that entry point. Do not create interfaces for every helper. [S13]
- **Martin Fowler and Kent Beck:** make a small behavior-preserving preparatory refactor when it makes the immediate feature easier. Application: extract the specific city-selection or descriptor-consumption call needed for Tokyo, verify existing behavior, then add Tokyo in a separate change. Avoid bundling a general main-branch cleanup with the feature. [S14]
- **Sandi Metz:** forced reuse can become an abstraction full of special-case parameters. Application: tolerate small experimental duplication until the common behavior is clear. Consolidate a proven shared rule without making every city pass dozens of flags through one renderer. [S15]
- **PStack's published workflow:** keep implementation, parity checking and independent verification explicit. Application: use a reviewable decision trail and compare behavior at actual revisions; do not adopt its entire orchestration stack. This source is workflow evidence, not proof that a particular architecture is optimal. [S16]

The synthesis is a few cohesive modules with explicit inputs, stable identities and local changes. Flexibility comes from hiding volatile details and keeping outputs testable, not from maximizing the number of abstraction layers.

| Likely change | Where it should stay | Evidence that integration is local |
| --- | --- | --- |
| New geometry encoding or resolution policy | Codec/packing implementation | Same canonical descriptor contract; renderer unchanged |
| Different Tokyo names or landmark priorities | City data/profile | Shared selection/rendering logic unchanged |
| Updated room traits or furnishings | Current interior-generation entry point and profile data | Street geometry and building identity unchanged |
| New stair/lift strategy | Circulation generation and its traversal behavior | No city-loader or compression rewrite |
| Updated engine organization from the refactor thread | Production wiring or a necessary small adapter | Tokyo packet, deterministic generation and benchmark fixtures remain usable |
| Display, weather or input fix | Existing shared presentation/input components | Codec and source-data pipeline unchanged |

Use current project conventions and typed GDScript data structures where they make contracts clearer. Keep pure descriptor computation separate from node creation and browser/Godot lifecycle effects. Document units, ownership, ordering and version rules at the few meaningful boundaries. Version persistent byte formats and deterministic identities; avoid versioning every internal function or guessing future APIs.

Prefer small additive files for new Tokyo/codec experiments, with minimal edits to shared orchestration. Keep mechanical extraction, behavior changes and generated-data updates distinguishable in the diff. Avoid wholesale file moves or formatting churn across paths another thread owns. Test through meaningful behavior contracts rather than private helper layout, so later refactors do not require rewriting all tests.

Maintain a short integration note: pinned base SHA, stable inputs and why they were chosen, touched shared paths, included capabilities, pending improvements, and the call sites expected to absorb them. This is enough coordination; do not build another task scheduler or claim system. The separate technical-debt effort can proceed independently. When its changes land, adapt to its actual interfaces rather than implement its anticipated architecture in advance.

At a planned checkpoint, rehearse replacing one real experimental codec or profile with another already-built variant. Measure changed call sites and rerun the contract checks. If that replacement requires rewriting unrelated rendering or controls, fix the demonstrated coupling. Do not invent a second implementation solely to justify an adapter.

## 3. Research: what transfers and what does not

### Small executable demos: store the construction recipe

Fabian Giesen's first-party account of Farbrausch's tools and the original source repository describe operator-based procedural mesh and texture construction. The transferable idea is to reuse a small vocabulary of construction operations with compact parameters. It does not prove that a measured real-world street network will fit the same budget as an authored demo. [S1, S2]

Apply this to facade bays, roof profiles, stairs, railings, pipes, window assemblies, curb materials and original signage. Reuse computed motif definitions at several scales. Author a small generic material vocabulary, then derive variation from stable IDs. Avoid recreating a general-purpose node editor or importing proprietary game assets.

### Shared topology and columnar encoding

TopoJSON represents shared geometry through arcs and can quantize and delta-encode coordinates. MapLibre's tile specification separates geometry and properties into streams and describes delta, dictionary, Morton and run-length encodings. Borrow the mechanisms in a bounded game format. A full generic GIS schema may cost more than it saves here. [S3, S4]

Roads meet more often than they share long identical arcs. Junction sharing and compatible degree-two chain merging may therefore matter more than generic polygon-style arc deduplication. This is a hypothesis to measure on Tokyo.

### Fit a grammar, then pay for its mistakes

Wu, Yan, Dong, Zhang and Wonka's 2014 facade-layout paper investigates recovering split grammars from observed layouts. It supports trying inverse procedural fitting for repeated facades. Applying the same idea to Tokyo streets is a separate experiment, not a published result from that paper. The paper's abstract was retrievable; full-text retrieval failed during this pass. [S5]

For each region compare literal encoding against `recipe + dictionary references + residuals + exceptions`. Select the smaller valid representation. An attractive generated road is not a valid replacement for a mismatched source road. Use held-out districts and other cities to detect dictionaries that merely hide Tokyo data in the engine.

### Topology protection beats appearance-only simplification

CGAL documents simplification that preserves intersections and polygon nesting. Pinning endpoints alone is weaker: a shortcut can create a crossing between previously separate streets. Tokyo additionally requires bridge/tunnel/layer semantics; a 2D intersection is not automatically a junction. Use these as offline validation principles. Library adoption would need its own dependency/license review. [S6]

### Runtime geometry has a different optimization problem

Meshoptimizer documents quantization, spatial ordering and mesh compression. Those techniques are useful references, but shipping compressed render meshes would undermine the current procedural-data architecture. Consider its ordering ideas first; add a dependency only for a measured need. [S7]

Godot's MultiMesh guidance supports batching repeated instances, with all-or-none visibility per MultiMesh. Use spatially bounded batches, not one whole-Tokyo batch. The 4.6 page carries an update warning, so confirm behavior in the installed 4.6.1 runtime. Web export remains WebGL2 Compatibility; its documentation favors single-threaded export for compatibility. [S8, S9]

### Evidence synthesis

The durable pattern is to remove repeated structure before approximating unique structure. The tradeoff is startup work and decoder complexity. Geography codecs preserve data but may not achieve the requested ratio; generative art can be very compact but cannot independently reproduce arbitrary geography. No source supports a blanket 28× claim. Do not spend the budget on a new neural decoder, runtime seed search, or a wholesale engine rewrite before simpler experiments have measured results.

## 4. Compression experiment order

Use identical frozen inputs and a canonical decoded representation. Save each encoder version, parameters, output hash, per-section bytes, runtime decode time and fidelity results. Report an ablation, not just the smallest final file.

| Experiment | Mechanism | Required falsification check |
| --- | --- | --- |
| E0: honest baseline | Reproduce Tokyo audit; normalize production-compatible descriptors | Counts/hash/scope agree; source and codec error are separated |
| E1: topology first | Share junctions; merge compatible degree-two chains; encode branch structure and ordered bends separately | No missing edges, dead ends, property boundaries or grade distinctions |
| E2: better integer streams | Local origins, signed deltas, delta-of-deltas, fixed bit widths versus varints, Morton ordering, RLE/default flags | Whole-file savings after references/directories; independent decoder agrees |
| E3: labels | UTF-8 dictionary, prefix sharing/front coding, repeated suffixes, compact references | All retained strings round-trip exactly, including Japanese text |
| E4: reversible geometric recipes | Repeated steps, local orientation/spacing motifs, canonical junction templates with literal residuals | Exact decoded baseline unless explicitly classified as lossy |
| E5: fitted district rules | Fit genuine regularities; preserve irregular residuals; choose per-region literal fallback | Coverage/topology/error gates still pass; all fitted bits counted |
| E6: conservative approximation | Topology-aware vertex reduction within predeclared tolerance | No new crossings; unchanged semantic graph; geometric and route limits pass |

Do E1–E3 before expensive fitting. Benchmark representatively across dense irregular streets, major roads, waterfront, sparse outskirts and remote islands, then confirm the whole JP-13 packet. Do not select a flattering downtown sample as the final benchmark.

Creative candidates worth a bounded experiment:

- **Junction graph plus bend streams:** preserve the graph explicitly, then encode shape between decision points. Road classes and names can attach to runs. This targets fragmented source records without deleting roads.
- **Reference-line residuals:** describe a chain relative to its dominant direction or a nearby parallel chain. Pay for offsets and exceptions; do not infer a missing road merely because a parallel one exists.
- **Minimum-description-length tile choice:** each spatial block selects literal, chain, repeated-step or fitted representation based on complete encoded cost. Include its mode tag and seam links. Stop adding modes when marginal savings do not justify the decoder.
- **Procedural history shared across buildings:** generate wear, repairs, lighting and furnishing from purpose, age and stable building/floor IDs. A shared event can influence several rooms without storing every object. Preserve traversable space before adding clutter.
- **Landmarks as residuals over generic forms:** store footprint, height, roof/form parameters and only the silhouette details a generic grammar cannot express. Keep 100 descriptors distinct from 100 labels or 100 fully bespoke models.
- **Compute survey LOD from the same atlas:** derive distant geometry at load time, share cached descriptors, and keep relevant ground chunks resident. If any extra LOD is shipped, include its bytes.

A seed only chooses among outputs of the frozen generator. Extra hash-derived child seeds do not add descriptive capacity. Store unusual geometry explicitly when the generator cannot reproduce it economically.

### Updated name selection and discovery contract

Sam clarified after the initial plan that Tulsa survey labels felt too dense, and recognizable shapes and areas often supplied enough orientation. Preserve room for players to discover and identify places themselves. This supersedes any implication that every source street name must ship or appear on the map.

Separate three decisions:

- **Geometry:** retain the real street network and meaningful landmark shapes. An unnamed street still exists, renders and connects normally. This clarification authorizes name selection, not removal of lesser-known streets.
- **Stored names:** ship a curated set of recognizable or navigationally useful street and landmark names. Keep the full source names in offline provenance. Retain the 100 landmark descriptors as the existing geometry/content target; each does not require a stored name or visible label.
- **Displayed labels:** show fewer names than we store. Use zoom, available screen space, collision avoidance and stable priority. Regional views emphasize a few major anchors; closer views may expose selected street names. Avoid labels flickering or reshuffling during small camera movements. Ordinary locations may remain unnamed.

Proposed experiment, not a fixed content quota: compare 25, 50 and 100 named anchors total across streets and landmarks. Select for recognizable forms, route decisions, geographic spread and distinctiveness, rather than filling quotas or reproducing a tourism ranking. Compare a label-free view, sparse default and optional richer view on matched Tokyo and Tulsa routes. Evaluate whether the player can orient and find a destination without the labels obscuring intersections or silhouettes. Discovery-triggered labels remain optional; do not add a progression subsystem just to reduce clutter.

Report retained name count, exact encoded name/reference bytes, omitted source-name count and geometry coverage separately. Compare both all-name and curated-name packets against the original benchmark so editorial savings are not presented as codec improvements. Do not claim the 69,339-byte historical name table disappears in full; selected strings and references still cost bytes. Shared selection/rendering logic belongs in the shared-code ledger, while Tokyo's selected strings, positions and priority overrides count toward 100 KB.

### Resolution experiments: measure the actual tradeoffs

Sam explicitly requested experiments across unit and data resolutions. Run the following sweep before selecting a shipping precision. Coarse candidates are diagnostic experiments, including where they fail existing fidelity gates; this does not authorize silently weakening the production contract.

Keep authoritative physical scale in metres. Encoding an integer as quarter-metres or two metres is a storage choice, not a change to player size, movement speed, gravity, world scale or collision margins. Do not globally rescale Godot units as a compression shortcut. Report coordinate step, integer width and origin separately. Coarser values save bytes only if the encoded distribution, bit width or representation actually becomes smaller.

| Dimension | Proposed experimental values | What it tests |
| --- | --- | --- |
| Horizontal coordinate grid | 0.125, 0.25, 0.5 baseline, 1, 2, 4 m | Byte cost versus bends, narrow streets, landmark placement and seams |
| Polyline simplification tolerance | 0, 0.5, 1, 2, 4, 8 m, plus existing class-dependent baseline | Benefit of removing vertices independently of coordinate rounding |
| Ordinary building dimensions and height | 0.25, 0.5, 1, 2 m | Envelope size, skyline steps, clearance and collisions; distinguish inferred from sourced values |
| Ordinary building orientation | Existing representation versus 0.5, 1, 2, 5 degree steps | Savings versus rotation and corner displacement, especially on long buildings |
| Interior/opening local grid | Existing precision versus 0.01, 0.025, 0.05, 0.1 m | Door/window seams, stairs, lift alignment and traversal; shared-generator cost rather than automatically Tokyo payload |
| Codec spatial block edge | 96, 384, 1,536, 6,144 m | Origin/directory overhead, integer range, decode locality and seams; rendering chunks remain independently sized |
| Procedural surface detail | 0.5×, 1×, 2× baseline sampling density | Runtime texture/cache cost, shimmer and visual detail, usually without geography-byte savings |
| Render pixel grid | 320×180 and 640×360 at matched 16:9; TV/aspect modes tested separately | Whether map errors are visible at target resolution; avoid changing aspect and data precision together |

Values are experimental starting points, not accuracy claims about source data. Finer quantization does not create new information. Regenerate each candidate directly from the same frozen source rather than requantizing the half-metre packet or chaining lossy conversions. Keep an unsimplified reference, the existing production-style baseline, and each candidate distinguishable. Fix rounding, tie-breaking, projection and canonical IDs. Keep Japanese strings exact; resolution changes apply to numeric/spatial fields, not text fidelity.

Use a staged experiment rather than the full Cartesian product:

1. Freeze the encoder, source, selected names, weather, capture seed and reference routes. Sweep coordinate grids with simplification held constant, then simplification with the grid held constant. Run ordinary-building and local-interior sweeps separately. Preserve landmark precision initially.
2. Record the best tradeoff candidates and failure thresholds for each dimension. Combine only promising settings to measure interactions; retain one coarse stress case to expose failure modes. Compare no-simplification and class-dependent baseline controls.
3. Test mixed precision: fine junctions, openings and landmark geometry; coarser safe segments and ordinary envelopes. Account for precision tags, reference tables and exceptions. Protect close approaches and grade separation as well as shared endpoints. Decode shared boundaries from one canonical record to avoid cracks.
4. Use stratified representative regions for discovery, including narrow alleys, long rotated buildings, parallel streets, bridge/tunnel crossings, coastline, islands and chunk seams. Validate finalists on the entire frozen JP-13 dataset and representative Hamburg/Tulsa regressions.
5. Inspect paired, fixed-camera images and short traversal clips in Sharp and the selected CRT mode. Add a zoomable geographic error overlay for diagnosis. A prettier filter cannot override a failed topology or collision gate.

For every candidate record raw city bytes by section, transport bytes, shared-code growth, point/feature counts, coordinate bounds, geometric error distribution and maximum, new/lost graph connections, route deviations, footprint overlap, minimum traversable clearance, landmark silhouette/position errors, decode latency, generated RAM/GPU cost and browser frame times. Use repeated matched runs for latency/frame results and report variation; repeat only enough to resolve candidate differences. Separate name-selection savings, codec savings, precision loss and rendering changes.

Deliver a machine-readable result table and a size-versus-error plot with pass/fail markers, plus worst-case annotated captures. Show the non-dominated choices: candidates for which no alternative is both smaller and more faithful at comparable runtime cost. Identify the point where additional byte savings begin causing meaningful gameplay or visual damage. Report actual measurements rather than estimating all resolutions from one sample.

The intended outcome is a justified choice of precision per feature type, with a small number of format modes. A universal coarse grid or unlimited per-feature overrides should not win by default. If a small precision exception preserves an alley or recognizable silhouette cheaply, measure it. If mixed-precision metadata consumes the savings, prefer the simpler uniform candidate. Keep candidates outside the existing added-error limits in the report as failed diagnostics; any relaxed shipping threshold is a separate decision for Sam.

### Full-city coordinate prerequisite

PWVECT2 currently expands into signed i16 coordinates at half-metre resolution. That gives only about 32.768 km of span per axis and cannot represent the complete Tokyo scope around one origin. Use a versioned directory of local coordinate blocks with wider authoritative origins and explicit seam ownership. Preserve absolute-world identity through rebasing. Validate projection error for the widely separated islands rather than reusing one urban tangent-plane approximation unquestioned.

Choose codec block size by measuring directory overhead against decode locality. It need not equal the existing 96 m rendering chunk. Do not create a header per rendering chunk unless the savings survive that overhead.

## 5. Visual and gameplay parity

Render source-preserving streets with the shared road/sidewalk/intersection rules. Ordinary buildings may remain deterministic approximations, consistent with Tulsa; disclose that distinction. Preserve mapped landmark positions and available authoritative shapes/heights. Never present inferred architecture as surveyed geometry.

Use a small, original Tokyo material/form vocabulary, with candidates such as narrow facade bays, mixed storefront/residential stacks, service piping, tiled surfaces, shutters and rooftop equipment. These are art hypotheses for reference review, not claims that every Tokyo district looks alike. Research the actual selected neighborhoods and landmarks before fixing their profiles.

Generate texture detail once into reusable runtime resources or evaluate cheap shaders as appropriate. Do not regenerate every material per frame. Bound texture-cache size and eviction. Japanese labels need an explicit glyph strategy: use an existing shared font where available, measure any new subset as incremental content, and do not promise exact readable Japanese lettering from decorative procedural strokes.

Use the current merged interior behavior now; integrate the Snow 2 shared 60-trait catalog when available through the smallest necessary call boundary. The full catalog is not a prerequisite for a playable Tokyo. Resolve entrances, circulation cores and outside openings before rooms and furnishings. Derive independent random streams from version/city/building/floor/room/stage identity. Adding a chair must not move windows or regenerate the building next door. Retain older city outputs through versioned behavior or explicit migration.

Floor connections should depend on building purpose, era, dimensions and playable clearance. Use supported current traversal choices and incorporate stable researched additions when available, with literal exceptions where needed. Keep later circulation changes local; do not build a competing universal traversal framework. Elevators need aligned stops, interior surfaces and usable panels. All traversal types need a connected walkability proof and actual player traversal.

Keep weather, time, city ambience and preferred art palette local to profiles. Preserve global user display/input preferences. Exact screenshot-selected defaults were not recoverable here: use the latest verified upstream configuration and flag that single unresolved comparison when the screenshot or committed values become available.

## 6. Acceptance gates

| Gate | Evidence required |
| --- | --- |
| Scope and bytes | Frozen source hash and query; retained feature coverage; complete city-specific payload manifest; raw/transport totals; no runtime map fetch after preload |
| Exact codec | Independently implemented offline and Godot decoders agree on canonical values; useful malformed-input and version tests; declared decode/allocation limits |
| Geography | All baseline road coverage accounted for; no lost junctions or new at-grade connections; layers/bridges/tunnels preserved; polygon holes and landmark coordinates valid |
| Approximation | Measure source-to-baseline and baseline-to-candidate separately; initial proposed added-error ceiling of 2 m and route-length deviation of 1%, tightened where clearances require it; publish worst cases and sampling limitations |
| Routes | Full semantic graph comparison after compatible chain normalization; deterministic route sample across districts, short alleys, bridges and tile seams; no reachability change |
| Determinism | Identical canonical descriptors across offline/Godot implementations, chunk visitation order, reloads, travel and origin rebases; visual shader differences need not be bit-identical |
| Visual parity | Fixed-camera reference routes in all three cities, same resolution/FOV/time/weather/display mode where applicable; inspect Sharp as well as selected CRT mode so softness cannot hide missing geometry |
| Interaction | Walking/collision agree with visible surfaces; windows line up; floors reachable; no furnishing blocks required access; travel restores city effects and user input state |
| Browser performance | Record hardware, browser, resolution and mode; median/p95/p99 frame time, stalls, RAM/GPU allocation, decode time and cold time-to-play; bounded chunk generation |
| Regression | Required `./scripts/verify.sh`, existing Tulsa/winter/input/loader suites, focused Tokyo gates and actual exported-game captures at the final revision |

The 2 m and 1% figures above are proposed experimental ceilings, not prior user-approved tolerances or measured Tokyo accuracy. Exact codec candidates add zero error. If these ceilings would violate the original baseline or observed visual parity, retain the stricter baseline rather than relax it to reach 100 KB.

Use provisional performance targets of at least 30 FPS on the physical phone route, preserve any stricter existing gate, and keep the prior 60 FPS aim visible. Allow no more than a 10% p95 frame-time regression on matched Hamburg/Tulsa routes. These are planning targets, not hardware measurements. A software-rendered capture does not certify iPhone Safari or MacBook performance.

The final review route should include a landmark street, narrow ordinary street, elevated/grade-separated segment, waterfront, furnished interior with multiple floors, and street-to-survey-to-orbit travel. Produce real gameplay GIFs at the project's accepted capture rate and provenance requirements. Headless tests cannot replace visual acceptance.

## 7. Astra Ultra work structure

OpenAI's current Astra guidance recommends explicit initiative/follow-through, clear instruction priority, specified delegation behavior, concise reporting and testing proportionate to actual risks. Its model guidance describes Ultra as parallel subagent work, while Max allocates more reasoning to a single task. Select Astra and Ultra in the product controls; wording in a prompt does not itself configure the session. [S10, S11]

Application here: use one integrator and up to four independent lanes within the available capacity. Freeze their inputs and contracts first. This is our project design, not a claim that OpenAI prescribes a particular agent count.

| Lane | Output and ownership |
| --- | --- |
| Integrator | Live revision inventory, shared runtime changes, sequential merges, final regression and handoff |
| Codec investigator | E1/E2 candidates in isolated experiment paths; section sizes and decode results |
| Labels/recipe investigator | E3/E4/E5 candidates; dictionary/exception accounting; no runtime shared-file edits |
| Fidelity reviewer | Independent oracle and counterexamples; tests source loss, topology and budget claims |
| Rendering reviewer | Inspect baseline captures, identify reusable city/interior hooks, propose measured runtime work; later write only explicitly assigned independent files |

Keep a short file-based checkpoint at each phase: base/head, source hashes, decisions, measured results, failures, next action. Matt Pocock's skill repository provides first-party examples of handoffs and dividing large work into decision tickets; it is workflow evidence, not Astra-specific performance evidence. [S12]

### Phases and stop conditions

1. **Reconcile and baseline.** Pin current main, record available capabilities and stable optional inputs, and establish the working baseline, frozen Tokyo reference and measurement contract. Begin implementation without waiting for independently owned art, input or refactor merges.
2. **Prove the best representation.** Run E0–E3 and the staged unit/data-resolution sweep, then promote promising recipe experiments. Compare 100/200/500 KB targets and the best fidelity-preserving result. Record inability to meet a budget; never fill a graph with misleading zeros.
3. **Playable Tokyo.** Decode into the existing city pipeline, add city profile and landmarks, reuse shared procedural/interior features. The whole-city atlas can exist while only nearby chunks are materialized.
4. **Close visual and runtime gaps.** Capture and inspect the review route. Correct missing geometry, repetitive art, blocking clutter and load/frame spikes. Benchmark actual browser export.
5. **Free Space.** Use the geography, byte-accounting, resolution and gameplay results to invent and run new experiments. Preserve the best verified candidate first. Follow the evidence beyond the original experiment list, then promote measured improvements.
6. **Integrate final upstream changes and review.** Reconcile once more, run required checks, save reproducible evidence, update project docs and present the PR/build with exact bytes and remaining acceptance items.

If no tested candidate meets both fidelity and 100 KB, retain a faithful playable build and the measured size/fidelity frontier. Present the smallest passing size and the concrete losses required by smaller candidates. A separately labeled reduced-area demo is an option only after Sam chooses it; it does not satisfy the full JP-13 goal. Lack of success in these experiments is not proof of mathematical impossibility.

### Free Space: experiments suggested by what we learned

This is an intentional exploration period near the end, after the core measurements and playable candidate exist and before final integration and verification. Run it even if 100 KB has already been achieved: better recognizability, simpler shared code, smoother streaming or stronger cross-city reuse are valuable outcomes too. If the target remains unmet, the best measured candidate is the starting point.

Do not preselect every experiment now. First review surprising patterns, expensive residuals, failed precision settings, visual recognition cues and runtime bottlenecks. Propose a short slate of new hypotheses tied to those observations. Include an unconventional idea when there is a concrete, cheap way to test it. The original E0–E6 list is a starting point, not an exhaustive boundary.

For each experiment write a brief record: observation, hypothesis, smallest useful prototype, comparison baseline, measured result and disposition. Useful learning can be a demonstrated failure or a simpler explanation, not only a shipping optimization. A new naming strategy, geometry representation, procedural motif, precision policy or streaming approach is eligible when it follows from the results.

Default scope: one round of three to five small experiments, followed by one refinement round for the strongest result. This is a planning default, not an invented wall-clock or spending allowance. Fit the work to actual remaining session capacity and existing resource limits, retaining time for final verification. Stop when another round would mainly repeat findings; record worthwhile unfinished ideas for later.

Keep the best verified build and its packet immutable as the comparison baseline. Use isolated experiment paths or branches and the same inputs wherever the hypothesis permits. Independent experiments may use the existing Ultra lanes; the integrator remains the only writer to shared runtime files. Record any deliberately changed source, name set or visual setting so comparisons remain interpretable.

Experiments may probe coarser or otherwise failing candidates to learn where the limits are. Promote results only when they improve the measured tradeoff, pass the applicable geography/gameplay gates and retain honest shared-versus-city byte accounting. Test a claimed reusable improvement on Hamburg or Tulsa as well as Tokyo where applicable. If nothing improves the baseline, preserve the learning and ship the baseline through the normal review process.

Deliver a compact experiment gallery or table with before/after bytes, fidelity/runtime effects, captured evidence where relevant, and a clear outcome: adopt, reject with reason, or revisit with a specific next test. Carry adopted results through phase 6 and the final regression suite; exploratory outputs do not bypass acceptance.

## 8. Entry points for the implementing session

Resolve these paths in the live repository, then follow current symbols/call sites. Some newer paths were found only in the local Tulsa source and may require integration.

- Contract: `AGENTS.md`, `docs/agent-workflow.md`, `docs/site/src/content/docs/features/tulsa-phone-preview.md`, `docs/site/src/content/docs/features/winter-district.md`.
- Existing codec: `src/city_geometry_codec.gd`, `src/city_vector_target.gd`, `src/city_target.gd`, `scripts/pack_tulsa_geometry.py`, `scripts/test_city_geometry_codec.py`, `scripts/measure_tulsa_source_fidelity.py`.
- Existing source/evidence: `docs/research/tulsa-compression-2026-09-05.md`, `docs/research/city-vector-atlas/PWVECT2-format.md`, `docs/research/city-vector-atlas/`.
- Prior design: `docs/site/src/content/docs/research/geohash-seed-fitting-and-procedural-interiors-2026-07-19.md`, `docs/site/source-data/city-seed-benchmark/`.
- Rendering and integration: `src/city_surface_rules.gd`, `src/world_stream.gd`, `src/game.gd`, `src/tulsa/city_distant_view.gd`, `src/winter/winter_layout.gd`, `src/winter/winter_interiors.gd`, `src/display/`.
- Web regressions: `web/desktop-controls.mjs`, `web/engine-loader.mjs`, `web/shell.html`, `src/web_game_bridge.gd`, `src/web_desktop_bridge.gd`, `tests/desktop_controls.test.mjs`, `tests/engine_loader.test.mjs`, `tests/tulsa_shell.test.mjs`.

At planning time the full Tokyo experiment was available under `/workspace/scratch/56eb2a400cdc/planetwide-data-audit/`, including `README.md`, `benchmark.py`, `tokyo-full-osm.json.gz`, `tokyo-full-osm-measurement.json` and the compact outputs. This is a recovery hint, not a durable path guarantee. Locate its retained artifact or exact hash before reacquiring data; a newer download is a new benchmark version.

## 9. Copyable launch prompt

```text
Implement the attached Tokyo at 100 KB plan in ThatGuySam/planetwide.
I am selecting Astra with Ultra in the session controls.

Goal: a playable Tokyo using the combined Hamburg/Tulsa runtime and its
visual fidelity, targeting at most 100,000 bytes of city-specific encoded
data before general-purpose compression. Include the frozen JP-13 road
scope, curated recognizable street/landmark names, 100 landmark descriptors,
city rules, dictionaries, indices
and exceptions. Report engine/generator growth, transport bytes, decoded
memory and runtime performance separately.

Start by reading the attached plan, live AGENTS.md and relevant handoffs.
Resolve and pin live main and the available Hamburg/Tulsa source revisions.
Start from merged behavior; do not wait for Snow 2, Tulsa fixes or the
separate technical-debt refactor to merge. Reuse stable unmerged work only
selectively at a verified immutable revision with narrow, revertible changes.
Keep codec, city descriptors and rendering concerns separate through small
existing call boundaries. Add adapters only for real incompatible interfaces;
avoid speculative frameworks. Integrate newly landed improvements at planned
checkpoints and record pending capabilities as follow-up work, not blockers.
Verify which defaults and interior/input features are actually included.

Use one shared-runtime integrator. Delegate independent codec experiments,
label/recipe work, fidelity checking and rendering investigation to bounded
subagents on frozen inputs and disjoint paths. No concurrent writers to
game.gd, world_stream.gd, tests/run_tests.gd or shared shell/input wiring.
Honor the repository's current coordination rules without redesigning them.

Reproduce the source baseline, then test topology/chain encoding, integer
streams and labels before expensive fitting. Sweep coordinate precision,
simplification, dimensions/heights, orientation and codec block size separately,
then test promising combinations and mixed precision. Keep physical world
scale fixed. Measure bytes, geographic error, connectivity, clearances, visual
changes and runtime cost. Preserve coarse failures as diagnostic evidence,
not shipping defaults. Count all parameters and
exceptions. Preserve real streets, intersections, layers, landmark positions,
source coverage and deterministic identities. Keep generated architecture
clearly distinguished from surveyed geography. Version the coordinate/codec
extension needed for full-city scope while retaining old decoders.

Proceed autonomously through implementation and ordinary repairs. Keep
compact file checkpoints with revisions, measurements and the next action.
Ask only for a consequential decision the existing contract cannot resolve.
Do not redefine the city boundary, remove minor streets or hide Tokyo data
in shared code to claim success. If 100 KB cannot be reached at passing
fidelity, finish the faithful playable candidate and show the measured
size/fidelity frontier and exact unresolved choice.

Verify independently decoded data, geography/topology, cross-city travel,
collision/interiors, desktop/trackpad/touch, loader behavior, and performance.
Run required repository gates at the final integrated revision. Capture and
inspect actual game output in Sharp and the selected CRT mode. Report
physical-device checks separately from runner tests.

Before final integration, run the plan's Free Space phase. Review what the
geography, precision, compression and gameplay experiments taught us, then
invent and test a small set of new ideas beyond the initial list. Preserve
the best verified baseline, give independent experiments separate outputs,
and compare measured results. Keep useful negative findings. Promote only
improvements that pass the existing contract, then perform final verification.

Deliver the coherent review branch/PR, reproducible packet and byte ledger,
source/codec hashes, format/rebuild instructions, benchmark results, real
game captures and a concise handoff. Preserve current merge/publication
approval boundaries. Do not stop after another plan or an attractive mockup.
```

## Sources and provenance

All external sources checked September 5, 2026. Primary technical sources were preferred. Search-only snippets were not used to establish implementation claims. No retrieved source demonstrates Tokyo at 100 KB.

- [S1: Fabian Giesen, Debris: Opening the box, February 13, 2012](https://fgiesen.wordpress.com/2012/02/13/debris-opening-the-box/). First-party procedural tooling account.
- [S2: Farbrausch original demo/tool source](https://github.com/farbrausch/fr_public). Reference architecture, not permission to import assets or assume all code shares one license.
- [S3: TopoJSON specification](https://github.com/topojson/topojson-specification). Shared topology and delta-encoded coordinates.
- [S4: MapLibre Tile specification](https://maplibre.org/maplibre-tile-spec/specification/). Columnar streams and encoding choices; live spec, pin any adopted version.
- [S5: Inverse procedural modeling of facade layouts, 2014](https://doi.org/10.1145/2601097.2601162). Abstract checked; full text unavailable during this pass.
- [S6: CGAL polyline simplification](https://doc.cgal.org/latest/Polyline_simplification_2/index.html). Topology preservation and error costs.
- [S7: meshoptimizer](https://meshoptimizer.org/). Quantization, ordering and runtime mesh considerations.
- [S8: Godot 4.6 MultiMesh guidance](https://docs.godotengine.org/en/4.6/tutorials/performance/using_multimesh.html). Spatial batching and culling caveat.
- [S9: Godot 4.6 web export](https://docs.godotengine.org/en/4.6/tutorials/export/exporting_for_web.html). Compatibility/WebGL2 and export constraints.
- [S10: Official Astra model and prompting guidance](https://developers.openai.com/api/docs/guides/latest-model). Astra-specific behavior and prompting recommendations.
- [S11: Official model/effort guidance](https://learn.chatgpt.com/docs/models). Ultra parallelism and Max distinction.
- [S12: Matt Pocock's skills](https://github.com/mattpocock/skills). First-party engineering workflow examples.
- [S13: Matt Pocock, codebase-design](https://github.com/mattpocock/skills/blob/main/docs/engineering/codebase-design.md). Small interfaces, hidden implementation complexity and evidence for adapters.
- [S14: Martin Fowler, An example of preparatory refactoring, January 5, 2015](https://martinfowler.com/articles/preparatory-refactoring-example.html). Scoped refactoring before adding behavior; includes Kent Beck's formulation.
- [S15: Sandi Metz, The Wrong Abstraction, January 20, 2016](https://sandimetz.com/blog/2016/1/20/the-wrong-abstraction). Avoid preserving a shared abstraction when changing requirements expose incompatible behavior.
- [S16: Cursor PStack source](https://github.com/cursor/plugins/tree/main/pstack). Review, visual parity and independent verification workflow reference.
- [Hamburg integration commit](https://github.com/ThatGuySam/planetwide/commit/f18614c2c93a1e794cde7b7fdff6efcdda9af049).
- [Tulsa integration commit](https://github.com/ThatGuySam/planetwide/commit/5c6d5027c11cad87f65ef53fa05139fd3cfb62a4).
- [CRT and desktop commit](https://github.com/ThatGuySam/planetwide/commit/e1b942da64ccaeae56b1c9bdc64f753a6ed4c7bd).
- Local full Tokyo measurement and audit README, plus the Tulsa compression, browser-performance, winter v5/v6 verification and seed-fitting documents listed above, were read directly. Their measurements are inherited evidence, not rerun benchmarks from this planning session.

Compression Check: Preserved the geographic scope, raw-byte definition, baseline measurements, real-street constraint, 100 landmarks, procedural identity, tracked art/input follow-ups and device limitations. The full source-name inventory is intentionally excluded from runtime requirements following Sam’s clarification; street geometry is unchanged. The exact screenshot defaults and latest Snow 2 completion state remain unresolved. This plan does not include a full Tokyo landmark catalog or new compression results; those belong to execution.
