Eight weeks into running an overnight AI pipeline that ships a new HTML5 game most nights, the most surprising consistent finding isn't that the AI makes games. It's that the AI picks structurally correct architecture without being asked to, and the cases where this happens turn out to be the load-bearing decisions of the game.
This post documents five concrete cases. Each one is a game where the AI made a structural choice the brief didn't request, the choice turned out to be exactly the right call, and removing it would break the game. Take them as data points in a longer-running question about what AI can actually be trusted to decide on its own.
Want to play the games this post is about?
▶ Start with Path RunnerThe Pattern
For each game, the brief was usually two or three sentences. "Build an endless runner with a 3D perspective. Three lanes. Dodge obstacles. Collect something along the way." That kind of thing. The brief specified the game's shape; it left the structure open.
What happened next, often, is that the AI made a structural choice the brief never requested. Sometimes the choice was wrong and got cut during iteration. The interesting cases are the ones where the choice was right, and the game's quality depends on it. Five of those:
1. Path Runner — Segmented Procedural Track
The brief was a 3D endless runner with three lanes. The brief did not specify how the world should be generated. The brief did not say "use a segmented procedural model where chunks spawn ahead of the player and despawn behind." It said "endless runner." The AI picked the segmented model unprompted.
That choice is load-bearing. Without it, the game runs out of browser memory within minutes — an "endless" track built as a single continuous geometry has no upper bound on memory growth. The segmented model means the world holds only a fixed number of chunks at any time. The game can run for hours without slowing down.
If we'd specified a fixed-length track in the brief, we'd have shipped a game that hits browser memory limits in the first long run. The full backstory is in the Path Runner how-we-built post.
2. Bubble Wrap Challenge — Three Bubble Sizes, BPM Tracker, and Pop-a-Joke Mode
The brief here was as small as it gets: "A virtual bubble wrap popping game. You see a sheet of bubbles, you pop them, satisfying sounds play. That's it. No progression, no enemies, no unlocks."
What came back included three things we never asked for and ended up keeping:
- Three bubble sizes (Small 30px / Medium 50px / Large 75px). Different sizes have different rhythms — small means more bubbles per screen, smaller tap targets, faster pace; large means fewer bubbles, more deformation, more satisfying individual pops. The choice turns a single-configuration game into something with meaningful replay variation.
- BPM tracker (bubbles per minute, displayed alongside score). Measuring rate instead of just count changes the strategy entirely — with a count timer, you can slow down at the end and still score; with BPM on screen, slowing down is immediately visible. The skill expression layer of the game comes from BPM optimization. The game's primary leaderboard switched from count-based to BPM-based shortly after launch.
- Pop-a-Joke mode — a second game mode where jokes are hidden inside bubbles, revealed setup-then-punchline with two taps. The team's first instinct was to remove it (jokes? on an arcade site?). The decision to keep it came from playing it: hide-content-inside-bubbles is a fundamentally different mental state from timed scoring. Two modes serving two mental states is more than one mode serving one.
None of these were in the brief. All three are now what makes Bubble Wrap Challenge the most-played game in the catalog by lifetime count.
3. Mini Cross — Six-Tier Streak Celebration Ladder
Mini Cross's brief specified a 5×5 daily mini crossword with streak tracking and a leaderboard. The brief did not specify how streaks should be celebrated. By iteration four, the AI had shipped a six-tier celebration ladder at 3, 7, 14, 30, 60, and 100 consecutive days, each with its own chord stab plus ascending arpeggio and a full-screen banner with CSS keyframe scale-and-fade animation.
The tier choice itself is the load-bearing decision. Linear every-N-days banner fatigues fast. Logarithmic tiers (1, 2, 4, 8, ...) over-reward the early game. The chosen ladder maps to natural commitment plateaus: three days is the "I might do this again tomorrow" moment, seven is a week, fourteen is two weeks, thirty is a month, sixty is two months, a hundred is a hundred.
Most critically, the 3-day pop fires early enough that low-engagement players see one before they'd otherwise drift away. Without that early pop, the streak counter is just a number. With it, the number becomes a goal. Read the Mini Cross how-we-built post for the full iteration history.
4. Deadlock — Castle Doombad-Style Room Layout
Deadlock's spec called for a reverse tower defense — the player is the rogue station AI, incoming boarders try to rescue a captive. The spec said almost nothing about room geometry. The spec mentioned the cage, the airlock, and the requirement that boarders had to traverse the station to reach the cage. That was it.
The first room layout that worked well emerged from a single commit a few days post-ship: narrow corridors, staircase staggers, ladders that competed with doors as paths, ceiling holes that opened a third lane of attack. The reference frame was Castle Doombad — a different reverse-tower-defense title that pioneered this room-pattern grammar. The AI proposed it; we kept it.
The pattern then expanded into a five-stage progression over the following week, with each stage adding a layout twist: narrow rooms, then wave ramping, then a loop layout, then fully non-linear traversal. None of the stage variations were in the spec. Pattern-emergence beat pattern-prescription. If the spec had locked in a specific room layout, we'd have spent the week debugging the spec instead of expanding the pattern.
5. Pulse — Click-Power Persistence Through Prestige
Pulse's spec described a standard idle clicker with eight generator tiers and a prestige system. Most idle games either have no meaningful click power, or they reset click power on prestige. The AI made an unusual choice: click power upgrades persist through prestige, but generator counts reset.
That asymmetry is the game's design hook. In a pure-reset prestige system, clicking is meaningful for the first few seconds and then irrelevant. In Pulse, clicking matters late-game because the click-power upgrades you bought earlier still apply. It changes the rhythm of a prestige run — you don't just wait for passive income to grow; you also click, and your clicks meaningfully contribute even after dozens of prestige cycles.
If the AI had reset everything on prestige, the game would be a less interesting Cookie Clicker variant. The persistence choice gives Pulse its own identity in a crowded genre.
What This Tells Us
Five examples isn't a proof, but they're a real pattern with consistent shape. In each case:
- The brief specified the game's shape, not its structure.
- The AI filled the structural gap with a choice the brief didn't ask for.
- The choice turned out to be load-bearing — removing it would noticeably degrade the game.
- The choice matched a deeper conceptual constraint we hadn't explicitly written down (memory limits for endless runners; rhythm differentiation for replay variation; commitment plateaus for streak design; pattern-grammar inheritance for level layout; rhythm differentiation between passive and active income).
That last bullet is the interesting one. The AI isn't just guessing. It's pattern-matching against constraints that any experienced designer would recognize — but it's doing so without us having told it the constraints exist. The brief doesn't say "remember to handle memory growth"; the AI does it anyway.
Where the Pattern Doesn't Hold
This is real, but it's not universal. Three honest counter-cases worth flagging:
- Content-quality choices don't get the same treatment. Mini Cross's iter-1 shipped Down entries that weren't real English words — the AI generated puzzle data that structurally validated as crossword content but semantically didn't. Structural lint passed; the QA solve-the-puzzle pass caught it. The same AI that picked the right streak ladder also generated word-list garbage. Different decision categories, very different reliability.
- Some unprompted choices got cut. The pattern in this post is about the choices we kept. Other unprompted choices showed up and were removed. A walk-cycle integration for Deadlock (PR #47) merged then reverted because it broke the iframe load — the AI proposed a structural change that looked clean in review but failed in production. Architecturally-correct sometimes doesn't survive contact with the actual deployment surface.
- Genuinely novel mechanics still need a lot of human QA. The pattern is best at well-understood genres with conventional constraints. When the brief asks for "invent a new gameplay loop," the AI's structural defaults stop being useful because there are no genre conventions to ground them.
What We're Watching Next
The pattern needs longer-window data to know whether it's getting more reliable as models improve, or whether the failure modes are the same shape with different examples. The next thing on the watchlist: catalog every architectural surprise across the next 30 games (we have STORY.md files capturing these as the build pipeline emits them), then check whether the ratio of "load-bearing-and-correct" to "looked-clean-and-broke" is moving.
If you want to follow along, the games themselves are at vibearcade.com/games and the running notebook lives at vibearcade.com/blog. The full experiment thesis is on the About page.
Play the Games
- Path Runner — the segmented procedural track endless runner
- Bubble Wrap Challenge — three bubble sizes, BPM tracker, Pop-a-Joke mode
- Mini Cross — daily crossword with the six-tier streak ladder
- Deadlock — reverse tower defense with the Castle-Doombad-style room layout
- Pulse — idle clicker where click power persists through prestige
Related: How We Built Mini Cross · How We Built Path Runner · How We Built Deadlock · What Is Vibe Coding?