Manifesto

The Gauntlet Loop, and what it is actually good for

One prompt, eleven adversarial critics, four rounds, and a playable shooter. It never once beat its own reference. The record is more useful than the demo.

16 Aug 202613 min read
Unispawn share card on a near-black ground. A volt-yellow eyebrow reads Blog, the headline reads The gauntlet loop: games are made by argument now, and a volt rule runs along the bottom edge.

In July 2026, Matt Shumer published a short prompt and the 55,000 lines of code it produced. The prompt asks for a first-person shooter "at the level of the most recent Call of Duty games", in ThreeJS, and then says almost nothing about how. A few days later he wrote up the method behind it and named it: the Gauntlet Loop.

Both the name and the pattern are his. What follows is a reading of his method and, more usefully, of his own record of running it — because the repository contains an honest self-assessment, and it disagrees with the method in two specific and instructive places.

Key takeaways

  • The Gauntlet Loop is a build-and-critique cycle: a lead agent breaks a goal into pieces, each piece gets a builder and a separate blind critic, and the critic compares the rendered artifact against a concrete reference until it wins or you stop it.
  • The bar is the load-bearing part, and Shumer is explicit that it need not be reachable. In the original it never was: eleven critics, four rounds, and in every blind comparison every critic still picked the real Call of Duty frame.
  • His own repo found that parallel fan-out lost to sequential single ownership — the opposite of what the prompt emphasises, and a finding the generalising essay does not carry forward.
  • The critics were reliably right that something was wrong and reliably wrong about what. The single most valuable fix came from an agent contradicting its own brief.
  • It has been demonstrated on browser games and, so far as the public record goes, only browser games. No one has run it against a same-budget, no-critics baseline.

What the loop actually is

Strip the essay down and there are three moves that matter.

One: give a goal, not an implementation. The prompt names a destination and refuses to specify the route. No architecture, no file layout, no engine loop. The lead agent decides all of it, including how to break the work up.

Two: set a bar the critic can actually fetch and compare against. This is the part Shumer insists on hardest — "The bar is the most important part" — and his negative examples are the useful ones: "'Make it amazing' is not a bar. Neither is 'make it production-ready' or 'keep improving it.'" A bar has to be a concrete artifact. Real screenshots of the game you want to match. The best sites in your category. Paragraphs whose clarity you want to reach. And, counter-intuitively: "A hard bar does not need to be realistically reachable."

Three: separate the builder from the critic, and keep the critic blind. Shumer's reasoning here is the sharpest thing in the essay:

The builder has seen every decision it made. It remembers why it made them. That makes it very good at explaining why its work is reasonable. You do not want reasonable. You want an independent judgment.

So the critic never sees the builder's history or its explanations. It inspects "the real pixels, running product, rendered page, test results, or finished writing", compares against the bar — blind A/B where possible — names the single biggest remaining gap, and sends it back.

Then you loop. Not for three rounds: "Do not tell it to do three rounds and stop." You stop when you like the result, when improvements stop mattering, or when you have spent what you are willing to spend.

One goal prompt fans out to three parallel builder agents. Each builder renders a screenshot which its own critic agent compares blind against reference material. A critic either rejects, looping the work back to its builder, or passes. Only when every critic passes is the bundle emitted.
The Gauntlet Loop as the prompt specifies it. In the actual run, the critics never all passed, and the parallel fan-out shown here performed worse than giving one agent sequential ownership.

Why the critic works, as far as anyone can tell

The honest answer is that nobody has separated the variables. Three things are bundled together — the critic is blind to the authoring, it inspects a rendered artifact, and it is adversarially mandated — and no source isolates any of them. Shumer reasons only about the blindness.

The weak evidence points at the rendering leg. Critic reliability in the original improved materially only once captures became deterministic, and the critics' verbal diagnoses turned out to be wrong while their detection of a defect was right. More on both below.

If you take one idea from this, take the general form: anywhere you can render a system's output into something a fresh agent can inspect, you can run a critic that never wrote the code and has no stake in defending it. That is not specific to games.

The record, which is better than the demo

Shumer's README is unusually honest, and it opens: "The goal was to match a modern Call of Duty. It does not."

Eleven independent adversarial critics scored frames against that bar across four rounds. The scores went 3.59 → 4.14 → 4.05 → 5.05 out of ten. Two shots reached "close"; the rest stayed "amateur". And the line that matters most: "In a blind A/B, every critic in every round picked the real Call of Duty frame."

The loop never won. Not once, in any round.

That is not a debunking. The thing is genuinely playable and it runs in a browser. Every texture, mesh, animation and sound is generated procedurally at load time; there are no art assets at all. The point is narrower and more useful: the stopping condition in practice was the human, not the bar. The run stopped after four rounds, with the score still climbing. An unreachable bar is a feature precisely because it never terminates the loop, which means the loop terminates when your patience or your budget does.

The finding the essay leaves behind

The process note in the same README contains the single most actionable result in the whole corpus, and the generalising essay does not carry it forward:

Sequential single-owner passes beat parallel fan-out decisively. Three rounds of six agents each owning one directory moved the score +0.46 and left frame-ruining defects higher than they started (60 → 47 → 66), because tonemapping, sky and indirect light are one coupled system and isolated agents kept breaking each other's assumptions. One sequential pass with a single owner per coupled concern moved it +1.00 and cut defects 66 → 26.

Fan-out is the part of the pattern everyone quotes: parallel agents are what make it look like a machine rather than a chat. On the repo's own evidence it is also the part that hurt, wherever the subsystems were coupled — and the essay leaves the decision to the lead agent without mentioning the result.

So: fan out across genuinely independent pieces, and give coupled systems one owner. That is the practical rule, and it comes from the primary source rather than from anyone's theory.

The critics were confidently wrong

The second finding is subtler and, if you plan to run one of these, more important. Also verbatim:

The most valuable single result came from an agent contradicting its own brief. Every critic for three rounds reported the weapon as "untextured". It wasn't — it was specular-dominated, with the diffuse term measured at L=26 against a shipped L=67. Prior rounds had been crushing albedos to fight bright-part complaints, which killed diffuse and made it worse. The fix was the opposite of what was asked for.

Three rounds of critics agreed on a diagnosis. The diagnosis was wrong. Builders that complied made the artifact worse, and the fix came from an agent that refused the instruction.

The lesson is not that critics are useless — they correctly and persistently detected that the weapon was wrong. It is that a critic's judgement that something is bad is far more reliable than its account of why. Treat the verdict as a signal and the explanation as a hypothesis.

The harness was the real work

The README says it plainly: "The interesting part of this repo is arguably the harness, not the game." Two findings from building it invalidated earlier measurements outright.

"Median frame time hides the actual problem" — a static-camera benchmark reported 94 fps on a build that ran at 12–17 fps in motion. And "Captures were not reproducible": the capture tool reused one page, so particle age, decal buffers and exposure leaked between shots, and two identical runs differed on ten of eleven shots.

That second one is the load-bearing fact of the entire exercise. Until captures were deterministic, the blind A/B was comparing against noise. The critics became trustworthy only after a human built a rig that made their input stable. Skip the harness and you have built a machine for generating confident opinions about randomness.

What running one actually costs

Neither primary source says. The README carries no token count, no wall-clock time and no bill, and the essay carries none either. Every figure in circulation for the original run — including the one below, which is ours — is somebody's reconstruction.

We modelled a tuned version while designing our own — cheap models for critics, mid-tier builders, the expensive one only for architecture passes, aggressive prompt caching, screenshots capped and downsampled:

Our costing, 2026-08-04. Roughly $15 to $35 per accepted game, worst runs two to three times that, priced against then-current token rates. Five components at about three iterations each, 200–500K cumulative input and 30–60K output per builder session; critic verdicts at four capped screenshots plus rubric, 25–40K input each, fifteen to thirty verdicts per run; one frontier-class orchestrator across the session. The number in this post most likely to be wrong by the time you read it, in the direction of cheaper.

Grouped bar chart comparing the estimated low and high cost in US dollars of three parts of a tuned gauntlet run per accepted game: builder sessions, critic verdicts, and the orchestrator.
Source: Unispawn internal costing, 2026-08-04 — an estimate against then-current token rates, not a billed figure.

The split is the useful part: the critics, which are the mechanism, are the cheapest thing in the run. The money is in the builders.

Where it breaks

The critic can only see what survives a screenshot. This is the strongest objection and it comes from Karpathy, who notes models "can't easily audit their work because they aren't able to efficiently and natively perceive videos or play games within them". The original's scoring set was eleven static shots. A frame cannot express pacing, difficulty curve, game feel, audio, or whether a level is boring. Everything the loop optimised was, by construction, the part that photographs well. That is also, precisely, the part the critics could compare against Call of Duty stills.

The genre may have been in the training data. Three.js ships pointer-lock camera controls as an official example. Mouse-look plus WASD plus raycast gunfire has been forked and tutorialised for over a decade. As Decrypt puts it, that "doesn't make Claude of Duty fake, but it makes 'built from scratch' a harder claim to fully credit."

Nobody has run the obvious experiment. There is no same-budget, no-critics baseline anywhere in the public record — nothing that gives one agent the same budget, the same reference material and no critics at all, and reports what comes out. Until someone runs one, the null hypothesis stands: a frontier model with a large budget and a good spec may simply be very good at Three.js shooters.

And it has only been demonstrated on browser games. The method is claimed for code, websites, product design, marketing campaigns, writing and research. First-hand accounts with specifics, on any non-game task, do not appear to exist yet.

So where does that leave the method

Roughly here. The blind critic inspecting a rendered artifact is a real and portable idea, and the cheapest part of the run. The unreachable bar is a genuinely clever trick for never terminating early. The fan-out is oversold and actively harmful on coupled systems. The verdicts are trustworthy and the explanations are not. And none of it works until something deterministic exists to render against.

Run one. The prompt and the essay are both public and it costs whatever your tokens cost.

How Unispawn came out of this

We are a marketplace for browser games made this way, and we got here by trying to build the other thing first.

The obvious product is a text box: describe a game, we run the loop, you get something playable. We designed it — an inspectable bar the author supplies rather than a hidden rubric, hard budget caps, resumability, a live progress page, and the generated bundle going through the same security pipeline as any human upload, because a bundle we produced is not more trustworthy than one we received.

Three of the four things we found are in this article already: what it costs, that the bar has to be somebody's property, and that the output cannot be trusted. That last one has a sharp edge in a product. Reference images are model input, and anything a model reads can hide an instruction; a poisoned reference could persuade a builder to write a credential into the bundle, which a progress page would then publish as a playable preview, promptly and helpfully. The fix would not be clever, only strict:

Three zones. The runner container holds builder and critic agents and an output directory and carries no platform credentials. An egress proxy outside the container holds the model API key. A host-side collector reads the output directory and hands the bundle to the standard review pipeline. A poisoned reference entering the container stops at its boundary because there is nothing inside worth stealing.
How a generation runner would have to be built. None of this exists; it is the design that would have been required.

The fourth finding is the one that isn't about the technique at all, and it is why this is a marketplace.

Safe harbours protect an intermediary that stores material at another person's direction. The DSA shelters neutral, technical handling of information provided by the recipient of the service, not information the provider supplies itself (Regulation (EU) 2022/2065, recital 18, operative rule in Article 6). The DMCA is blunter: storage has to be "at the direction of a user" (17 U.S.C. 512(c)).

Generate the game and you are the one providing the information. Not the host. The author. And you cannot hand that authorship back: the US Copyright Office concluded in January 2025 that "prompts alone do not provide sufficient human control to make users of an AI system the authors of the output" (Part 2: Copyrightability). So a site that generates your game and tells you that you own it is saying something untrue, in the one place where a creator most needs the truth.

Four routes out, what each one makes you, and why each one fails:

  • Generate it, keep the authorship. A studio with a suggestion box. The people using it are requesters, not authors — nothing to own, nothing to be paid for.
  • Generate it, tell the user they own it. A marketplace built on a false statement. The consequence lands on the author, not the platform, and at the worst moment.
  • Generate it, say honestly the raw output may not be protectable. Accurate, and unattractive. Almost nobody would ship it, but it exists.
  • Don't generate it. A host. Costs you the feature everyone asks for.

We took the last one. This is our own reading, not legal advice, and it is US-centric: the UK still deems the author of a computer-generated work to be the person by whom the arrangements necessary for its creation are undertaken (CDPA 1988 s.9(3)), though the UK Government's March 2026 report proposed removing that protection. Our conclusion follows from wanting to be a host. It does not follow from wanting to make games.

So you run the loop, on your machine, under your direction, and the game is yours. What we do is the part that is hard to do alone. A home in a browser tab with nothing to install. Isolation on its own origin, so nobody has to trust the code. A person in front of every build before it goes public. And payment out of the subscriptions of the people who actually played it. How games rank is published in full, formula by formula, and placement cannot be bought.

The obvious questions

Do you generate games?

No. Games are made by their authors, using whatever tools those authors choose, and published here. The platform reviews, hosts, isolates, ranks and pays. It does not write.

Can I publish a game I built with Claude, Cursor, Bolt or Replit?

Yes, and that is the normal case. The tool is not the question. The question is whether you directed the work and whether the result is yours to publish.

If AI wrote most of the code, do I still own the game?

Probably, and the reason matters. Current US Copyright Office guidance holds that purely prompt-generated output is not protected, while human contributions — selection, arrangement, modification — can be. A game whose structure you designed, whose generated parts you edited, and whose assembly is your own creative arrangement looks very different from one line of prompt. Note the Office's warning: re-running prompts and picking a favourite output is expressly not enough on its own. Our own reading, not legal advice.

Will you add generation later?

There is no plan to. What is above is an abandoned design, not a roadmap item.

More on publishing, review and payment is in the FAQ, and you can see what has been published so far.


This is the first post on this blog. It seemed right to start with the technique rather than the product, since the technique is what changed and the product is only one response to it.

All posts