# Synthetic playtest

- **Event:** [GPT-6 Astra Hackathon SF](https://cerebralvalley.ai/e/openai-gpt-6-astra-sf)
- **When:** Tue, Sep 8 at 9:00 AM – 10:00 PM (PDT)
- **Where:** San Francisco, CA
- **Team:** [Kartik Pandey](https://cerebralvalley.ai/u/kapie)
- **GitHub:** https://github.com/KARTIK-PANDEY-KP/synthetic-playtest
- **Demo video:** https://youtu.be/ObWNA_zUkZk
- **Gallery:** https://cerebralvalley.ai/e/openai-gpt-6-astra-sf/hackathon/gallery
- **Page:** https://cerebralvalley.ai/e/openai-gpt-6-astra-sf/hackathon/gallery/44

--Due to time crunch, the video has no audio--

The same thing happens every time production outruns testing.

When AI made coding fast, engineers started shipping faster than anyone could review, and buggy code went out the door — so testing had to become automated, and a whole industry of test infrastructure grew up to keep pace. When voice agents got easy to build, teams shipped them faster than any human could call them — so voice-AI testing became persona-driven: you write the impatient caller, the confused senior, the one who mumbles, and let them hammer the agent thousands of times. That's how that industry ships today.

Games are next, and the gap is bigger. With GPT-6 Astra, game development and 3D modeling are about to go from months to hours — more titles, more levels, more builds a day than there are people to play them. But games have a kind of testing code never needed: beta testers and early-access players, hundreds of them, playing the way real people play — skimming the tutorial, getting lost, getting bored, quitting at a door that won't open. That's the only way a studio learns whether a game is confusing, not just whether it's broken. It costs weeks, it costs real money, and it runs at human speed. Production is about to stop running at human speed. The bottleneck in game production is about to be testing — and it needs the same answer testing always gets: automate it, with agents that behave like the people they replace.

That was impossible until now, for one reason: no model could play a game it had never seen. Astra can. The model that makes games fast is the model that makes testing them possible.

Synthetic Playtest is the beta program, run by agents. You describe testers as people — Maya, a streamer who skims everything; Robert, 52, first PC game ever; Dana, who plays with the sound off. Each one gets its own cloud sandbox with a browser and an Astra agent and plays the game the way that person would — not because a prompt asked nicely, but because the harness enforces it: Maya's screenshots have long text blurred out before the model ever sees them; Dana has no "listen" tool at all. They see only pixels and press only keys. As they play they file what a beta tester would file — confusion, boredom, unfairness, bugs — in their own voice. Reports are deduplicated across the fleet, attributed to the game or to the agent using telemetry the agents never see, and every finding with a ground-truth signal is replay-verified by re-running the exact inputs at the same seed. Make as many personas as you want; watch every one of them live.

To prove it works we built the test subject too. Station Kepler: you wake alone on a dead orbital research station as its last signal, and have to bring it back to life — find the fuse and the keycard, restore power, revive the greenhouse, restart the reactor before the air runs out. Six rooms, first person, and seeded with fifteen deliberate flaws and two decoys, so we hold an answer key and can score ourselves. First real run — five personas, $33: three planted flaws found, zero false alarms, and one real problem nobody planted — four silent, identical locked doors that stalled four of five testers, exactly what a hundred early-access players would have found in week one. Found before week one, in twenty minutes.

And this isn't only for games. The agents don't know they're in a game; they know they're a person in a 3D space, looking and trying. Make any 3D asset walkable — a building, a vehicle interior, a product model, a level block-out — and the same fleet walks it as a first-timer, a power user, someone who can't hear, someone who won't read, and reports what each of them noticed and would change. We didn't build that tonight, but nothing in the pipeline is game-specific. The pattern is human personas, enforced, at scale, doing a job that currently needs a room full of people to try something and say how it felt. Testing is the first such job. It won't be the last.

---

Markdown version of https://cerebralvalley.ai/e/openai-gpt-6-astra-sf/hackathon/gallery/44. Site index for agents: https://cerebralvalley.ai/llms.txt · full text: https://cerebralvalley.ai/llms-full.txt
