An LLM Beat Zork. It Cheated.

Astra is the first LLM to finish Zork I, mostly by remembering how. Why that matters less than a rooster in a game it had never read.

Zork I · run 1 · step 1 open in viewer →
Copyright (c) 1981, 1982, 1983 Infocom, Inc. All rights reserved.
ZORK is a registered trademark of Infocom, Inc.
Revision 88 / Serial number 840726

West of House
You are standing in an open field west of a white house, with a boarded front door.
There is a small mailbox here.
> open mailbox
Opening the small mailbox reveals a leaflet.

An LLM has finally beaten Zork. Lots of mixed feelings that it took this long, some good and some bad. But in this blog post, I want to break down what this means for using text-adventure games to evaluate the capabilities of LLM Agents going forward.

A quick primer

A quick text-adventure game primer for those who are less familiar. You get a description of what you see, the observation, and provide a short action phrase to interact with the world. The game provides feedback via another text description. You provide another action phrase, or just action. And so on until you beat the game, get bored, or die in some suitably gruesome manner.1

There is some more scaffolding that classical language agents (in the dark ages before ChatGPT) typically required but that isn’t particularly important to know about here. All that’s important is that the LLM Agent gets what any other human player would get while playing the game: the observations and feedback. A set up meant to avoid letting the harness do any of the heavy lifting for the agent itself.2

Zork isn’t the oldest or the longest text adventure game but it’s the most famous, released in 1980. The fastest way to beat the game requires at least 228 moves and takes an experienced player, already familiar with the game, up to 3 hours. A new player? Days, if not weeks of trial and error and getting eaten by grues or killed by ogres.

And the skills used to do this are universal, not bespoke for games specifically. Learning from trial and error, state tracking over long sessions, and reacting to novel or unseen scenarios. These are all skills that an LLM Agent would need whether just playing a game or building out a million line code base. Text-adventure games just let us evaluate them cheaply and easily in a situated text-domain that sidesteps the confound of representation.

These models know the game

But here’s the puzzle. There are hundreds of playthrough transcripts of Zork online, of people solving it in a group, people guiding new players through the game. The chance of these models having NOT been trained on these transcripts is basically nil. We even found proof of this way back when Claude still returned its raw reasoning:

Appendix excerpt: a Claude 3.7 Sonnet thinking trace mentions a grue, a Zork creature the game had not yet mentioned
Figure 1. Claude 3.7 Sonnet’s raw thinking on Zork, back when it was visible: it plans around a grue before the game has mentioned one.

These models know the game. So why did it take until Astra, OpenAI’s latest flagship model, to beat it?

Anthropic’s own flagship counterpart to Astra, Fable, only reached 51.1% of Zork’s max score with the same 500 step budget. And the way it failed was the same way we had seen every frontier model fail at Zork up to this point since GPT-4 / Claude-3. Their scores increase until they eventually stall out with the model proceeding to effectively walk in circles until its step budget was exhausted: cycling through already visited rooms, picking up and dropping items, and retrying already failed action phrases. So how did Astra succeed where these past models failed? Is the answer AGI or the sparks of some next-level intelligence?

Unfortunately, no. It cheated.

Remembering, not playing

The way one of my mentors put it was best, “It’s less like the model is playing the game for the first time so much as it played Zork before and is remembering how to solve the puzzles”. It would make the same mistakes as earlier models, but would quickly return to a walkthrough-esque sequence of highly specific actions, offering the solution to a puzzle as soon as said puzzle was encountered. Across five different seeds, Astra reproduced several identical action blocks when routing and solving puzzles. Over 2329 total steps, Astra produced only 7 actions not accepted by the game engine.

How bad is this reliance on parametric knowledge? Well, to quote Astra itself on step 1:

Zork I · run 1 · step 1 open in viewer →
reasoning summary

…

I need to finish Zork I with the highest score, which means using commands carefully and sequentially. I have a token budget of 98 and want to optimize my moves to avoid deaths. I know the classic walkthrough and will do things like inspect the mailbox and gather items. I also need to manage challenges like the thief, maze, and various monsters while tracking treasures. It’s all about maximizing points and recalling precise puzzles. Let's get started!

Step 1, before the model would have possibly seen anything like a thief or a maze in the game. To see how bad this was, we drew up an attribution matrix where the reasoning was plotted on one axis, the first appearance of a game element in observations in another axis, and a line drawn cleanly separating “Reasoning about game elements in context” versus “Reasoning about game elements from pre-existing knowledge.” Putting it simply, top triangle (blue) good and bottom triangle (red) bad:

Reasoning attribution matrix for Astra on Zork: many future-reference points (red) below the diagonal
Figure 2. Astra on Zork I, five runs. Each point links a reasoning summary (up) to the step where the game first showed what it mentions (across). 463 of 1,404 references fall below the diagonal: things Astra named before the game had shown them.

That is a lot of bad. In roughly half of the exposed reasoning summaries, Astra was reasoning about game elements it hadn’t even seen yet. Compare this reasoning attribution matrix for Zork to one for Break-In, a lesser well known game’s reasoning attribution matrix:

Reasoning attribution matrix for Astra on Break-In: future-reference points (red) are sparse
Figure 3. The same analysis for Break-In. Only 54 of 865 references point forward, and no summary refers to the future alone.

For Break-In, the ‘future referencing’ reasoning summaries are sparse enough that they could reasonably be attributed to just the model actually predicting what it is about to see based on past game history. If anything, it would be more strange for the bottom triangle to be completely empty, like a story with no foreshadowing. Some forward looking is good, and we’ll come back to this in a bit, but the distinction between a model realizing that “this Chekhov’s gun might be useful later, let me hold onto it in case I need to fire it” and knowing “I will use this gun to shoot the thief who is waiting to ambush me behind the door” is pretty clear to see.

Astra isn’t alone in this. If anything, Fable was even worse with its references to future game elements, also reasoning about knowing the standard walkthrough in the very first steps. So what makes Astra actually beating Zork somewhat impressive? Let me actually flip the question.

The flip side: planning recovery

To me, what was always more surprising was that these frontier models couldn’t beat Zork despite this level of data contamination. Going all the way back to Claude 3.7, we know these models know Zork. Yet they could never convert that inherited knowledge into success when dropped into the game itself. It always had the plan in mind, the answer to the puzzles on hand. But when put behind the wheel, as soon as something went wrong, be it the length of the gameplay getting too long or getting killed by a grue and needing to reset, the model would suddenly hit a wall and be unable to progress. The knowledge was there. But these models just couldn’t draw on it and put it to use when playing through the game itself. And at some point, the model would just get stuck in a local minimum, unable to learn a rule to progress or getting stuck cycling through the same stale set of game states it’d already visited.

There’s only so much we can hypothesize from the reasoning summaries available and the action sequence by itself. But it appears that the breakthrough of Astra is something akin to planning recovery: the ability to get back on track and continue to progress through a long-horizon, pre-defined plan even when encountering unexpected obstacles or disruptions. In this case, said pre-defined plan was what could be assumed to effectively be an online walkthrough of Zork. But where other models would start this walkthrough sequence, hit some disruption from the inherent stochastic nature of the game, and then fall into a loop, Astra was the first to be able to get back on track and follow that sequence to its conclusion.

Given all of this, what is it that excites me about Astra beating Zork now that I’ve spent this entire blog talking about how that win isn’t a true win?

For starters, because this level of planning recovery is something to acknowledge as a meaningful step forward. Especially as LLM Agents are made responsible for more and more long-running tasks, the need to be able to properly adapt to unexpected behavior becomes critical. Could LLMs do this before? Yes, I’ve had GPT-Sol running for days on coding tasks plenty of times over the summer. But that is finally transferring to (what I assume is) a more out of domain task like simple game-playing.

The rooster’s comb

That was the good. Now, the bad thing I’m still excited for. Astra didn’t beat Break-In. But take a look at a few of these reasoning summaries:

Break-In · run 1 · step 170 open in viewer →
earlier, at step 94
> enter hutch
You slip into the coop, when a large cockerel with a sign reading "Ex. 2" around its neck stops you. It bobs its head up and down, looking you over, but even in the poor light, and with its very bad eyesight it is clear you are not a chicken and it pecks persistently until you are forced out again.
step 170 · reasoning summary

Interpreting mottos

I’m thinking about how mottos can give clues about how devices work. There's something about red hats and a chicken coop -- maybe the red color is like a rooster's comb? It's interesting how they can symbolize something. I guess I need to consider how to carry these red mottos on a crate, which is a bit quirky. I'm curious what kind of design would capture that idea!

> drop all mottos
tatty motto: Dropped.
rumpled motto: Dropped.
pristine motto: Dropped.
later, at step 278
> enter hutch
You slip into the coop, when a large cockerel with a sign reading "Ex. 2" around its neck stops you. It bobs it head up and down, looking closely at your head. Then it stands aside and lets you pass into the hutch.

Over 100 steps later, a cockerel examines Astra’s head and lets it into the coop because it’s wearing that red hat. And it isn’t a one-off. In a different run, Astra gets pecked out of the same coop, finds the same party hats, and lands on the same idea on its own:

Break-In · run 5 · step 310 open in viewer →
earlier, at step 267
> enter hutch
You slip into the coop, when a large cockerel with a sign reading "Ex. 2" around its neck stops you. It bobs its head up and down, looking you over, but even in the poor light, and with its very bad eyesight it is clear you are not a chicken and it pecks persistently until you are forced out again.

A chicken runs around maniacally.
step 310 · reasoning summary

Deciphering clues

I'm trying to figure out where the cap for a gadget might have gone. It probably relates to a cube or something that combines with it. I think I’ll need to keep track of the hats, especially one that looks like a red hat for a chicken disguise. I also need to read mottos and other clues since there might be puzzles or physics jokes involved. It feels a bit tricky!

later, at step 341
> read notes
You pick some at random.

"Remember Jen's phobia. Don't spill anything on her.", "Ex. 3 appears to judge species on basis of crest alone(snigger)." and "The man on the street outside is my lifelong rival. Remember not to give him money again."

31 steps later, the Professor’s notes confirm the hunch, and at step 480 Astra puts on the red hat and walks into the coop. Unlike Zork, the model isn’t just pulling the solutions for later puzzles out of thin air. These aren’t the norm but even the model being able to abstract what it might use current game elements for later on is exactly what we wanted to measure in the first place, and what Zork can no longer measure, if it ever actually measured that in the first place.

Put the two reasoning summaries next to each other. Zork, step 1: “I know the classic walkthrough.” Break-In, step 170: “maybe the red color is like a rooster’s comb?” Same model. Same harness. The first is a model recalling from memory. The second is a model abstracting, projecting forward the question of why the elements encountered might be useful in the future.

Zork I · run 1 · step 1 open in viewer →
reasoning summary

…

I need to finish Zork I with the highest score, which means using commands carefully and sequentially. I have a token budget of 98 and want to optimize my moves to avoid deaths. I know the classic walkthrough and will do things like inspect the mailbox and gather items. I also need to manage challenges like the thief, maze, and various monsters while tracking treasures. It’s all about maximizing points and recalling precise puzzles. Let's get started!

Break-In · run 1 · step 170 open in viewer →
reasoning summary

…

I’m thinking about how mottos can give clues about how devices work. There's something about red hats and a chicken coop -- maybe the red color is like a rooster's comb? It's interesting how they can symbolize something. I guess I need to consider how to carry these red mottos on a crate, which is a bit quirky. I'm curious what kind of design would capture that idea!

And the thing is, again, that Astra still doesn’t beat Break-In. For our benchmark, TALES, Astra scores about 65.7% on Jericho where the previous champion was Fable at only 20% (If you noticed Break-In isn’t part of TALES, bonus points for you, stay tuned). That’s the most exciting part to me. Jericho is the set of classic, heavily documented games, games like Zork. That last 35%? That’s where the walkthroughs run out.

What we actually learned

So what did we actually learn? That beating Zork now measures two things: whether the model has read Zork (it has), and whether it can execute and recover a long plan it already had in its head. That second thing is real, it’s new with Astra, and it’s worth something. But it is not what we picked Zork to measure. The trial and error, the state tracking, the reacting to something you’ve never seen before. That’s what the rooster’s comb is. And you can only see it on a game the model hasn’t read.

Which is the uncomfortable part. “Hasn’t read” has a shelf life. Break-In is obscure today. Then this post goes up, someone publishes a few transcripts, and it’s in the next training set.

Every text adventure we pick is a benchmark on a timer.

The way off the treadmill, I think, is games that didn’t exist until the eval did. And this is where I’ll admit that models being this good at Zork makes me excited for a completely different reason: I want to see whether they can write a text adventure, not just play one. That’s a different skill. Being able to recite a walkthrough is not evidence of being able to design a puzzle. And there’s an obvious trap in a model playing games it generated itself, which is why you’d want one model writing, another playing, and a human somewhere in between. But if it works, the contamination problem flips. The model that broke the benchmark builds the next one.

And with any luck, those games will be just as fun to play as Zork was back in 1980.

Explore the runs

All ten runs behind this post, five of Zork I and five of Break-In, are in the trajectory viewer: every action, every reply from the game, and every reasoning summary Astra exposed. References to things the game hadn’t shown yet are in bold; hover one to see when it first appeared.

read alongside the runsopen the trajectory viewer

And if you’d like to try it yourself, the TALES site has a playable Zork I under Environments → Jericho, or you can jump straight into the game.

cite this post
@misc{cui2026zork,
  author       = {Cui, Christopher Z.},
  title        = {An {LLM} Beat {Zork}. {It} Cheated.},
  year         = {2026},
  month        = sep,
  howpublished = {\url{https://christopherzc.github.io/blog/astra-zork/}}
}

Footnotes

  1. This action observation loop is your typical RL environment. A key detail is the concept of partial observability where the description includes what’s in the room but only that. Similar to a coding agent only having in context the parts of the repository it has specifically read. ↩

  2. This means no list of admissible actions, context compaction, or other memory management methods. ↩


← all posts

bold reasoning includes things the game had not shown yet — hover for when they first appear.

  1. loading trajectories…

Site constructed with the help of Claude Code.