// Games

Solve today's Wordle

Play the NYT Wordle end to end: read the tile colors after every guess and pick the next word until it is solved.

stagehand code-mode agent (mcp facade), browserbase recording at 2× speed, no sound

16 agent tool calls, click to jump

++

Why it’s hard

A game is a loop of acting, reading state and deciding, where every guess depends on the last board.

  • Tile colors only appear after a flip animation settles
  • Guesses outside Wordle's word list are rejected and must be retried
  • Duplicate letters follow Wordle's own scoring rules

Stagehand vs. Playwright MCP

Given the same goal and model, the Stagehand code-mode agent finished for 32% less than an agent driving Playwright MCP.

stagehand code mode agent3/3 runs succeeded
$0.180
playwright mcp agent3/3 runs succeeded
$0.263

medians over successful runs, or over all runs when none succeeded; lower is better

Then run it as a script

Once the flow works, the cookbook runs it as plain Stagehand calls. With caching on, a re-run cost 19% more than the first run. 1 of 3 cached re-runs failed, so treat that figure with care.

cookbook, first run2/3 runs succeeded
$0.035
cookbook, cached re-run2/3 runs succeeded
$0.042

medians over successful runs, or over all runs when none succeeded; lower is better

anthropic/claude-sonnet-5 at $2 input · $0.2 cached · $2.5 cache write · $10 output per 1m tokens · stagehand 4.1.0 · playwright mcp 0.0.82 · agents stop after 60 llm calls · run on 2026-09-28 · methodology

++

Run it yourself

Give the same goal to Claude Code with Stagehand’s code-mode MCP server, which exposes three tools: run, snapshot and screenshot. Once the flow works, the cookbook runs it as a plain Stagehand script.

the goal both agents were given

Play today's Wordle at nytimes.com/games/wordle and solve it in six guesses or fewer. Report whether you solved it, the answer (null if not solved), and each guess in order with its tile feedback as five emoji (🟩 correct, 🟨 present, ⬛ absent).

++