// QA & monitoring
Catch UI regressions with one test across app variants
Run the same checkout test against a working store and two deliberately buggy builds, and report what broke.
41 agent tool calls, click to jump
Why it’s hard
A selector-based test has to be rewritten per variant; one written as intent runs unchanged and reports the bugs it finds.
- The buggy builds keep the page structure but break behaviour
- Failures must be reported as findings, not crash the run
- Each account needs a clean cart, stored in the browser
Stagehand vs. Playwright MCP
Given the same goal and model, the agent driving Playwright MCP did not complete the task in any of its runs; the Stagehand code-mode agent completed 2 of 3.
- stagehand code mode agent2/3 runs succeeded
- $0.328
- playwright mcp agent0/3 runs succeeded, median of failed runs
- $0.536
medians over successful runs, or over all runs when none succeeded; lower is better
Then run it as a script
Once the flow works, the cookbook runs it as plain Stagehand calls. With caching on, a re-run cost 100% less, because Stagehand served unchanged steps from its cache.
- cookbook, first run3/3 runs succeeded
- $0.255
- cookbook, cached re-run3/3 runs succeeded
- $0
medians over successful runs, or over all runs when none succeeded; lower is better
anthropic/claude-sonnet-5 at $2 input · $0.2 cached · $2.5 cache write · $10 output per 1m tokens · stagehand 4.1.0 · playwright mcp 0.0.82 · agents stop after 60 llm calls · run on 2026-09-27 · methodology
Run it yourself
Give the same goal to Claude Code with Stagehand’s code-mode MCP server, which exposes three tools: run, snapshot and screenshot. Once the flow works, the cookbook runs it as a plain Stagehand script.
On saucedemo.com (password secret_sauce), run the same checkout test for each of standard_user, problem_user and error_user, starting each one with an empty cart: log in, add "Sauce Labs Backpack" and "Sauce Labs Bike Light" to the cart, open the cart, check out as Ada Lovelace with ZIP 94103, continue and finish. For each user, report pass/fail for these checks, stopping a user's test at the first failed check: "both items in cart", "checkout info accepted", "order confirmed"; and list any issues found, in plain words.