// QA & monitoring
QA a checkout flow end to end
Log in to a demo store, buy two items and check the cart math and the order confirmation.
1 agent tool call, click to jump
Why it’s hard
A QA test that survives redesigns has to describe intent, not selectors.
- Credentials must stay out of the model prompt
- The flow spans five pages of state
- Assertions check business rules like totals, not DOM
Stagehand vs. Playwright MCP
Given the same goal and model, the Stagehand code-mode agent finished for 85% less than an agent driving Playwright MCP.
- stagehand code mode agent3/3 runs succeeded
- $0.018
- playwright mcp agent3/3 runs succeeded
- $0.121
medians over successful runs, or over all runs when none succeeded; lower is better
Then run it as a script
Once the flow works, the cookbook runs it as plain Stagehand calls. With caching on, a re-run cost 100% less, because Stagehand served unchanged steps from its cache.
- cookbook, first run3/3 runs succeeded
- $0.080
- cookbook, cached re-run3/3 runs succeeded
- $0
medians over successful runs, or over all runs when none succeeded; lower is better
anthropic/claude-sonnet-5 at $2 input · $0.2 cached · $2.5 cache write · $10 output per 1m tokens · stagehand 4.1.0 · playwright mcp 0.0.82 · agents stop after 60 llm calls · run on 2026-09-27 · methodology
Run it yourself
Give the same goal to Claude Code with Stagehand’s code-mode MCP server, which exposes three tools: run, snapshot and screenshot. Once the flow works, the cookbook runs it as a plain Stagehand script.
On saucedemo.com, log in as standard_user with password secret_sauce, add "Sauce Labs Backpack" and "Sauce Labs Bike Light" to the cart and check out as Ada Lovelace, ZIP 94103. From the checkout overview report each item and price, the subtotal, tax and total; finish the order and report the confirmation heading. Also report these checks as pass/fail: cart has both items, subtotal equals item prices, total equals subtotal plus tax, order confirmed.