Pull the latest quarterly results of a public company
investors.digitalocean.com
// Research & public data
Read two benchmark leaderboards on stagehand.dev/evals and return the top 10 models of each as JSON.
5 agent tool calls, click to jump
The data lives in an interactive client-side view, not a file you can download.
Given the same goal and model, the Stagehand code-mode agent finished for 52% less than an agent driving Playwright MCP.
medians over successful runs, or over all runs when none succeeded; lower is better
Once the flow works, the cookbook runs it as plain Stagehand calls. With caching on, a re-run cost 100% less, because Stagehand served unchanged steps from its cache.
medians over successful runs, or over all runs when none succeeded; lower is better
anthropic/claude-sonnet-5 at $2 input · $0.2 cached · $2.5 cache write · $10 output per 1m tokens · stagehand 4.1.0 · playwright mcp 0.0.82 · agents stop after 60 llm calls · run on 2026-09-28 · methodology
Give the same goal to Claude Code with Stagehand’s code-mode MCP server, which exposes three tools: run, snapshot and screenshot. Once the flow works, the cookbook runs it as a plain Stagehand script.
On stagehand.dev/evals, read the leaderboard for the "Browserbase Benchmark v2" benchmark and then for the "Online Mind2Web" benchmark. For each, report the top 10 rows with rank, model, harness, accuracy in percent, cost per task in USD and seconds per task (null if not shown).
investors.digitalocean.com