When a developer asks Claude Code or Codex to build their app and use your product, the agent is the one reading your docs, running your CLI, and deciding what to do when a command fails. We're all iterating on our MCP and Skills to improve this experience, but it's hard to actually evaluate this.
We built a tool that runs these agents against real tasks in isolated sandboxes, then changes one thing at a time: CLI or MCP, skills on or off, this week's CLI release or last week's. It also separates the help an agent genuinely needs (like signing in) from the unnecessary human intervention due to a gap in the experience.
We'll use those runs to look at how to design for agents: docs they can actually find, responses that point to the next step, and output they don't have to dig through.
This talk has been presented at JSNation US 2026, check out the latest edition of this JavaScript Conference.





















