Build
Create an evaluation suite with realistic scenarios, edge cases, and checks. Start with:Run
Simulate realistic user interactions and evaluate how your agent behaves across different paths.Debug
Trace failures back to conversations, tool calls, and decisions.Improve
Verify fixes, catch regressions, and continuously improve your agent.CLI reference
Command reference for automation and CI/CD pipelines.
Evaluation concepts
Core mental model and key evaluation terms.