Evals that run in production
Most evals test individual LLM calls. Agentuity evals test the complete input and output of your agent. Deploy evals alongside your agent code and run them on real production traffic, with results appearing directly in your OpenTelemetry traces.
Production-grade evals, not just CI checks
Evals deploy with your agents and evaluate real user sessions, not just curated test cases. Configure evals in code and deploy with a single command.
Deployed Alongside Your Agents
Evals are packaged and deployed with your agent code. No separate eval service to manage. Run on live traffic and catch regressions in production. Each eval run appears as a span in your OpenTelemetry traces.
Preset Eval Library
Start with ready-to-use evals for safety, PII detection, adversarial attacks, politeness, and more. Each preset returns binary pass/fail or 0-1 scores. Configure thresholds and models to match your requirements.
Custom Evaluators
Define domain-specific quality criteria. Use LLM-as-judge, deterministic checks, or custom scoring functions. Compare model configurations using the same evals on live traffic.
Close the loop on model changes
Gate rollouts based on eval outcomes. Compare different model configurations on the same live traffic. Use eval results to tune prompts, tools, and routing strategies with real data.
See it in action
Watch this 1-minute overview to see how to use Agentuity's Evals to test your agents in production and catch regressions.
Resources
Adding Evaluations
Automatically test and validate agent outputs with binary pass/fail checks or 0-1 quality scores. Create custom evaluators or use LLM-as-judge patterns.
Evaluations API
Reference for createEval(), preset evals from @agentuity/evals, schema middleware, and result types.
Frequently Asked Questions
How are Agentuity evals different from other eval tools?
What preset evals are available?
Do evals impact agent response time?
Ready to build reliable agents?
Ship agents you can trust with continuous evaluation.