Evals that run in production

Most evals test individual LLM calls. Agentuity evals test the complete input and output of your agent. Deploy evals alongside your agent code and run them on real production traffic, with results appearing directly in your OpenTelemetry traces.

Production-grade evals, not just CI checks

Evals deploy with your agents and evaluate real user sessions, not just curated test cases. Configure evals in code and deploy with a single command.

Deployed Alongside Your Agents

Evals are packaged and deployed with your agent code. No separate eval service to manage. Run on live traffic and catch regressions in production. Each eval run appears as a span in your OpenTelemetry traces.

Preset Eval Library

Start with ready-to-use evals for safety, PII detection, adversarial attacks, politeness, and more. Each preset returns binary pass/fail or 0-1 scores. Configure thresholds and models to match your requirements.

Custom Evaluators

Define domain-specific quality criteria. Use LLM-as-judge, deterministic checks, or custom scoring functions. Compare model configurations using the same evals on live traffic.

Close the loop on model changes

Gate rollouts based on eval outcomes. Compare different model configurations on the same live traffic. Use eval results to tune prompts, tools, and routing strategies with real data.

See it in action

Watch this 1-minute overview to see how to use Agentuity's Evals to test your agents in production and catch regressions.

Resources

Adding Evaluations
Automatically test and validate agent outputs with binary pass/fail checks or 0-1 quality scores. Create custom evaluators or use LLM-as-judge patterns.

Evaluations API
Reference for createEval(), preset evals from @agentuity/evals, schema middleware, and result types.

Frequently Asked Questions

How are Agentuity evals different from other eval tools?

What preset evals are available?

Do evals impact agent response time?

Ready to build reliable agents?

Ship agents you can trust with continuous evaluation.