eval9 · Evals Planned concept

How well did the agent do?

We are exploring ways to evaluate agent behavior with test cases, explicit scoring criteria, and quality monitoring. The aim is to compare changes and investigate regressions with evidence tied to the task.

Scope is subject to change. No public release or API is announced.

Test casesScoring criteriaQuality comparisonsRegression review
proposed workflow

Define what good looks like, then compare.

The concept starts with representative cases and scoring criteria chosen for your workflow. Results could be compared by case and category, with selected production examples informing future tests. Data collection and integrations are still being designed.

offline suites results by case quality review
eval9 · concept
1. Define representative cases and scoring criteria.

2. Run an evaluation against a candidate change.

3. Compare results and investigate regressions.
areas we're exploring

Test. Compare. Learn.

The proposed direction connects offline evaluation with evidence from real agent work.

01 · Test

Representative evaluation suites

Explore repeatable test cases and task-specific scoring, so teams can compare candidate changes before release.

02 · Compare

Results with context

Explore results by case and category, with enough detail to understand why a score changed rather than relying on one aggregate number.

03 · Learn

Quality monitoring

Explore ways to review selected production examples for changing behavior. Sampling, data handling, and alerting rules remain to be defined.

share your requirements

What would a useful eval measure?

Tell us about the behaviors, datasets, and review criteria that matter in your agent workflows.