How well did the agent do?
We are exploring ways to evaluate agent behavior with test cases, explicit scoring criteria, and quality monitoring. The aim is to compare changes and investigate regressions with evidence tied to the task.
Scope is subject to change. No public release or API is announced.
Define what good looks like, then compare.
The concept starts with representative cases and scoring criteria chosen for your workflow. Results could be compared by case and category, with selected production examples informing future tests. Data collection and integrations are still being designed.
1. Define representative cases and scoring criteria. 2. Run an evaluation against a candidate change. 3. Compare results and investigate regressions.
Test. Compare. Learn.
The proposed direction connects offline evaluation with evidence from real agent work.
Representative evaluation suites
Explore repeatable test cases and task-specific scoring, so teams can compare candidate changes before release.
Results with context
Explore results by case and category, with enough detail to understand why a score changed rather than relying on one aggregate number.
Quality monitoring
Explore ways to review selected production examples for changing behavior. Sampling, data handling, and alerting rules remain to be defined.
Data and execution for quality work.
These products provide separate building blocks for agent workflows. An eval9 integration is not available today.
Session records
Capture and review native transcripts from supported agent runtimes with collection enabled.
owl9 → run9Sandboxes
Create Boxes, execute commands, and fork filesystem snapshots with run9.
run9 → db9Postgres database
Create and query Postgres databases, with database branching for separate work.
db9 → mem9Agent memory
Store and retrieve context for agent workflows through mem9’s memory interfaces.
mem9 →What would a useful eval measure?
Tell us about the behaviors, datasets, and review criteria that matter in your agent workflows.