Agentic evaluation is where crowd models break. Task authoring, environment engineering and trajectory grading are engineering disciplines, and they are the founding team’s specialty.
Agent trajectoryGraded step by step, not pass or fail
A run that ends green can still be wrong in the middle. Scoring the trajectory is what separates an agent benchmark from a unit test.
What you receive
Benchmarks that hold up under scrutiny, with gold solutions and grading pipelines you can re-run.
Services in this line
Terminal Agent Evaluation
Per task / per trajectory
Terminal-Bench-style task authoring, environment setup and trajectory grading.
Dockerised environments, verified solutions and step-level trajectory grading for CLI agents. Authored by engineers who have built and graded these tasks in production pipelines.
Deliverables
Task + Dockerfile
Verified gold solution
Test harness
Graded trajectories
SWE-Bench Evaluation
Per instance
Repository-level task curation and patch verification.
Test-harness validation and failure-mode analysis for software-engineering agents, with reproducible pass/fail evidence per instance.
Deliverables
Curated instances
Patch verification record
Harness validation
Failure-mode analysis
Browser Agent Evaluation
Per task / per episode
Web-task design, human demonstration recording and trajectory review.
Success-criteria grading for browser-use agents, including recorded human demonstrations as reference trajectories.
Deliverables
Task specifications
Demonstration recordings
Success-criteria grades
AI Agent Evaluation
Per trajectory
Tool-use correctness and multi-step reasoning-trace scoring.
Rubric-based trajectory scoring and safety-boundary probing for arbitrary agent frameworks, independent of your orchestration layer.
Deliverables
Trajectory scores
Tool-call correctness
Safety-boundary findings
Benchmark Programs
Per program
End-to-end benchmark construction as a managed program.
Task taxonomy → environment engineering → gold solutions → grading pipeline → leaderboard-grade reporting, run as a single accountable workstream.