Skip to content
Line B · Capabilities

The evaluation work that needs real engineers

Agentic evaluation is where crowd models break. Task authoring, environment engineering and trajectory grading are engineering disciplines, and they are the founding team’s specialty.

Line
Line B
Services
5
Acceptance
≥95%
Agreement
κ ≥ 0.7
Agent trajectoryGraded step by step, not pass or fail

A run that ends green can still be wrong in the middle. Scoring the trajectory is what separates an agent benchmark from a unit test.

What you receive

Benchmarks that hold up under scrutiny, with gold solutions and grading pipelines you can re-run.

Services in this line

Terminal Agent Evaluation

Per task / per trajectory

Terminal-Bench-style task authoring, environment setup and trajectory grading.

Dockerised environments, verified solutions and step-level trajectory grading for CLI agents. Authored by engineers who have built and graded these tasks in production pipelines.

Deliverables

  • Task + Dockerfile
  • Verified gold solution
  • Test harness
  • Graded trajectories

SWE-Bench Evaluation

Per instance

Repository-level task curation and patch verification.

Test-harness validation and failure-mode analysis for software-engineering agents, with reproducible pass/fail evidence per instance.

Deliverables

  • Curated instances
  • Patch verification record
  • Harness validation
  • Failure-mode analysis

Browser Agent Evaluation

Per task / per episode

Web-task design, human demonstration recording and trajectory review.

Success-criteria grading for browser-use agents, including recorded human demonstrations as reference trajectories.

Deliverables

  • Task specifications
  • Demonstration recordings
  • Success-criteria grades

AI Agent Evaluation

Per trajectory

Tool-use correctness and multi-step reasoning-trace scoring.

Rubric-based trajectory scoring and safety-boundary probing for arbitrary agent frameworks, independent of your orchestration layer.

Deliverables

  • Trajectory scores
  • Tool-call correctness
  • Safety-boundary findings

Benchmark Programs

Per program

End-to-end benchmark construction as a managed program.

Task taxonomy → environment engineering → gold solutions → grading pipeline → leaderboard-grade reporting, run as a single accountable workstream.

Deliverables

  • Task taxonomy
  • Environment suite
  • Grading pipeline
  • Reporting harness