Skip to content
Capabilities

Seventeen services, four lines, one delivery standard

Services are productised, not invented per engagement. Each line carries a named owner, a versioned rubric library and a calibration protocol that runs before production begins.

Rubric scoringAnchored, not impressionistic

21/25

Item score

Mandatory dimensions all cleared

Any mandatory dimension below 3 caps the item regardless of the rest. That rule is written into the SOW, not left to reviewer judgment.

Line B5 services

Agent Benchmarking

The evaluation work that needs real engineers.

Agentic evaluation is where crowd models break. Task authoring, environment engineering and trajectory grading are engineering disciplines, and they are the founding team’s specialty.

Benchmarks that hold up under scrutiny, with gold solutions and grading pipelines you can re-run.

Line detail
Line C4 services

Human Data & RLHF

Preference data from raters who are actually calibrated.

Reward-model quality is bounded by rater agreement. We run pods against a versioned rubric with weekly calibration, blind gold seeding and a documented escalation path for every ambiguity.

Training data with the agreement statistics attached, so you know what you are training on.

Line detail
Line D3 services

AI Data Operations

The operating layer between your spec and a finished dataset.

Recruiting, vetting, training, scheduling, reviewing, auditing and reporting. Run as one accountable system, inside your tooling wherever your policy requires it.

A delivery organisation you can point at a problem, without building one.

Line detail