Seventeen services, four lines, one delivery standard
Services are productised, not invented per engagement. Each line carries a named owner, a versioned rubric library and a calibration protocol that runs before production begins.
21/25
Item score
Mandatory dimensions all cleared
Any mandatory dimension below 3 caps the item regardless of the rest. That rule is written into the SOW, not left to reviewer judgment.
Model Evaluation
Judgment your researchers would sign off on.
Rubric-anchored scoring of model behaviour, executed by raters who understand the domain and calibrated against your own acceptance bar before a single production item is scored.
Scored datasets, calibration reports and disagreement analysis, not raw opinion.
Line detailLLM Evaluation
Per response / per comparisonRubric-based scoring across helpfulness, accuracy, safety and instruction-following.
Code Generation Evaluation
Per solutionFunctional-correctness review and code-quality rubric scoring of model-written code.
Expert Code Review
Per PR / per fileLine-level review of model or human code for training signal.
Human Preference Ranking
Per pair / per setPairwise and k-way ranking with trained, calibrated raters.
Human Verification
Per itemHuman sign-off layers over synthetic or model-generated data.
Agent Benchmarking
The evaluation work that needs real engineers.
Agentic evaluation is where crowd models break. Task authoring, environment engineering and trajectory grading are engineering disciplines, and they are the founding team’s specialty.
Benchmarks that hold up under scrutiny, with gold solutions and grading pipelines you can re-run.
Line detailTerminal Agent Evaluation
Per task / per trajectoryTerminal-Bench-style task authoring, environment setup and trajectory grading.
SWE-Bench Evaluation
Per instanceRepository-level task curation and patch verification.
Browser Agent Evaluation
Per task / per episodeWeb-task design, human demonstration recording and trajectory review.
AI Agent Evaluation
Per trajectoryTool-use correctness and multi-step reasoning-trace scoring.
Benchmark Programs
Per programEnd-to-end benchmark construction as a managed program.
Human Data & RLHF
Preference data from raters who are actually calibrated.
Reward-model quality is bounded by rater agreement. We run pods against a versioned rubric with weekly calibration, blind gold seeding and a documented escalation path for every ambiguity.
Training data with the agreement statistics attached, so you know what you are training on.
Line detailRLHF & Human Feedback
Per item / per pod-monthPreference data, critique writing and demonstration data from trained rater pods.
Prompt Engineering & Red-Team Sets
Per prompt setAdversarial prompt design and evaluation-prompt curation.
Synthetic Dataset Generation
Per datasetModel-assisted generation with human seed curation and verification.
Expert Data Annotation
Per itemExpert text and code annotation, not commodity labelling.
AI Data Operations
The operating layer between your spec and a finished dataset.
Recruiting, vetting, training, scheduling, reviewing, auditing and reporting. Run as one accountable system, inside your tooling wherever your policy requires it.
A delivery organisation you can point at a problem, without building one.
Line detailQA Pipelines
Retainer / per itemMulti-tier review systems designed and operated over any human-data workflow.
Enterprise Workforce Management
Per seat-monthClient-dedicated contributor programs run under your brand and tooling.
Expert Contributor Network
Per expert-hourOn-demand access to a vetted specialist bench.