A 1,000-item coding evaluation, delivered a day early with the audit attached
The standard six-week play: pilot scoped against the client’s own acceptance bar, pod calibrated to an agreement gate before production, batch delivered with full QA evidence, converted to a dedicated pod inside two weeks.
A documented program shape with illustrative volumes, not a named client reference.
The problem
The platform had demand it could not staff without diluting its own quality story. Prior overflow vendors had returned volume on time and quality below bar, pushing verification cost back onto the client’s internal reviewers.
How it runs
- 1
Feasibility before contract
A two-page memo with the staffing plan and pilot quote went back before pricing was discussed. Acceptance criteria and rejection definitions were written into the SOW.
- 2
Calibration gate before production
Two days of rubric training on 60 gold tasks. The agreement gate cleared on round two. Three clarification questions were logged with the client and folded into rubric v1.1 rather than guessed.
- 3
Live quality, not a delivery surprise
Tier 1 review at 100%, Tier 2 at 30%, blind gold seeds at 4%. A mid-batch checkpoint was shared at item 500. One contributor fell below the gold-task floor at day three and was rotated out. The bench replacement calibrated the same day.
- 4
Evidence with the delivery
The package shipped with a per-item review trail, agreement data, gold-task results and a defect taxonomy. Two Category-A defects surfaced in the client spot-check were reworked within 48 hours at our cost.
Outcome
- Acceptance
- 96.8%
- Client spot-check, against a 95% bar
- Schedule
- −1 day
- Delivered ahead of committed date
- Rework
- 48h
- Two Cat-A defects, at our cost
- Expansion
- 3 months
- Dedicated pod signed within two weeks
What makes it stick
The retro produced an unsolicited memo on three systematic model-failure patterns observed across the batch. That memo, not the acceptance rate, is what opened the pod conversation.
Other patterns
Terminal-agent benchmark built end to end, from task taxonomy to grading pipeline
Agentic evaluation is where crowd models break. Task authoring, Dockerised environment engineering, verified gold solutions and step-level trajectory grading, run as one accountable workstream.
ReadAI Data OperationsAuditing someone else’s output, the highest-trust way to start
A review and audit layer operated over a third party’s delivery: the client’s own crowd, or another vendor. We publish the metrics either way, including the ones that are inconvenient.
Read