Acceptance rate
The share of delivered items a client accepts without requesting rework. It is the headline quality metric in most evaluation contracts because it measures the thing the buyer actually experiences.
An acceptance rate is only meaningful alongside the acceptance criteria it was measured against. A 99% rate on a loose rubric says less than a 95% rate on a strict one.
How we apply itAgreement gate
A checkpoint that blocks production until a team of raters demonstrates they interpret the rubric the same way, usually by clearing an inter-rater agreement threshold on consecutive calibration rounds.
Placing the gate before production is the point. Discovering that raters disagreed after a batch ships makes the whole batch suspect.
How we apply itBlind gold seeding
Gold tasksMixing known-answer items into a live work queue so they are indistinguishable from real tasks. Scores on those items measure rater accuracy continuously without anyone being able to see the test coming.
Typically 3 to 5 percent of a queue. If raters can identify the gold items, the measurement is worthless.
How we apply itCalibration
The process of training a group of raters to apply a rubric consistently, usually by having everyone score the same items and arguing the disagreements to resolution.
Calibration is recurring, not one-off. Rater judgment drifts, and weekly recalibration is what keeps a long programme stable.
How we apply itCohen's kappa
κA statistic measuring agreement between two raters, corrected for the agreement you would expect by chance. Ranges from below 0 to 1, where 1 is perfect agreement.
Chance correction is why it beats raw percent agreement. Two raters who both pick "pass" 90% of the time will agree 82% of the time knowing nothing at all.
How we apply itDefect taxonomy
A structured classification of the ways an item can be wrong, used to tag failures consistently so that patterns become visible across a batch.
Without one, a quality report says how many items failed. With one, it says why, which is the part that changes the rubric.
How we apply itEscaped defect
A defect the client finds after delivery, having passed every internal review. Usually expressed as a rate per thousand delivered items.
The most honest quality metric a vendor can publish, because it counts only the failures their own QA missed.
How we apply itFirst-pass yield
The share of items that clear the first review stage unchanged. It measures how well the production layer is performing before any correction.
A rising acceptance rate with a falling first-pass yield means the reviewers are absorbing a problem rather than the rubric fixing it.
How we apply itInter-rater agreement
IAAThe degree to which independent raters assign the same label to the same item. Measured with statistics such as Cohen’s kappa or Krippendorff’s alpha rather than raw percent agreement.
It functions as a ceiling on your reward model. If trained raters disagree on a third of comparisons, no volume of additional data recovers the lost signal.
How we apply itKrippendorff's alpha
αAn agreement statistic that handles any number of raters, missing data and different measurement scales, which makes it more flexible than Cohen’s kappa on real annotation projects.
How we apply itThree-tier review
A QA structure where every item is peer reviewed, a sample plus all flagged items go to a senior reviewer, and a smaller sample is audited independently of the delivery line.
Independence of the third tier is what makes it meaningful. An audit run by the team that produced the work measures compliance, not quality.
How we apply it