Skip to content
Glossary

The vocabulary of AI evaluation

Plain definitions of the terms that appear in evaluation contracts, quality reports and vendor conversations. Written by the people who use them on delivery, not lifted from a textbook.

Terms
25
Groups
4
Written by
Delivery
Updated
Aug 2026
Quality

Acceptance rate

The share of delivered items a client accepts without requesting rework. It is the headline quality metric in most evaluation contracts because it measures the thing the buyer actually experiences.

An acceptance rate is only meaningful alongside the acceptance criteria it was measured against. A 99% rate on a loose rubric says less than a 95% rate on a strict one.

How we apply it

Agreement gate

A checkpoint that blocks production until a team of raters demonstrates they interpret the rubric the same way, usually by clearing an inter-rater agreement threshold on consecutive calibration rounds.

Placing the gate before production is the point. Discovering that raters disagreed after a batch ships makes the whole batch suspect.

How we apply it

Blind gold seeding

Gold tasks

Mixing known-answer items into a live work queue so they are indistinguishable from real tasks. Scores on those items measure rater accuracy continuously without anyone being able to see the test coming.

Typically 3 to 5 percent of a queue. If raters can identify the gold items, the measurement is worthless.

How we apply it

Calibration

The process of training a group of raters to apply a rubric consistently, usually by having everyone score the same items and arguing the disagreements to resolution.

Calibration is recurring, not one-off. Rater judgment drifts, and weekly recalibration is what keeps a long programme stable.

How we apply it

Cohen's kappa

κ

A statistic measuring agreement between two raters, corrected for the agreement you would expect by chance. Ranges from below 0 to 1, where 1 is perfect agreement.

Chance correction is why it beats raw percent agreement. Two raters who both pick "pass" 90% of the time will agree 82% of the time knowing nothing at all.

How we apply it

Defect taxonomy

A structured classification of the ways an item can be wrong, used to tag failures consistently so that patterns become visible across a batch.

Without one, a quality report says how many items failed. With one, it says why, which is the part that changes the rubric.

How we apply it

Escaped defect

A defect the client finds after delivery, having passed every internal review. Usually expressed as a rate per thousand delivered items.

The most honest quality metric a vendor can publish, because it counts only the failures their own QA missed.

How we apply it

First-pass yield

The share of items that clear the first review stage unchanged. It measures how well the production layer is performing before any correction.

A rising acceptance rate with a falling first-pass yield means the reviewers are absorbing a problem rather than the rubric fixing it.

How we apply it

Inter-rater agreement

IAA

The degree to which independent raters assign the same label to the same item. Measured with statistics such as Cohen’s kappa or Krippendorff’s alpha rather than raw percent agreement.

It functions as a ceiling on your reward model. If trained raters disagree on a third of comparisons, no volume of additional data recovers the lost signal.

How we apply it

Krippendorff's alpha

α

An agreement statistic that handles any number of raters, missing data and different measurement scales, which makes it more flexible than Cohen’s kappa on real annotation projects.

How we apply it

Three-tier review

A QA structure where every item is peer reviewed, a sample plus all flagged items go to a senior reviewer, and a smaller sample is audited independently of the delivery line.

Independence of the third tier is what makes it meaningful. An audit run by the team that produced the work measures compliance, not quality.

How we apply it
Evaluation

Agentic evaluation

Assessing an AI agent on a multi-step task rather than a single response. The grader judges the whole trajectory: which tools were called, in what order, and whether the final state is correct.

A run that ends in the right answer can still be wrong in the middle. Scoring only the endpoint misses reward hacking, wasted steps and lucky recoveries.

How we apply it

LLM evaluation

Scoring language-model outputs against defined criteria such as factual accuracy, instruction following, reasoning quality and safety, usually with a rubric and trained human raters.

How we apply it

Rubric

The written scoring guide raters apply: the dimensions being judged, the scale for each, and anchored examples of what each score looks like.

Ambiguity in a rubric is the single largest cause of failed evaluation programmes. Every clarification should be versioned and logged rather than answered in a chat thread.

How we apply it

SWE-Bench

A benchmark that evaluates whether a model can resolve real GitHub issues by producing a patch that passes the repository’s own tests. Instances are drawn from actual open-source projects.

Curating an instance is engineering work: the environment must be reproducible and the tests must genuinely verify the fix rather than pass trivially.

How we apply it

Terminal-Bench

A benchmark for agents operating in a terminal, where tasks are defined in containerised environments and scored on whether the agent reaches the correct end state.

Authoring a task means building the environment, the solution and the verification, which is closer to test engineering than to annotation.

How we apply it

Trajectory grading

Scoring each step an agent takes through a task rather than only the final result, so that a correct outcome reached by a flawed path is still marked down.

How we apply it
Human data

Expert bench

A pool of credential-verified professionals, typically in medicine, law, finance or specific STEM fields, held on call for evaluation work that requires domain qualification.

Bench capacity is usually held above deployed demand so a programme can be staffed without a fresh recruitment cycle.

How we apply it

Preference ranking

Pairwise comparison

Having a human choose which of two or more model responses is better. The resulting comparisons train the reward model that guides reinforcement learning from human feedback.

How we apply it

Red teaming

Deliberately attempting to make a model produce unsafe, policy-violating or otherwise undesirable output, in order to find failure modes before deployment.

How we apply it

RLHF

Reinforcement learning from human feedback

A training method where human preferences between model outputs are used to train a reward model, which then guides reinforcement learning to align the model with those preferences.

The quality ceiling of RLHF is set by the consistency of the human preferences underneath it, which is why agreement measurement matters more than raw volume.

How we apply it

SFT

Supervised fine-tuning

Training a model on curated input and output pairs written or approved by humans, usually as the stage before reinforcement learning from human feedback.

How we apply it

Synthetic data verification

Human checking of machine-generated training data to confirm it is correct, diverse and free of the artefacts that cause model collapse when synthetic data is trained on unchecked.

How we apply it
Delivery

AI data operations

AI ops

The managed running of data workflows that feed AI development: evaluation pipelines, RLHF programmes, benchmark harnesses and the quality systems around them.

Deliberately broader than data labelling. Operations remains a durable service even as raw annotation commoditises.

How we apply it

Managed pod

A named, fixed team assigned to one client programme, typically contributors plus reviewers plus a lead who owns the service-level agreement.

The opposite of a marketplace draw, where each batch may be worked by different people who never build context on the rubric.

How we apply it

These are the terms that end up in a statement of work. If one of them is doing heavy lifting in a contract you are about to sign, it is worth agreeing on the definition first.

Talk it through