Insights
Notes from inside the evaluation pipeline
Written by the people who run the batches. Methodology, failure modes and the operational detail that decides whether an evaluation programme produces signal or noise.
Methodology14 July 2026 · 6 min read
Inter-rater agreement is the ceiling on your reward model
If two trained raters disagree on a third of your comparisons, no amount of data volume recovers the signal. Agreement is not a quality report line item, it is the bound.
Read itAgent benchmarking7 min read
Why agentic evaluation breaks the crowd model
Scoring a model response is a rating task. Grading an agent trajectory is an engineering task. The economics that make crowds work stop applying at exactly that boundary.
21 July 2026Procurement5 min read
A buyer’s checklist for evaluation capacity
Eleven questions that separate a vendor running a quality system from one running a spreadsheet. Most can be answered in a first call.
27 July 2026