Rubric-anchored scoring of model behaviour, executed by raters who understand the domain and calibrated against your own acceptance bar before a single production item is scored.
Any mandatory dimension below 3 caps the item regardless of the rest. That rule is written into the SOW, not left to reviewer judgment.
What you receive
Scored datasets, calibration reports and disagreement analysis, not raw opinion.
Services in this line
LLM Evaluation
Per response / per comparison
Rubric-based scoring across helpfulness, accuracy, safety and instruction-following.
Side-by-side comparisons and regression evaluations between checkpoints, scored against a versioned rubric with logged clarifications.
Deliverables
Scored dataset
Calibration report
Disagreement analysis
Rubric version log
Code Generation Evaluation
Per solution
Functional-correctness review and code-quality rubric scoring of model-written code.
Execution-verified where environments permit, across languages. Compile/run status, test pass counts and runtime errors recorded verbatim. Never scored by reading alone when execution is available.
Deliverables
Verdict + evidence per solution
Defect taxonomy tags
Execution logs
Failure-mode summary
Expert Code Review
Per PR / per file
Line-level review of model or human code for training signal.
Bugs, security issues, style violations and spec deviations, each tagged to a defect taxonomy code with a written justification a reviewer can verify without redoing the task.
Deliverables
Structured findings
Severity classification
Referenced justifications
Human Preference Ranking
Per pair / per set
Pairwise and k-way ranking with trained, calibrated raters.
Inter-rater agreement tracked continuously at Cohen’s κ ≥ 0.7 on defined dimensions, with weekly calibration sessions resolving disagreement to rubric updates.
Deliverables
Ranked pairs or sets
Per-dimension IAA
Weekly calibration notes
Human Verification
Per item
Human sign-off layers over synthetic or model-generated data.
Factuality checks, harmful-content screens and gold-set validation applied as an independent pass over machine-produced output.