Skip to content
Line A · Capabilities

Judgment your researchers would sign off on

Rubric-anchored scoring of model behaviour, executed by raters who understand the domain and calibrated against your own acceptance bar before a single production item is scored.

Line
Line A
Services
5
Acceptance
≥95%
Agreement
κ ≥ 0.7
Rubric scoringAnchored, not impressionistic

21/25

Item score

Mandatory dimensions all cleared

Any mandatory dimension below 3 caps the item regardless of the rest. That rule is written into the SOW, not left to reviewer judgment.

What you receive

Scored datasets, calibration reports and disagreement analysis, not raw opinion.

Services in this line

LLM Evaluation

Per response / per comparison

Rubric-based scoring across helpfulness, accuracy, safety and instruction-following.

Side-by-side comparisons and regression evaluations between checkpoints, scored against a versioned rubric with logged clarifications.

Deliverables

  • Scored dataset
  • Calibration report
  • Disagreement analysis
  • Rubric version log

Code Generation Evaluation

Per solution

Functional-correctness review and code-quality rubric scoring of model-written code.

Execution-verified where environments permit, across languages. Compile/run status, test pass counts and runtime errors recorded verbatim. Never scored by reading alone when execution is available.

Deliverables

  • Verdict + evidence per solution
  • Defect taxonomy tags
  • Execution logs
  • Failure-mode summary

Expert Code Review

Per PR / per file

Line-level review of model or human code for training signal.

Bugs, security issues, style violations and spec deviations, each tagged to a defect taxonomy code with a written justification a reviewer can verify without redoing the task.

Deliverables

  • Structured findings
  • Severity classification
  • Referenced justifications

Human Preference Ranking

Per pair / per set

Pairwise and k-way ranking with trained, calibrated raters.

Inter-rater agreement tracked continuously at Cohen’s κ ≥ 0.7 on defined dimensions, with weekly calibration sessions resolving disagreement to rubric updates.

Deliverables

  • Ranked pairs or sets
  • Per-dimension IAA
  • Weekly calibration notes

Human Verification

Per item

Human sign-off layers over synthetic or model-generated data.

Factuality checks, harmful-content screens and gold-set validation applied as an independent pass over machine-produced output.

Deliverables

  • Verification verdicts
  • Flagged-item register
  • Gold-set validation results