Skip to content
Work

Engagement patterns, documented in full

Newbieget Labs is an early-stage firm, so we publish the shape of our programs rather than borrowed logos. Volumes below are illustrative. Named client studies publish once engagements clear for publication.

Agent trajectoryGraded step by step, not pass or fail

A run that ends green can still be wrong in the middle. Scoring the trajectory is what separates an agent benchmark from a unit test.

01Model Evaluation6 weeks · pilot to recurringPattern

A 1,000-item coding evaluation, delivered a day early with the audit attached

The standard six-week play: pilot scoped against the client’s own acceptance bar, pod calibrated to an agreement gate before production, batch delivered with full QA evidence, converted to a dedicated pod inside two weeks.

Read the pattern
Acceptance
96.8%
Client spot-check, against a 95% bar
Schedule
−1 day
Delivered ahead of committed date
Rework
48h
Two Cat-A defects, at our cost
Expansion
3 months
Dedicated pod signed within two weeks
02Agent Benchmarking10 weeks · programPattern

Terminal-agent benchmark built end to end, from task taxonomy to grading pipeline

Agentic evaluation is where crowd models break. Task authoring, Dockerised environment engineering, verified gold solutions and step-level trajectory grading, run as one accountable workstream.

Read the pattern
Reproducibility
100%
Gold solution passes from clean container
Grader agreement
κ ≥ 0.7
Tracked per rubric dimension
Shortcut audit
Every task
Adversarially reviewed before inclusion
Handover
Full
Harness runs in your CI, not ours