Skip to content
Pilot Planner

Size the pilot before you send the email

Set your batch, your window and your acceptance bar. The planner applies the same coverage rules and staffing ratios we publish, and gives you the pod shape, the review load and a scope brief you can send straight to us.

Pilot length
2 weeks
Calibration
1 week
Response
≤3 days
Rework
Our cost
Your batch
Service line

Terminal-Bench, SWE-Bench, browser and tool-use agents

Batch volume5,000 items
Production window4 weeks
Items per contributor per day40

Yours to set. Throughput follows task complexity, so we will not guess it for you. It is recorded in the brief as your assumption.

Acceptance bar
Data posture

Our environment, engagement-specific NDA, KYC-verified pod.

Pod shape

7
Contributors

Named, NDA signed

2
Reviewers

1 per 6

1
Senior reviewers

1 per 15

3
Calibrated reserve

1.3x bench

Plus one project lead who owns the SLA. 11 people on your program, every one of them named before production opens.

QA coverage on this batch

  • Tier 1 peer review5,000100%
  • Tier 2 senior review1,60025–35%
  • Tier 3 independent audit50010%
  • Blind gold seeding2003–5%
Total review touches7,300

The bar we work to

You asked for ≥95%. We run internal QA at ≥97%, because a batch that only just clears your bar has no margin left. Allowance across this batch is 25 client-found defects, at ≤5 per 1,000.

Time to first delivery

1 week of calibration to clear the agreement gate on two consecutive rounds, then 4 weeks of production against a live dashboard. 5 weeks in total, with a mid-batch quality snapshot you see before the batch lands.

Your scope brief

Ready to send.

Nothing here is a quote. It is the shape of the work at our published rules, which is the part most vendors will not put in writing until after the call.

PILOT SCOPE DRAFT

Service line      Agent Benchmarking
Batch volume      5,000 items
Production window 4 weeks after calibration
Acceptance bar    ≥95% (Newbieget Labs works to ≥97%)
Data posture      Standard. Work in our environment under an engagement-specific NDA.

POD SHAPE AT THE PUBLISHED RATIOS
  Contributors          7
  Reviewers             2 (1 per 6)
  Senior reviewers      1 (1 per 15)
  Project lead          1
  Calibrated reserve    3 (1.3x bench)

QA COVERAGE ON THIS BATCH
  Tier 1 peer review    5,000 items (100%)
  Tier 2 senior review  1,600 items (25–35%)
  Tier 3 independent audit500 items (10%)
  Blind gold seeding    200 items (3–5%)
  Total review touches  7,300

Escaped-defect allowance: 25 client-found defects across the batch, at ≤5 per 1,000.
Timeline: 1 week calibration to clear the agreement gate, then 4 weeks of production. 5 weeks total.

My throughput assumption: 40 items per contributor per day.

STILL TO SEND
  - The rubric, or the closest draft you have
  - Written acceptance criteria and what counts as a rejection
  - Ten sample items, ideally including two you find borderline
Where the numbers come from

No figure here is invented for the planner.

Every rule the planner applies is published elsewhere on this site, so you can check the output against the framework rather than take it on trust.

Coverage comes from the framework

Tier 1 at 100%, Tier 2 on a 25 to 35 percent sample plus every flagged item, Tier 3 audited at 10 percent independent of the delivery line.

Quality framework

Ratios hold at every scale

One reviewer per six contributors, one senior reviewer per fifteen, one project lead per program, and a calibrated bench at 1.3 times deployed demand.

Staffing ratios

The gate is before production

A pod trains on gold tasks and must clear the agreement gate on two consecutive rounds before production opens. That week is in the plan, not hidden in it.

Engagement workflow

What the planner cannot tell you

Price. Throughput per contributor drives cost and it follows task complexity, which is why the planner asks you for it rather than assuming one. Send the brief with these three things and we come back with a feasibility memo, a staffing plan and a paid pilot quote, usually inside three working days.

  • 01The rubric, or the closest draft you have
  • 02Written acceptance criteria and what counts as a rejection
  • 03Ten sample items, ideally including two you find borderline