Skip to content
AI Operations · AI Data Services

Building the humanintelligence layerbehind frontier AI

Newbieget Labs is the managed execution partner for AI labs and data platforms. Complete, quality-assured evaluation and training-data projects. Delivered under SLA, with the audit attached.

Delivery pipeline
1,000 items processed

Produce

1,000

Contributors work the rubric

Named pod

Tier 1

1,000

Every item peer reviewed

100%

Tier 2

310

Senior review plus all flags

25–35%

Audit

100

Independent of delivery

10%

Deliver

970

Review trail attached

+ evidence

Rework loop. Items failing Tier 1 return to the contributor with tagged findings. Batches predicted to miss your bar are reworked before delivery, at our cost.

≥95%

Acceptance rate SLA

Fee-at-risk option available

κ ≥ 0.7

Inter-rater agreement gate

Cleared before production starts

100%

First-pass peer review

Every item reviewed at least once

≤5

Escaped defects / 1,000

Client-found, tracked per batch

Built for the teams defining frontier AI

Frontier research labsAI data platformsEvaluation toolingEnterprise AI teamsBenchmark consortiaSafety & alignment teams

Client logos withheld under NDA · category placeholders shown

Operating standardCommitments carried into every engagement, not historical averages
95%

Acceptance rate

Written into every SOW

κ ≥ 0.7

Agreement gate

Cleared before production

100%

First-pass review

Every item, no sampling

~10%

Applicant pass rate

Screen, trial, calibration

1 : 6

Reviewer ratio

Invariant at any scale

5

Escaped defects / 1k

Client-found, per batch

Azure marks the system: acceptance, agreement, review coverage. Violet marks the people who produce it.

The thesis

Frontier AI isbottlenecked bytrustworthy judgment.

The companies supplying that judgment are constrained by quality, not demand. Newbieget Labs wins by making quality the product. Measured, audited and delivered as evidence with every batch.

Frontier AI development is bottlenecked not by compute or capital, but by trustworthy human judgment at scale.

Where it breaks todayCost to the buyer
  • 01

    Freelancer variance

    Individual contributors on marketplaces vary wildly in skill, effort and honesty.

    20–40% of budget spent on redundancy and rework

  • 02

    Churn and retraining

    Gig contributors leave mid-project. Calibration is lost and ramp-up repeats endlessly.

    Schedule slip and quality regression mid-dataset

  • 03

    QA burden stays with the buyer

    Marketplaces deliver raw output, not verified output, so labs staff large internal review teams.

    Researcher time diverted to reviewing

  • 04

    Security exposure

    Thousands of anonymous individuals touching confidential model data. NDA enforcement is impractical.

    Leak risk on unreleased capabilities and benchmarks

  • 05

    Expertise scarcity

    Agentic evaluations and repository-grade review need real engineers, not raters.

    Roadmap delay on evaluation-gated releases

The cost of variance

Unmanaged judgmentis not cheaper.

20–40%

of budget spent on redundancy and rework, because freelancer variance means individual contributors on marketplaces vary wildly in skill, effort and honesty.

  • Churn and retraining

    Gig contributors leave mid-project. Calibration is lost and ramp-up repeats endlessly.

    Schedule slip and quality regression mid-dataset

  • QA burden stays with the buyer

    Marketplaces deliver raw output, not verified output, so labs staff large internal review teams.

    Researcher time diverted to reviewing

  • Security exposure

    Thousands of anonymous individuals touching confidential model data. NDA enforcement is impractical.

    Leak risk on unreleased capabilities and benchmarks

  • Expertise scarcity

    Agentic evaluations and repository-grade review need real engineers, not raters.

    Roadmap delay on evaluation-gated releases

None of this is a sourcing problem. It is a management problem, and it is the one Newbieget Labs is built to absorb.

Model Evaluation

Judgment your researcherswould sign off on.

Rubric-anchored scoring of model behaviour, executed by raters who understand the domain and calibrated against your own acceptance bar before a single production item is scored.

  • LLM evaluation, code generation review and expert code review
  • Human preference ranking with agreement tracked per dimension
  • Versioned rubrics with every clarification logged, never guessed
Rubric scoringAnchored, not impressionistic

21/25

Item score

Mandatory dimensions all cleared

Any mandatory dimension below 3 caps the item regardless of the rest. That rule is written into the SOW, not left to reviewer judgment.

Agent Benchmarking

The evaluation work thatneeds real engineers.

Agentic evaluation is where crowd models break. Task authoring, environment engineering and trajectory grading are engineering disciplines, and they are the founding team's specialty.

  • Terminal-Bench style tasks with Dockerised, reproducible environments
  • SWE-Bench instance curation with verified patches and harness validation
  • Benchmark programs handed over complete, running in your CI
Agent trajectoryGraded step by step, not pass or fail

A run that ends green can still be wrong in the middle. Scoring the trajectory is what separates an agent benchmark from a unit test.

Human Data & RLHF

Preference data from raterswho are actually calibrated.

Reward-model quality is bounded by rater agreement. We run pods against a versioned rubric with weekly calibration, blind gold seeding and a documented escalation path for every ambiguity.

  • Preference pairs, written critiques and SFT demonstration data
  • Adversarial prompt sets built against an explicit coverage matrix
  • Synthetic pipelines with human seed curation and verification layers
A named podRatios hold at every scale

1.3×

Calibrated bench

Held against deployed demand

100%

KYC verified

Before task access is granted

AI Data Operations

The operating layer betweenyour spec and a finished dataset.

Recruiting, vetting, training, scheduling, reviewing, auditing and reporting, run as one accountable system, inside your tooling wherever your policy requires it.

  • QA pipelines operated over our work or a third party’s
  • Client-dedicated workforce programs under your brand and tooling
  • On-demand expert bench across medicine, law, finance and STEM
The delivery packageThe work and the proof, together
You never have to ask what you received
The gap we fill

Between the marketplaceand the in-house team.

Marketplaces are scalable but unaccountable. In-house teams are accountable but unscalable. Between them sits the managed services firm. Vetted teams, owned QA, contractual accountability, elastic capacity. In AI data, that layer is still thin.

Individuals suppliedOutcomes delivered
Expert tierGeneralist tier

Expert marketplaces

Vetted individuals, client-managed

Newbieget Labs

Managed pods, owned QA, SLA-backed

Crowd platforms

Volume, high variance

Generalist managed ops

Process depth, not AI-native

We target the expert-tier, outcomes-delivered quadrant, where willingness to pay is highest and crowd models cannot follow.

Boundaries

  • Not a marketplace

    No self-serve gig pool. Every contributor is sourced, tested, trained, calibrated and managed under a named lead.

  • Not a body shop

    We sell accountable deliverables under SLAs, not resumes at an hourly markup. Scope, acceptance criteria and rework terms are contractual.

  • Not a model company

    We build no competing IP. There is no scenario where your unreleased capabilities become our product.

Quality Framework

QA is not our process.It is our product.

Marketplaces deliver raw output and leave verification with the buyer. We deliver verified output with the evidence attached.

Review coverage

  • Tier 3. AuditHead of Quality

    Random 10% plus risk-weighted sample

  • Tier 2. Senior reviewSenior Reviewers

    25–35% sample plus all flagged items

  • Tier 1. Peer reviewReviewers

    100% of contributor output

  • ProductionContributors

    Self-check against rubric checklist before submit

MetricTargetMature
  • Acceptance rate

    Share of delivered items accepted by the client without rework.

    ≥95%≥98%
  • First-pass yield

    Share of contributor items passing Tier 1 unchanged.

    ≥85%≥90%
  • Inter-rater agreement

    Cohen’s κ / Krippendorff’s α on the double-rated sample.

    ≥0.7≥0.75
  • Escaped-defect rate

    Client-found defects per 1,000 delivered items.

    ≤5≤2
  • Gold-task accuracy

    Contributor score on blind, seeded known-answer tasks.

    ≥90%≥93%
  • Review accuracy

    Tier 1 decisions overturned at Tier 2 or Tier 3.

    ≤8%≤5%

Rule of the house

No batch leaves Newbieget Labs below the client’s acceptance bar. If internal QA predicts a miss, the batch is reworked before delivery. Schedule slips are negotiable. Quality misses are not.

Engagement Workflow

Eight stages.Five gates. One standard.

Every engagement runs the same pipeline. Standardisation is what lets a small leadership team run many programs at once without quality drifting between them.

  1. 01

    Qualify

    Fit, scope, data sensitivity, feasibility.

    We check the work against our service lines and against your data-handling constraints before anyone talks price. Work we cannot staff to the acceptance bar is declined at this stage rather than discovered at delivery.

    • Feasibility read
    • Data-handling assessment
  2. 02

    Scope & pilot

    Acceptance criteria, rubric draft, pilot pricing.

    A two-page feasibility memo with a staffing plan and a paid pilot quote. Acceptance criteria and rejection definitions are written down before signature. Ambiguity here is the single largest cause of failed engagements.

    • Feasibility memo
    • Pilot SOW
    • Rubric v1 draft
  3. 03

    Staff

    Pod selection from the bench, skills-matched, NDA’d.

    A named pod. Contributors, reviewers and a lead. Is drawn from the calibrated bench and matched on domain. Every member signs an engagement-specific NDA before task access is granted.

    • Named pod roster
    • NDA register
  4. 04

    Calibrate

    Gold tasks, rubric training, IAA gate before production.

    The pod trains on gold tasks and must clear the agreement gate on two consecutive calibration rounds before production begins. Clarification questions are logged and folded into the rubric, versioned.

    • Calibration report
    • Rubric v1.1
    • Clarification log
  5. 05

    Produce

    Daily quotas, live dashboards, blocker escalation.

    Production runs against daily targets with a live quality dashboard shared with you. Contributors below the gold-task floor are moved to recalibration and replaced from the bench the same day.

    • Live dashboard
    • Mid-batch quality snapshot
  6. 06

    QA pipeline

    Peer review → senior review → independent audit.

    Every item is peer-reviewed. A senior sample and all flagged items go to Tier 2. Ten percent is audited independently of the delivery line. Internal bar sits two points above yours.

    • Per-item review trail
    • Audit sample results
  7. 07

    Deliver

    Batch plus QA evidence, audit sample and metrics.

    The delivery package contains the work and the proof: per-item review trail, agreement data, gold-task results and defect taxonomy. You never have to ask what you received.

    • Delivery package
    • Quality report
    • Defect taxonomy
  8. 08

    Review & expand

    Retro, quality report, systematic-failure insight memo.

    A retro with your ops lead, rubric proposals for the next version, and an unsolicited memo on the systematic model-failure patterns we observed. The insight memo is the part clients remember.

    • Retro notes
    • Insight memo
    • Expansion proposal
The Human Intelligence Layer

Trained people, wrapped in a systemthat makes excellence repeatable.

Roughly one in ten applicants reaches our bench. The ones who do are paid for their training, calibrated weekly against a versioned rubric, scored on blind gold tasks, and given a career ladder no gig platform offers. That is not an HR programme. It is the reason the numbers hold.

Contributor funnel

  1. Source100%

    Campus networks, referrals, engineering and Kaggle communities, IIT/NIT/IIIT alumni groups.

  2. Screen~30%

    Structured application plus an async, domain-matched skills test.

  3. Assess~40%

    A paid trial batch on gold tasks and a live rubric interview.

  4. Onboard~90%

    KYC, NDA and a paid calibration week before any client work.

  5. Bench~10% e2e

    Skills-tagged, score-tracked and deployable. Roughly one in ten applicants.

~10%

End-to-end applicant pass rate

Screen, paid trial, live rubric interview, calibration week

1 : 6

Reviewer to contributor ratio

Invariant at every scale we operate

3–5%

Blind gold seeding per queue

Indistinguishable from live work

Weekly

Pod calibration sessions

Disagreement argued to a versioned rubric change

Enterprise Security

Most vendors fail procurement,not the pilot.

Every engagement is treated as if it covers unreleased model data, because it usually does. The compliance posture is built before it is asked for, so a security review is a document exchange rather than a project.

  • Two-layer NDA chain, signed before any task access is granted
  • Work performed inside your environment by default, so data never leaves
  • KYC-verified contributors, named accounts, same-day deprovisioning
Compliance postureStated, not implied

Nothing here is gated behind a sales call. Certifications in progress are labelled as in progress, because a procurement reviewer will find out either way.

Questions

The ones that decide a first call.

Including the answers that are inconvenient for us, because a buyer finds those out anyway and would rather find them out now.

All questions

With a paid pilot, usually two weeks. You send a rubric and the acceptance bar you already use. We return a two-page feasibility memo, a staffing plan and a fixed quote. If we miss the bar, the rework is at our cost.

It does not ship. Internal QA runs two points above your bar, so a predicted miss is reworked before delivery rather than after. If something escapes and you find it, the rework is at our cost within the window written into the SOW.

Four models: fixed-scope project delivery, a monthly retainer for a dedicated pod, per-item or retainer pricing for QA over someone else’s output, and project pricing for benchmark engineering. Rates depend on domain, expertise tier and volume, and are quoted per engagement.

Work is performed inside your environment by default, so the data never leaves your systems. Where we do host, it is encrypted at rest and in transit under least-privilege role-based access with 90-day access reviews and deletion certificates at project close.

Yes, and it is often the best way to start. We publish the metrics either way, including findings that are commercially inconvenient for us.

Not yet. The security policy pack, incident-response plan and pre-drafted vendor questionnaire responses exist today, and SOC 2 Type I is on the roadmap with Type II following. We would rather state the position plainly than imply a certification we do not hold.

What buyers sayPlaceholder quotes. Attributions publish once clients approve them.
  • The delivery arrived with a per-item review trail and an audit sample. It was the first vendor batch we did not have to re-verify ourselves.

    Head of Data Operations

    AI data platform

  • They logged three rubric ambiguities instead of guessing. That single behaviour told us more about the team than the acceptance rate did.

    Evaluation Lead

    Frontier research team

  • We asked them to audit another vendor’s output. They reported findings that were commercially inconvenient for them. We expanded the contract.

    VP, Vendor Management

    Prime platform