RLHF & Human Feedback Per item / per pod-month Preference data, critique writing and demonstration data from trained rater pods.
Rubric-anchored reward-model data produced by named pods under weekly calibration, with per-dimension agreement reported per batch.
Deliverables
Preference pairs Written critiques SFT demonstrations Agreement statistics
Prompt Engineering & Red-Team Sets Per prompt set Adversarial prompt design and evaluation-prompt curation.
Instruction datasets and red-team prompt sets built against an explicit coverage matrix, so gaps are visible rather than assumed.
Deliverables
Prompt sets Coverage matrix Adversarial variants Curation rationale
Synthetic Dataset Generation Per dataset Model-assisted generation with human seed curation and verification.
Generation pipelines with human-curated seeds, deduplication and a verification layer. Delivered with a quality-audit sample rather than a raw dump.
Deliverables
Generated dataset Seed curation log Dedup report Audit sample
Expert Data Annotation Per item Expert text and code annotation, not commodity labelling.
Entity labelling, classification, reasoning-step tagging and error-taxonomy application performed by domain-matched specialists.
Deliverables
Annotated corpus Taxonomy definitions Edge-case register