Human data for frontier AI
Five kinds of human data.
Three ways to buy.
Speech, expert judgment, private records, enterprise data and coding traces, from real people. License it, commission it, or have experts label it.
Pick the data. Pick how you buy it. See a sample before anything scales.
Every order runs the same way, whether it is forty cardiologists or ten thousand hours of conversation.
- Agree the spec. Task, target people, format, volume, usage rights and acceptance criteria.
- Review a sample. Real contributions, so the spec stops being abstract.
- Scale. Batch delivery with a quality report per contributor.
Five data lines
Choose the data
your model needs.
Speech and conversation
Two-speaker and group recordings with overlap, pauses and barge-ins, plus labels for tone, emotion and delivery.
Learn more →Expert judgment
Doctors, lawyers, teachers, engineers and hundreds of other professions answering domain tasks and explaining their decisions.
Learn more →Private records
Consumer chats, AI histories, photos and purchases, and expert notes, contributed with consent and profile context.
Learn more →Enterprise data
Non-public emails, chats, documents, meetings and private codebases from consenting companies, sourced against a brief.
Learn more →Coding traces
Real developer-agent sessions from Cursor, Claude Code and Codex: prompt, tool calls, edits and outcomes.
Learn more →Three ways to buy
License it, commission it,
or have it reviewed.
License
Records people and organizations have already made. We agree the domain, formats and permitted use before a dataset changes hands.
- Private personal and professional records
- Enterprise material and codebases
- Source context, metadata and usage rights
Commission
We recruit the people your brief describes and capture new answers, explanations, audio, video or real sessions.
- Consumer and expert responses
- Directed speech and natural dialogue
- Task rationale and profile context
Label and review
Experts annotate and grade against your rubric, calibrated on reviewed samples first.
- Annotations and response grading
- Preference pairs, rubric scores, red-team prompts
- Review notes with every batch
Compared with a typical vendor
Everything you would expect.
Four things you would not.
- Everyday people, not only credentialed experts. 170M+ reachable, recruited by behavior as well as profession.
- A live interview before anyone starts. Every contributor talks to us on voice and video first. Credentials are checked in that conversation.
- Checks after every session. Automated consistency checks when a session ends, then review by trained annotators.
- A way for people to sell their own data. Cookiy Earn and an open-source skill, so supply keeps growing.
Based on the public websites of 27 human data companies we reviewed in September 2026.
Trusted by builders and teams at









Questions teams ask.
Is this simulated data?
No. What we sell comes from real people. Our work on simulated humans and RL environments is separate research, and labelled as such.
Can we start from an incomplete spec?
Yes. Bring the model problem or an example of the output you want. We use it to agree the cohort, the task and the sample criteria.
Is every dataset ready now?
It depends on the cohort, rights, domain and format. Existing material and new collection are scoped separately.
Who are the contributors?
Consumers and professionals outside China, recruited through 30,000+ local partners, plus developers who sell through Cookiy Earn. Everyone consents to the agreed use.
Can several teams share the cost?
Yes. When more than one team needs the same data, we can co-fund collection with agreed exclusivity and usage terms.
Send the spec.
We'll send a sample back.
Video on this site is licensed stock footage of real people, used to illustrate the kinds of data we collect.