Speech and conversation

Speech,
beyond the words.

Real conversations between real people, labelled for tone, pauses, overlap and emotion. In many languages.

A script read in a booth does not sound like people. We record the real thing and label what the words leave out.

Every recording comes with an aligned transcript, speaker turns and timestamps, and the labels your rubric asks for.

  • Real conversations, recorded with everyone's consent.
  • Three pairings: person to person, expert to person, expert to expert.
  • Rubric agreed on a sample before anything scales.

Three kinds of conversation

Who is talking changes
how they talk.

Friends, family, couples

Person to person

Friends, couples and families on a call or at the table. Overlap, laughter, backchannels and topic jumps.

Doctor and patient

Expert to person

Doctor and patient, teacher and student, therapist and client. Explanation, reassurance and clarifying questions.

Two practitioners

Expert to expert

Two lawyers, two engineers, a clinical team. Dense vocabulary, interruptions, disagreement and how it resolves.

Labels

Same words.
Different meaning.

  • Voice. Tone, emotional intensity, pace, pauses and emphasis.
  • Interaction. Speaker turns, overlap, barge-ins and backchannels, with timestamps.
  • Video, when included. Facial expression and gesture alongside the speech.
  • Context. Scenario, intended delivery and speaker profile.
"I SAID IT'S FINE." · THREE DELIVERIES, THREE LABEL SETSreassurancetone: warm · pace: slow · pause: 0.4s · emphasis: "fine"frustrationtone: tense · pace: fast · pause: 0.0s · emphasis: "said"ironytone: flat · pace: even · pause: 1.1s · pitch fall on "fine"DELIVERY: audio/video + aligned transcript + labels + speaker and scenario metadata
Illustrative labels. Project rubrics are agreed on real samples.

Many languages

Recruited by language, accent and market.

A dataset should sound like the people who will use the model. We recruit by country, language, accent and background, including low-resource languages, and review every cohort on a sample before it grows.

East Asia
South Asia
Africa
Latin America
Middle East
Europe
Southeast Asia

For voice and interactive models

Building a full-duplex or agentic voice model?

Voice teams have their own page: a 70-second film, the stack we label for, action timelines for agentic voice, and how we test your model against real conversations.

Go to the voice page →

Questions teams ask.

Is it real conversation or acted?

Both, labelled separately. Natural sessions are recorded with consent between people who actually know each other or work together. Directed sessions follow a scenario you set, such as an emotion or an interruption pattern.

Can a label tell us what someone really thinks?

No. Labels describe what can be heard and seen, plus the agreed scenario or stated intent. They are not a reading of anyone's private mental state.

Why is Cookiy good at this?

We have run 190,000+ voice and video research interviews and built tooling to review tone, pace and expression in them. The same tooling, with trained annotators, produces these labels.

Which languages?

English, Mandarin and major European and Asian languages as standard. Low-resource languages and regional accents on request, checked on a sample before the cohort grows.

Which voices,
which scenarios?

Video on this site is licensed stock footage of real people, used to illustrate the kinds of data we collect.