For voice and interactive models
Your model can talk.
Can it hold a real conversation?
Real conversations, labelled turn by turn, by a team that builds voice agents.
The short version
You buy our data.
We test your model.
We sell conversation data to voice teams, and we run our own interviews and research with real people. So we can also test your model against real conversations and show you where it breaks.
Pull That Up · research preview
We build voice agents.
We know the data they need.
Our research preview is a real-time assistant for solo and group conversations. Building it taught us how to source, prepare and label data for the whole stack.
Speech to text
Speakers, timestamps, overlapping speech and instructions that change mid-sentence.
LLM, agent harness, tools
Act or wait? Which tool? Track context, interruptions and results that arrive later. Decide when to speak and when to show a card.
Text to speech
Response timing, turn-taking, intended delivery and expressive speech.
Four timelines, labelled together
What you can commission
Recording, transcripts,
labels, review.
- Natural and directed recordings, solo, two-speaker and group
- Person to person, expert to person and expert to expert
- Many languages and accents, including low-resource ones
- STT transcripts with speaker turns and timestamps
- Paralinguistic labels: tone, pauses, emotion, emphasis, delivery
- TTS performance labels and expert review against your rubric
- Action and tool-call timelines for agentic voice
Agree the spec and acceptance criteria, review a sample, then scale. Details on the speech data page.
Questions voice teams ask.
Is this real conversation or acted?
Both, labelled separately. Natural conversations are recorded with consent between people who actually know each other or work together. Directed sessions follow a scenario you specify, for example an emotion or an interruption pattern. Synthetic or augmented data is always scoped and labelled as such.
Which languages?
English, Mandarin and major European and Asian languages as standard, and low-resource languages and regional accents on request. We confirm recruiting and a sample before scaling.
How is quality checked?
Every session gets automated consistency checks when it ends. Trained annotators, including linguists where the brief needs them, review what gets flagged.
Do you train speech models?
No. Our focus is agent orchestration and data pipelines, not end-to-end speech model training. Pull That Up is a research preview that taught us what interactive voice agents need from data.
Can we share the cost of a dataset?
Yes. If several teams need the same kind of data, we can co-fund collection with agreed exclusivity and usage terms.
Bring us your hardest conversation.
We'll run the sample.
Video on this site is licensed stock footage of real people, used to illustrate the kinds of data we collect.