For voice and interactive models

Your model can talk.
Can it hold a real conversation?

Real conversations, labelled turn by turn, by a team that builds voice agents.

The short version

You buy our data.
We test your model.

We sell conversation data to voice teams, and we run our own interviews and research with real people. So we can also test your model against real conversations and show you where it breaks.

See the conversations we record →

70 seconds: what voice teams buy from us, and what we buy back.

Pull That Up · research preview

We build voice agents.
We know the data they need.

Our research preview is a real-time assistant for solo and group conversations. Building it taught us how to source, prepare and label data for the whole stack.

01 / Hear

Speech to text

Speakers, timestamps, overlapping speech and instructions that change mid-sentence.

02 / Orchestrate

LLM, agent harness, tools

Act or wait? Which tool? Track context, interruptions and results that arrive later. Decide when to speak and when to show a card.

03 / Respond

Text to speech

Response timing, turn-taking, intended delivery and expressive speech.

Four timelines, labelled together

TranscriptWhat was said, by whom and when
ContextTopic, goals and changing requests
ActionTool triggers, results and response timing
ExpressionTone, pauses and scenario-defined intent

What you can commission

Recording, transcripts,
labels, review.

  • Natural and directed recordings, solo, two-speaker and group
  • Person to person, expert to person and expert to expert
  • Many languages and accents, including low-resource ones
  • STT transcripts with speaker turns and timestamps
  • Paralinguistic labels: tone, pauses, emotion, emphasis, delivery
  • TTS performance labels and expert review against your rubric
  • Action and tool-call timelines for agentic voice

Agree the spec and acceptance criteria, review a sample, then scale. Details on the speech data page.

Questions voice teams ask.

Is this real conversation or acted?

Both, labelled separately. Natural conversations are recorded with consent between people who actually know each other or work together. Directed sessions follow a scenario you specify, for example an emotion or an interruption pattern. Synthetic or augmented data is always scoped and labelled as such.

Which languages?

English, Mandarin and major European and Asian languages as standard, and low-resource languages and regional accents on request. We confirm recruiting and a sample before scaling.

How is quality checked?

Every session gets automated consistency checks when it ends. Trained annotators, including linguists where the brief needs them, review what gets flagged.

Do you train speech models?

No. Our focus is agent orchestration and data pipelines, not end-to-end speech model training. Pull That Up is a research preview that taught us what interactive voice agents need from data.

Can we share the cost of a dataset?

Yes. If several teams need the same kind of data, we can co-fund collection with agreed exclusivity and usage terms.

Bring us your hardest conversation.
We'll run the sample.

Video on this site is licensed stock footage of real people, used to illustrate the kinds of data we collect.