David AI | High-Quality Audio Datasets for Speech & Conversational AI

Mission

We are an audio data research company. Our mission is to bring AI into the real world through voice, the most important interface to human interaction.

Process

We develop audio datasets with the same rigor researchers bring to models.

1. Hypothesize

Determine an audio AI capability we wish to unlock.

2. Design

Architect a shape of data to teach models that capability.

3. Experiment

Launch a targeted data collection.

4. Evaluate & Iterate

Measure data quality and tune the collection until a small, high-signal set is achieved.

5. Productionize

Scale the dataset to thousands of hours.

6. Release

Publish the dataset, and continuously improve it over time.

Our datasets are used by Fortune 100 companies and research labs that work with speech recognition, translation, synthesis, and conversational AI.

Featured Datasets

A dataset suite designed for speech-to-speech, multilingual, and voice interaction systems

Converse

Our flagship English dataset consists of channel-separated, natural two-speaker conversations spanning a wide range of topics.

Atlas

A multilingual dataset spanning 15+ languages. It includes metadata on dialects and accents and follows the same format as Converse.

Chorus

A dataset of conversations involving three or more speakers. Originally designed for training speaker-separation and diarization models.

Dialog

A collection of expert conversations across a range of domains.

We offer additional proprietary datasets not listed here. Contact us to request a sample, explore more options, or collaborate on a new dataset.

How to access our datasets

1. Request samples

We will set up a quick call to understand your use case and then send you relevant data samples.

2. Purchase access

Enter a data license agreement for the dataset and use-cases your team needs.

3. Receive data

For off-the-shelf datasets, we will grant your team access within one to two days.

Bonus: Experiment with us

We frequently partner with research teams to design new shapes of data for any use case.

Careers

Join us to shape the future of audio AI

We’re hiring for research, engineering, and operations roles.

News

Updates on our progress

Announcing Our $50M Series B Led by Meritech

Oct 8, 2025

Announcing Our $25M Series A Led by Alt Capital

May 22, 2025

Announcing Our $5M Seed Round Led by First Round

Jan 16, 2025