David AI | High-Quality Audio Datasets for Speech & Conversational AI
Mission
We are an audio data research company. Our mission is to bring AI into the real world through voice, the most important interface to human interaction.
Process
We develop audio datasets with the same rigor researchers bring to models.
1. Hypothesize
Determine an audio AI capability we wish to unlock.
2. Design
Architect a shape of data to teach models that capability.
3. Experiment
Launch a targeted data collection.
4. Evaluate & Iterate
Measure data quality and tune the collection until a small, high-signal set is achieved.
5. Productionize
Scale the dataset to thousands of hours.
6. Release
Publish the dataset, and continuously improve it over time.
Our datasets are used by Fortune 100 companies and research labs that work with speech recognition, translation, synthesis, and conversational AI.
Featured Datasets
A dataset suite designed for speech-to-speech, multilingual, and voice interaction systems
Converse
Our flagship English dataset consists of channel-separated, natural two-speaker conversations spanning a wide range of topics.
Atlas
A multilingual dataset spanning 15+ languages. It includes metadata on dialects and accents and follows the same format as Converse.
Chorus
A dataset of conversations involving three or more speakers. Originally designed for training speaker-separation and diarization models.
Dialog
A collection of expert conversations across a range of domains.
We offer additional proprietary datasets not listed here. Contact us to request a sample, explore more options, or collaborate on a new dataset.
How to access our datasets
1. Request samples
We will set up a quick call to understand your use case and then send you relevant data samples.
2. Purchase access
Enter a data license agreement for the dataset and use-cases your team needs.
3. Receive data
For off-the-shelf datasets, we will grant your team access within one to two days.
Bonus: Experiment with us
We frequently partner with research teams to design new shapes of data for any use case.
Careers
Join us to shape the future of audio AI
We’re hiring for research, engineering, and operations roles.
News
Updates on our progress
Announcing Our $50M Series B Led by Meritech
Oct 8, 2025
Announcing Our $25M Series A Led by Alt Capital
May 22, 2025
Announcing Our $5M Seed Round Led by First Round
Jan 16, 2025