The speech data your models have never heard.
Proprietary speech corpora for low-resource languages: native audio, transcription, 14 annotation layers, commercial rights included. Produced in studio with native speakers.
- audio_file
- fon_04217.wav
- language
- fon-BJ
- speaker_id
- BEN-FON-083
- gender
- F
- age_band
- 30-39
- region
- Littoral
- speech_type
- Spontaneous
- split
- train
- domain
- Culture
- consent_id
- BEN-FON-083-C1
- transcription
- Mi ɖo gbɛ na mì...
- translation_en
- We have something to tell you...
- duration
- 6.42 s
- snr
- 41 dB
- sampling_rate
- 48 kHz
The languages we serve, produced continuously.
Languages nearly absent from the web, and therefore from your training data. We produce them in studio, with native speakers under contract.
- Fon
fon-BJ - Wolof
wo-SN - Beninese French
fr-BJ - Senegalese French
fr-SN - Ewe
ee-BJ - Yoruba
yo-BJ - Hausa
ha-BJ - Pulaar
fuc-SN - Bambara
bm-SN - Maninka
mlq-SN - Darija
ary-MA - Moroccan French
fr-MA - Swahili
sw-TZ - Tanzanian English
en-TZ - Dioula
dyu-CI - Baoulé
bci-CI - Ivorian French
fr-CI - Fulfulde
fub-CM - Pidgin
wes-CM - Cameroonian French
fr-CM
Every hour is annotated across 14 layers. One corpus, many voice systems.
ASR, TTS, speech translation, voice agents, voice servers: the same corpus powers the training, fine-tuning and evaluation of your models and machines.
Gold standard, measured quality
This data cannot be scraped. We produce it, end to end.
Local studios and mobile units, native narrators under contract, trained transcribers: we control every link, from speaker to delivered dataset.
Our field presence: Labari.io, our free listening app, streams stories and folk tales in local languages and sustains our narrator network. labari.io →
Three ways to access the data.
From a one-off test to language exclusivity: same quality standard, same rights, in every format.
Catalog license
Annual subscription access to one or more languages: search, export and delivery of structured datasets.
Custom dataset
Dedicated production for your needs: language, domain, register, speech type, volume. You specify, we produce in studio.
Temporary exclusivity
An exclusivity window on a language, domain or corpus, before it joins the shared catalog.
Test the data on your models.
Describe your use case (ASR, TTS, translation, agents) and the languages you need: we will prepare a sample.