The speech data your models have never heard.

Proprietary speech corpora for low-resource languages: native audio, transcription, 14 annotation layers, commercial rights included. Produced in studio with native speakers.

Fon
audio_file
fon_04217.wav
language
fon-BJ
speaker_id
BEN-FON-083
gender
F
age_band
30-39
region
Littoral
speech_type
Spontaneous
split
train
domain
Culture
consent_id
BEN-FON-083-C1
transcription
Mi ɖo gbɛ na mì...
translation_en
We have something to tell you...
duration
6.42 s
snr
41 dB
sampling_rate
48 kHz
20+ languages & local varieties500M+ speakers covered14 layers of annotationGold standard multi-annotatorRights secured at the source
Catalog

The languages we serve, produced continuously.

Languages nearly absent from the web, and therefore from your training data. We produce them in studio, with native speakers under contract.

500M+ speakers covered
  • Fonfon-BJ
  • Wolofwo-SN
  • Beninese Frenchfr-BJ
  • Senegalese Frenchfr-SN
  • Eweee-BJ
  • Yorubayo-BJ
  • Hausaha-BJ
  • Pulaarfuc-SN
  • Bambarabm-SN
  • Maninkamlq-SN
  • Darijaary-MA
  • Moroccan Frenchfr-MA
  • Swahilisw-TZ
  • Tanzanian Englishen-TZ
  • Diouladyu-CI
  • Baoulébci-CI
  • Ivorian Frenchfr-CI
  • Fulfuldefub-CM
  • Pidginwes-CM
  • Cameroonian Frenchfr-CM
Annotation

Every hour is annotated across 14 layers. One corpus, many voice systems.

ASR, TTS, speech translation, voice agents, voice servers: the same corpus powers the training, fine-tuning and evaluation of your models and machines.

Text
VerbatimNormalizedITN
Language & pronunciation
Code-switchingTones & geminationG2P
Time & speakers
AlignmentDiarization
Meaning
TranslationSemanticsProsody & emotion
Context & usage
Acoustic eventsIndexingDialogue acts

Gold standard, measured quality

Multi-expert annotation3 independent native speakers, blind, arbitration by a senior linguist
Reliability measured and publishedKrippendorff α, WER, CER, delivered in the datasheet
Strict linguistic complianceofficial orthography, tones and gemination respected
Legal compliancecommercial and AI consent traced per speaker
Linguistic signaturetonal error rate, which no competitor publishes
ASRTTSSpeech translationVoice agentsVoice serversAudio searchEvaluation
Production

This data cannot be scraped. We produce it, end to end.

Local studios and mobile units, native narrators under contract, trained transcribers: we control every link, from speaker to delivered dataset.

Narratorcontract, GDPR consent
Recordingstudios, mobile units
Annotation14 layers, linguist arbitration
Quality + rightsGold validation, IP assignment
Catalogdelivery by secure export
Your modeltraining, fine-tuning, evaluation

Our field presence: Labari.io, our free listening app, streams stories and folk tales in local languages and sustains our narrator network. labari.io →

Licensing

Three ways to access the data.

From a one-off test to language exclusivity: same quality standard, same rights, in every format.

Catalog license

Annual subscription access to one or more languages: search, export and delivery of structured datasets.

Custom dataset

Dedicated production for your needs: language, domain, register, speech type, volume. You specify, we produce in studio.

Temporary exclusivity

An exclusivity window on a language, domain or corpus, before it joins the shared catalog.

Data ready for production
GDPR consenttraced per recording
Intellectual propertycomplete chain of rights, from speaker to dataset
Clear commercial licenseno ambiguity for AI use
Contact

Test the data on your models.

Describe your use case (ASR, TTS, translation, agents) and the languages you need: we will prepare a sample.