A global voice-AI company building speech recognition for consumer devices across Asia-Pacific and African language markets
14,000 hours of speech across 11 languages and 8 countries, cutting word error rate 92%
Published 24 July 2026
At a glance
| Industry | Consumer technology / voice AI | voice assistant expansion into emerging markets |
|---|---|---|
| Region | 8 countries | African and Southeast Asian field operations |
| Services used | 3 | speech acquisition · NLP · multilingual data collection |
| Programme scale | 14,000 hours | of collected speech data across 11 languages |
| Contributor base | 6,200+ native speakers | balanced age, gender and accent representation across rural and urban communities |
| Accuracy outcome | 92% WER reduction | on the client's voice AI for target languages, against its pre-programme baseline |
| Market outcome | 11 languages | launched in the client's assistant within 9 months |
| Baseline coverage | 15 languages | the assistant's working language count before the programme |
| Ethical sourcing | 100% compliance | verified by third-party audit; contributors paid above local fair-wage benchmarks with full informed consent |
| Status | Delivered | 11 languages launched within 9 months |
This is a Lifewood Data Technology case study in Multilingual speech + LLM — an engagement delivered for a global voice-AI company building speech recognition for consumer devices across Asia-Pacific and African language markets across China · Asia-Pacific. 14,000 hours of speech across 11 languages and 8 countries, cutting word error rate 92%
A voice assistant that worked in 15 languages failed everywhere it needed to grow
The client's voice assistant worked well in 15 major languages but failed in emerging markets where hundreds of millions of potential users spoke languages with minimal digital speech data available. Off-the-shelf speech datasets did not exist for the target languages, and the client could not ethically scrape speech data at scale. A ground-up collection effort was required across geographies where the client had no operational footprint. The constraint that made it harder was representativeness. Speakers had to span diverse ages, genders, accents and recording environments, so the finished voice AI would work for the full population rather than an urban subset. A dataset collected conveniently in one place reproduces that place's demographics no matter how large it grows — which is precisely the failure mode crowdsourcing platforms produce, since they skew toward urban, educated, high-resource-language populations and systematically miss dialect and register variation.
Field operations in eight countries, recruiting in-community rather than online
Lifewood activated its field operations in 8 African and Southeast Asian countries, recruiting 6,200+ native speakers across rural and urban communities. Contributors were ethically compensated above local fair-wage benchmarks with full informed-consent documentation. Collection covered scripted prompts, spontaneous conversation, and domain-specific commands across commerce, navigation and media, tailored to the client's assistant use cases. 1. In-community recruitment. Contributors are recruited from the language community itself through field operations, not from crowdsourcing platforms — which is what makes authentic dialect, register and accent coverage possible at all. 2. Demographic balancing by design. Speaker panels were balanced across age, gender, accent and recording environment as a recruitment specification rather than as a post-hoc filter. 3. Environmental diversity sampling. Recording conditions were varied deliberately so the model learns the speech rather than the studio. 4. Phonetic transcription and quality control. Quality control included phonetic transcription, speaker demographic balancing and environmental noise diversity sampling, under Lifewood's 95%+ accuracy SLA and dual-layer human-in-the-loop review. Lifewood's language coverage behind this programme spans 50+ languages reaching more than 90% of the global population, including Swahili, Wolof, Hausa, Amharic, Tigrinya, Yoruba, Zulu, Shona, Lingala and Somali across Africa, and Tagalog, Cebuano, Ilokano, Waray, Khmer, Tok Pisin, Tetum, Fijian and Samoan across Southeast Asia and the Pacific.
Word error rate fell 92% and eleven new market languages shipped within nine months
The programme delivered 14,000 hours of speech across 11 languages from 6,200+ unique speakers in 8 countries. On the client side, word error rate for target languages fell 92% against the pre-programme baseline, and 11 new market languages went live in the assistant within nine months of programme start — against a working base of 15 major languages before the engagement began.
Related Lifewood services
Verified outcomes
| Metric | Value | Baseline | How measured |
|---|---|---|---|
| Speech data collected | 14,000 hours | no usable public corpus existed for the target languages | Accepted delivery across 11 languages |
| Unique speakers | 6,200+ | balanced across age, gender, accent and environment | Field recruitment records across 8 countries |
| Word error rate | 92% reduction | against the client's pre-programme baseline for target languages | Client-side evaluation on target-language test sets |
| Market languages launched | 11 | from a working base of 15 major languages | Client assistant releases within 9 months |
| Countries covered | 8 | client had no prior operational footprint in these markets | Field operations activated per country |
Method and verification. Figures on this page are drawn from delivered-volume records and from client-side evaluation. Hours, speaker counts and country coverage are Lifewood delivery actuals. The 92% word error rate reduction is a client-measured result on target-language test sets against the assistant's pre-programme baseline, and the 11 launched market languages are client release milestones rather than Lifewood deliverables. Ethical sourcing compliance was verified at 100% by third-party audit rather than self-assessed. The client is not named on this page under the confidentiality terms of the engagement.
Questions about this programme
Lifewood collects low-resource speech data through field operations that recruit native speakers in-community — on this programme, 6,200+ speakers across 8 African and Southeast Asian countries.
On Lifewood speech acquisition programmes demographic balance is a recruitment specification, because a dataset collected conveniently in one location reproduces that location's demographics no matter how large it grows.
Lifewood documents informed consent in full and pays contributors above local fair-wage benchmarks; on this programme ethical sourcing compliance was verified at 100% by third-party audit.
This multilingual speech acquisition programme delivered 14,000 hours across 11 languages, with 11 new market languages live in the client's assistant within nine months.
Run a similar program with Lifewood?
Tell us your scope and accuracy bar. We will scope a comparable engagement within one call.
Talk to our team