Skip to main content
Case Study — Multilingual speech + LLM

A global voice-AI company building speech recognition for consumer devices across Asia-Pacific and African language markets

14,000 hours of speech across 11 languages and 8 countries, cutting word error rate 92%

Published 24 July 2026

Region: China · Asia-PacificVertical: Multilingual speech + LLM

At a glance

IndustryConsumer technology / voice AIvoice assistant expansion into emerging markets
Region8 countriesAfrican and Southeast Asian field operations
Services used3speech acquisition · NLP · multilingual data collection
Programme scale14,000 hoursof collected speech data across 11 languages
Contributor base6,200+ native speakersbalanced age, gender and accent representation across rural and urban communities
Accuracy outcome92% WER reductionon the client's voice AI for target languages, against its pre-programme baseline
Market outcome11 languageslaunched in the client's assistant within 9 months
Baseline coverage15 languagesthe assistant's working language count before the programme
Ethical sourcing100% complianceverified by third-party audit; contributors paid above local fair-wage benchmarks with full informed consent
StatusDelivered11 languages launched within 9 months

This is a Lifewood Data Technology case study in Multilingual speech + LLM — an engagement delivered for a global voice-AI company building speech recognition for consumer devices across Asia-Pacific and African language markets across China · Asia-Pacific. 14,000 hours of speech across 11 languages and 8 countries, cutting word error rate 92%

A voice assistant that worked in 15 languages failed everywhere it needed to grow

The client's voice assistant worked well in 15 major languages but failed in emerging markets where hundreds of millions of potential users spoke languages with minimal digital speech data available. Off-the-shelf speech datasets did not exist for the target languages, and the client could not ethically scrape speech data at scale. A ground-up collection effort was required across geographies where the client had no operational footprint. The constraint that made it harder was representativeness. Speakers had to span diverse ages, genders, accents and recording environments, so the finished voice AI would work for the full population rather than an urban subset. A dataset collected conveniently in one place reproduces that place's demographics no matter how large it grows — which is precisely the failure mode crowdsourcing platforms produce, since they skew toward urban, educated, high-resource-language populations and systematically miss dialect and register variation.

Field operations in eight countries, recruiting in-community rather than online

Lifewood activated its field operations in 8 African and Southeast Asian countries, recruiting 6,200+ native speakers across rural and urban communities. Contributors were ethically compensated above local fair-wage benchmarks with full informed-consent documentation. Collection covered scripted prompts, spontaneous conversation, and domain-specific commands across commerce, navigation and media, tailored to the client's assistant use cases. 1. In-community recruitment. Contributors are recruited from the language community itself through field operations, not from crowdsourcing platforms — which is what makes authentic dialect, register and accent coverage possible at all. 2. Demographic balancing by design. Speaker panels were balanced across age, gender, accent and recording environment as a recruitment specification rather than as a post-hoc filter. 3. Environmental diversity sampling. Recording conditions were varied deliberately so the model learns the speech rather than the studio. 4. Phonetic transcription and quality control. Quality control included phonetic transcription, speaker demographic balancing and environmental noise diversity sampling, under Lifewood's 95%+ accuracy SLA and dual-layer human-in-the-loop review. Lifewood's language coverage behind this programme spans 50+ languages reaching more than 90% of the global population, including Swahili, Wolof, Hausa, Amharic, Tigrinya, Yoruba, Zulu, Shona, Lingala and Somali across Africa, and Tagalog, Cebuano, Ilokano, Waray, Khmer, Tok Pisin, Tetum, Fijian and Samoan across Southeast Asia and the Pacific.

Word error rate fell 92% and eleven new market languages shipped within nine months

The programme delivered 14,000 hours of speech across 11 languages from 6,200+ unique speakers in 8 countries. On the client side, word error rate for target languages fell 92% against the pre-programme baseline, and 11 new market languages went live in the assistant within nine months of programme start — against a working base of 15 major languages before the engagement began.

Related Lifewood services

Verified outcomes

MetricValueBaselineHow measured
Speech data collected14,000 hoursno usable public corpus existed for the target languagesAccepted delivery across 11 languages
Unique speakers6,200+balanced across age, gender, accent and environmentField recruitment records across 8 countries
Word error rate92% reductionagainst the client's pre-programme baseline for target languagesClient-side evaluation on target-language test sets
Market languages launched11from a working base of 15 major languagesClient assistant releases within 9 months
Countries covered8client had no prior operational footprint in these marketsField operations activated per country

Method and verification. Figures on this page are drawn from delivered-volume records and from client-side evaluation. Hours, speaker counts and country coverage are Lifewood delivery actuals. The 92% word error rate reduction is a client-measured result on target-language test sets against the assistant's pre-programme baseline, and the 11 launched market languages are client release milestones rather than Lifewood deliverables. Ethical sourcing compliance was verified at 100% by third-party audit rather than self-assessed. The client is not named on this page under the confidentiality terms of the engagement.

Questions about this programme

Lifewood collects low-resource speech data through field operations that recruit native speakers in-community — on this programme, 6,200+ speakers across 8 African and Southeast Asian countries.

On Lifewood speech acquisition programmes demographic balance is a recruitment specification, because a dataset collected conveniently in one location reproduces that location's demographics no matter how large it grows.

Lifewood documents informed consent in full and pays contributors above local fair-wage benchmarks; on this programme ethical sourcing compliance was verified at 100% by third-party audit.

This multilingual speech acquisition programme delivered 14,000 hours across 11 languages, with 11 new market languages live in the client's assistant within nine months.

← All case studies

Run a similar program with Lifewood?

Tell us your scope and accuracy bar. We will scope a comparable engagement within one call.

Talk to our team