Overview
A global consumer-electronics brand was preparing its first on-device generative assistant. English testing looked strong — but early pilots in other languages produced stilted phrasing, wrong honorifics and answers that missed local context entirely.
We built a fourteen-language data and evaluation program: native prompt collection, preference ranking, red-teaming and scored fluency audits, all feeding the client's training loop on a weekly cadence.
- Client
- Global device maker (consumer electronics)
- Scope
- 14 languages, data + evaluation
- Timeline
- Ongoing program, quarterly releases
- Services
- AI Data Services, Testing & QA
The challenge
The assistant needed to feel native, not translated: correct levels of formality in Japanese and Korean, dialect-aware Arabic, and culturally appropriate examples everywhere. Crowd-sourced data had produced generic prompts and inconsistent judgments, and internal reviewers couldn't cover the language matrix.
Our approach
Native collection
Contributors wrote realistic prompts per market — shopping, health, device help.
Preference ranking
Trained raters ranked responses for helpfulness, tone and accuracy.
Red-teaming
Adversarial testers probed safety and cultural edge cases per locale.
Scored audits
Weekly fluency scorecards tracked progress release over release.
The outcome
The assistant launched simultaneously in all fourteen languages with native-quality ratings from in-country reviewers. Fluency scores improved release over release, safety incidents in pilot markets dropped to near zero, and the client's team now plans launches with localization data from day one — not as a last-mile fix.
Related services
More solutions in action
Building a multilingual model?
We will design a data and evaluation pilot around your roadmap.