Healthcare AI
Multilingual clinical NLP training data for a US healthcare AI company
The client's growth strategy required expansion into non-English markets: Vietnam, Japan, Korea, Mexico, and Spain were prioritized based on inbound customer inquiry volume from...
91.6% Japanese clinical NLP F1 score achieved
-73% cost per annotated record vs US vendor
$9.1M Japanese hospital contract closed as a direct result
6 languages delivered in production
Client profile
The client is a Boston-based healthcare AI company that develops clinical natural language processing platforms for hospital systems, insurance payers, and pharmaceutical research organizations. Their core product extracts structured clinical data from unstructured physician notes, discharge summaries, and radiology reports. As of engagement start, the platform supported English only, serving 340+ US hospital systems and processing over 60 million clinical documents monthly.
The challenge
The client's growth strategy required expansion into non-English markets: Vietnam, Japan, Korea, Mexico, and Spain were prioritized based on inbound customer inquiry volume from those regions. The clinical NLP problem is genuinely harder than general NLP because of medical vocabulary, drug names, dosages, negation patterns, and clinical abbreviations that vary by country and language.
The client had experimented with two approaches internally before engaging Corpshore. First, they attempted machine translation of their English training corpus into target languages, then re-annotating. The result was an F1 score of 82.1% for Japanese clinical NLP versus 91.4% for English: commercially unusable in a healthcare setting. Second, they attempted contracting native-language annotators through a US-based platform. The quality was reasonable but throughput was one-tenth of what the roadmap required, and cost per annotated document was 4-5 times the English annotation cost.
Timeline pressure was severe. A large Japanese hospital system had signaled interest in a $8.4M annual contract contingent on a Japanese-language deployment ready for pilot by Q3 2026. The client's engineering roadmap required 240,000 clinically-annotated Japanese documents, 180,000 Korean, 220,000 Mandarin, 190,000 Vietnamese, and 210,000 Spanish, with a target accuracy of 90%+ F1 across all six languages.
Why Corpshore Vietnam
Three qualifications separated Corpshore Vietnam from competing bids: our multilingual annotation bench that already included native Vietnamese, Japanese, Korean, and Mandarin speakers with prior healthcare annotation experience; our Corpshore AI division's methodology for handling regulated PHI-adjacent data (even for anonymized synthetic corpuses), including SOC 2 Type II and HIPAA-aligned handling protocols; and our proven ability to scale annotation teams by 3-4x within 60 days when project timelines demanded it.
The client visited our Hanoi facility during evaluation. Their chief data officer specifically flagged the quality of our internal annotation QA process, which included dual-annotator agreement thresholds and clinician-reviewer escalation for ambiguous cases, as more rigorous than the two US-based vendors they had evaluated.
The engagement
Team scaled from 22 annotators at month one to 75 by month five. Team composition: 12 English-native annotators (health science background), 14 Vietnamese, 12 Japanese, 10 Korean, 10 Mandarin, 10 Spanish, plus a clinical review layer of 4 physicians (1 Vietnamese, 1 Japanese, 2 licensed abroad supporting cross-language reviews) and 3 QA leads managing calibration across language teams. All annotators completed a 6-week paid onboarding covering clinical vocabulary, the client's annotation schema, and PHI handling protocols before touching production data.
Approach and methodology
Original-language annotation, not translation. Every annotator worked exclusively in their native or near-native language, using clinical notes originally written in that language rather than translated from English. This is the single most important methodological decision in multilingual clinical NLP, and it drove the accuracy delta we ultimately delivered.
Dual-annotator agreement thresholds. Every clinical concept required agreement between two independent annotators. Disagreements above 3% within a batch triggered clinician review. Below 3%, the higher-confidence annotator's label was retained but flagged for weekly calibration review.
Weekly cross-language calibration. All language teams reviewed a shared batch of anonymized examples weekly to ensure consistent handling of clinically ambiguous concepts (e.g., 'mass' vs 'lesion' vs 'nodule' distinctions) across languages.
RLHF layer. Once the initial annotation phase closed at month 8, the team pivoted to reinforcement learning from human feedback, ranking model-generated outputs to fine-tune the client's downstream models.
Results
The client shipped the Japanese production model at month 9, three weeks ahead of the internal roadmap. The Japanese hospital system contract closed at $9.1M annual (above the initial $8.4M estimate) in month 11. The final Japanese F1 score was 91.6%, exceeding the 90% threshold. Vietnamese, Korean, and Mandarin models followed in months 10, 11, and 12, all above 90% F1. Spanish shipped in month 14 at 89.4% F1 (a known regional challenge with clinical Spanish variance across Mexico, Spain, and Latin America).
The engagement expanded to include an ongoing RLHF program (30 annotators, ongoing) and an AI safety red-team program (12 specialists, launching in month 15 for the client's LLM-based physician assistant product).
| Metric | Baseline (English-only) / Prior Approach | Post-Corpshore | Change |
|---|---|---|---|
| Japanese clinical F1 | 82.1% (translation-based) | 91.6% | +9.5 pts |
| Vietnamese clinical F1 | n/a (no prior model) | 92.4% | New capability |
| Korean clinical F1 | n/a | 91.2% | New capability |
| Mandarin clinical F1 | n/a | 90.8% | New capability |
| Annotation cost per record (US vendor comparison) | $1.42 | $0.38 | -73% |
| Peak weekly throughput per team | 820 records | 3,800 records | +363% |
Original-language annotation with clinician review was the difference. Every other vendor we evaluated wanted to translate our English corpus. Corpshore understood immediately why that would not work clinically. The Japanese pilot outperformed our internal English baseline in some benchmarks.
Enduring value
The methodology developed during this engagement has become the client's internal standard for future language expansions. The clinical annotation schema, dual-annotator protocols, and QA calibration cadence are now documented reference material. The client is negotiating hospital contracts in Vietnam, Thailand, and Malaysia based on the capability this engagement produced.
A comparable engagement to yours?
Bring this example to a discovery call and we will tailor it to your context.
