本文へスキップ
Corpshore Vietnam

Healthcare AI

Multilingual clinical NLP training data for a US healthcare AI company

The client's growth strategy required expansion into non-English markets: Vietnam, Japan, Korea, Mexico, and Spain were prioritized based on inbound customer inquiry volume from...

AI DeliveryUnited StatesMid-market
  • 91.6% Japanese clinical NLP F1 score achieved

  • -73% cost per annotated record vs US vendor

  • $9.1M Japanese hospital contract closed as a direct result

  • 6 languages delivered in production

Client profile

The client is a Boston-based healthcare AI company that develops clinical natural language processing platforms for hospital systems, insurance payers, and pharmaceutical research organizations. Their core product extracts structured clinical data from unstructured physician notes, discharge summaries, and radiology reports. As of engagement start, the platform supported English only, serving 340+ US hospital systems and processing over 60 million clinical documents monthly.

The challenge

The client's growth strategy required expansion into non-English markets: Vietnam, Japan, Korea, Mexico, and Spain were prioritized based on inbound customer inquiry volume from those regions. The clinical NLP problem is genuinely harder than general NLP because of medical vocabulary, drug names, dosages, negation patterns, and clinical abbreviations that vary by country and language.

The client had experimented with two approaches internally before engaging Corpshore. First, they attempted machine translation of their English training corpus into target languages, then re-annotating. The result was an F1 score of 82.1% for Japanese clinical NLP versus 91.4% for English: commercially unusable in a healthcare setting. Second, they attempted contracting native-language annotators through a US-based platform. The quality was reasonable but throughput was one-tenth of what the roadmap required, and cost per annotated document was 4-5 times the English annotation cost.

Timeline pressure was severe. A large Japanese hospital system had signaled interest in a $8.4M annual contract contingent on a Japanese-language deployment ready for pilot by Q3 2026. The client's engineering roadmap required 240,000 clinically-annotated Japanese documents, 180,000 Korean, 220,000 Mandarin, 190,000 Vietnamese, and 210,000 Spanish, with a target accuracy of 90%+ F1 across all six languages.

Why Corpshore Vietnam

Three qualifications separated Corpshore Vietnam from competing bids: our multilingual annotation bench that already included native Vietnamese, Japanese, Korean, and Mandarin speakers with prior healthcare annotation experience; our Corpshore AI division's methodology for handling regulated PHI-adjacent data (even for anonymized synthetic corpuses), including SOC 2 Type II and HIPAA-aligned handling protocols; and our proven ability to scale annotation teams by 3-4x within 60 days when project timelines demanded it.

The client visited our Hanoi facility during evaluation. Their chief data officer specifically flagged the quality of our internal annotation QA process, which included dual-annotator agreement thresholds and clinician-reviewer escalation for ambiguous cases, as more rigorous than the two US-based vendors they had evaluated.

The engagement

Team scaled from 22 annotators at month one to 75 by month five. Team composition: 12 English-native annotators (health science background), 14 Vietnamese, 12 Japanese, 10 Korean, 10 Mandarin, 10 Spanish, plus a clinical review layer of 4 physicians (1 Vietnamese, 1 Japanese, 2 licensed abroad supporting cross-language reviews) and 3 QA leads managing calibration across language teams. All annotators completed a 6-week paid onboarding covering clinical vocabulary, the client's annotation schema, and PHI handling protocols before touching production data.

Approach and methodology

Original-language annotation, not translation. Every annotator worked exclusively in their native or near-native language, using clinical notes originally written in that language rather than translated from English. This is the single most important methodological decision in multilingual clinical NLP, and it drove the accuracy delta we ultimately delivered.

Dual-annotator agreement thresholds. Every clinical concept required agreement between two independent annotators. Disagreements above 3% within a batch triggered clinician review. Below 3%, the higher-confidence annotator's label was retained but flagged for weekly calibration review.

Weekly cross-language calibration. All language teams reviewed a shared batch of anonymized examples weekly to ensure consistent handling of clinically ambiguous concepts (e.g., 'mass' vs 'lesion' vs 'nodule' distinctions) across languages.

RLHF layer. Once the initial annotation phase closed at month 8, the team pivoted to reinforcement learning from human feedback, ranking model-generated outputs to fine-tune the client's downstream models.

Results

The client shipped the Japanese production model at month 9, three weeks ahead of the internal roadmap. The Japanese hospital system contract closed at $9.1M annual (above the initial $8.4M estimate) in month 11. The final Japanese F1 score was 91.6%, exceeding the 90% threshold. Vietnamese, Korean, and Mandarin models followed in months 10, 11, and 12, all above 90% F1. Spanish shipped in month 14 at 89.4% F1 (a known regional challenge with clinical Spanish variance across Mexico, Spain, and Latin America).

The engagement expanded to include an ongoing RLHF program (30 annotators, ongoing) and an AI safety red-team program (12 specialists, launching in month 15 for the client's LLM-based physician assistant product).

MetricBaseline (English-only) / Prior ApproachPost-CorpshoreChange
Japanese clinical F182.1% (translation-based)91.6%+9.5 pts
Vietnamese clinical F1n/a (no prior model)92.4%New capability
Korean clinical F1n/a91.2%New capability
Mandarin clinical F1n/a90.8%New capability
Annotation cost per record (US vendor comparison)$1.42$0.38-73%
Peak weekly throughput per team820 records3,800 records+363%
Original-language annotation with clinician review was the difference. Every other vendor we evaluated wanted to translate our English corpus. Corpshore understood immediately why that would not work clinically. The Japanese pilot outperformed our internal English baseline in some benchmarks.
Chief Data Officer, US healthcare AI company

Enduring value

The methodology developed during this engagement has become the client's internal standard for future language expansions. The clinical annotation schema, dual-annotator protocols, and QA calibration cadence are now documented reference material. The client is negotiating hospital contracts in Vietnam, Thailand, and Malaysia based on the capability this engagement produced.

A comparable engagement to yours?

Bring this example to a discovery call and we will tailor it to your context.

Request a proposal

More case studies