En-ViMedNERAn English–Vietnamese Parallel Biomedical Corpus with UMLS Semantic Type Annotations

1VinUniversity 2University of Technology Sydney 3Monash University

*Equal contribution.

For Vietnamese biomedical NER, multilingual pretraining and frontier LLM prompting are no substitute for target-language annotation: a Vietnamese encoder fine-tuned on En-ViMedNER reaches 52.70 F1 — 16.16 points above English-supervised multilingual transfer on the same test set, and 17.31 above the best 3-shot frontier LLM on the human-audited mini-test.

4,392
PubMed abstract pairs
44,892
En–Vi sentence pairs
202,949
Aligned entity-mention pairs
21
UMLS semantic types
Vietnamese biomedical NER — best system per setting mini-test F1 Full test F1
Vietnamese-supervised encoder — ViPubMedDeBERTa† 53.78 52.70
English-supervised multilingual encoder — XLM-R-large 37.95 36.54
Best prompted LLM — gemini-3.1-flash-lite-preview, 3-shot 36.47 —
Best open-source LLM — Qwen-SEA-LION-v4-8B-VL, 3-shot 13.09 —

Vietnamese supervision beats English-supervised transfer by 16.16 F1 and the strongest prompted frontier LLM by 17.31 F1 — so En-ViMedNER is a training resource, not only a benchmark. Prompt-based LLMs are evaluated on the mini-test only. †The best checkpoint differs by split: ViPubMedDeBERTa-xsmall on the mini-test (53.78), ViPubMedDeBERTa-base on the full test set (52.70).

Summary

Biomedical named entity recognition underpins clinical decision support and medical information extraction. English has UMLS-annotated corpora such as MedMentions; Vietnamese has nothing comparable. We present En-ViMedNER, the first English–Vietnamese parallel biomedical NER corpus annotated with UMLS semantic types — language-neutral codes that give both sides a shared label space and keep the resource directly comparable with existing UMLS-based corpora.

The corpus contains 4,392 PubMed abstract pairs, 44,892 sentence pairs, and 202,949 aligned entity-mention pairs across the 21 semantic types of MedMentions ST21pv. To balance quality against scale, we built it through automatic translation, medical expert post-editing, LLM-assisted label projection, and human verification and adjudication. We characterise it as a large-scale silver-standard corpus with a human-audited, consensus-corrected 450-sentence mini-test subset.

We evaluate it in two settings. In Vietnamese-input/Vietnamese-output NER, the best model reaches 52.70 F1 on the test set and 53.78 on the mini-test. In English-input/Vietnamese-output cross-lingual NER, the best model reaches 45.44 F1 on the mini-test. We release the corpus, the construction pipeline, and the baseline models, and we argue that similar semi-automated efforts are worth making for other low-resource languages.

How we built the corpus

We start from MedMentions ST21pv, drop overlapping and unlabeled sentences, and carry the English UMLS annotations across to Vietnamese in three stages.

En-ViMedNER corpus construction pipeline English MedMentions sentences and their UMLS spans are machine translated and post-edited by medical experts; an LLM then projects each English entity span onto the post-edited Vietnamese text; automatic checks and human review produce the released corpus. SOURCE TRANSLATE & POST-EDIT PROJECT VERIFY & RELEASE MedMentions ST21pv 4,392 PubMed abstracts 21 UMLS semantic types English sentences 44,892 sentences with UMLS entity spans Machine translation vinai-translate-en2vi-v2 Expert post-editing Vietnamese doctors and final-year med students LLM label projection gpt-5.2 fills the blank Vi fields of a JSON template Checks + human review 1.77% of pairs flagged, fixed in Label Studio En-ViMedNER 202,949 aligned entity-mention pairs English spans + labels Vietnamese text Mini-test audit — 450 sentences, two independent reviewers Cohen’s κ = 0.83 · 98.32% exact match before correction English source and labels Vietnamese text Automated + human processing Released corpus

The English labels never pass through the translator: they are projected onto the already-finalized, expert-post-edited Vietnamese text. That rules out marker-based projection, which needs markers in the English source before translation, and weakens token-alignment tools, because post-editing often paraphrases non-literally (animals → động vật thí nghiệm, "laboratory animals").

The audit measures the projection itself: 98.32% entity-level exact match before any manual correction, with pooled Cohen's κ = 0.83 between the two independent reviewers. The mini-test we release is the corrected version; the 98.32% figure describes what the LLM projection got right on its own.

SplitAbstract pairsSentence pairsAligned entity-mention pairs
train2,63527,013122,017
dev8788,93240,844
test8798,94740,088
mini-test†—4502,085
Total4,39244,892202,949

†The mini-test is sampled from the test set for human evaluation; it is not an additional split.

What we found

Vietnamese supervision cannot be replaced by English supervision. Fine-tuned on the Vietnamese split, ViPubMedDeBERTa-base reaches 52.70 F1 on the full test set — 16.16 points above XLM-R-large fine-tuned on the English split and evaluated on the same Vietnamese text. Multilingual pretraining does not transfer Vietnamese entity boundaries, terminology, or labels for free.

Few-shot examples help LLMs, but not enough. Prompting improves steeply with examples — gemini-3.1-flash-lite-preview goes from 12.70 to 36.47 F1 between zero-shot and 3-shot, and Qwen-SEA-LION-v4-8B-VL from 1.59 to 13.09 — yet the best prompted model still trails the best fine-tuned encoder by 17.31 F1. In-context examples teach the BIO format, not the biomedical task.

Parallel supervision also carries cross-lingual NER. On English-input/ Vietnamese-output generation, fine-tuned NLLB-M1 reaches 45.44 Judge-F1, ahead of gpt-5.4-3shot by 2.80 and gemini-3.1-flash-lite-preview-3shot by 3.89 points.

End-to-end and cascaded pipelines win under different metrics. Direct translate-and-tag generation is best under Judge-F1 (which accepts a valid Vietnamese paraphrase of a gold entity), while the tag-first NER → translate cascade is best under the stricter Exact-F1, reaching 15.98 for NLLB-M3. Which pipeline is better depends on what you are willing to count as correct.

Higher BLEU does not imply better NER. NLLB-M2 has the best BLEU of the NLLB systems but a lower F1 than NLLB-M1; the same inversion appears for umT5. Fluent translation is not the same as preserving entity spans and labels through translation.

Open-source LLMs mostly fail to emit valid output. Between 114 and 292 of 450 cross-lingual outputs come back unparseable or entirely untagged, which caps their Judge-F1 at 18.43 even for the strongest of them.

Cross-lingual NER: English in, tagged Vietnamese out

Three cross-lingual NER pipelines M1 generates tagged Vietnamese directly from plain English. M2 translates first, then tags the Vietnamese. M3 tags the English first, then translates the tagged sentence. M1 Direct English sentence plain text Direct Trans.+NER Vietnamese sentence with NER tags Best Judge-F1 — NLLB 45.44 M2 Trans. → NER English sentence plain text Vietnamese sentence plain text Trans. only NER only Vietnamese sentence with NER tags Best BLEU — NLLB 51.83 M3 NER → Trans. English sentence plain text English sentence with NER tags NER only Tag-aware Trans. Vietnamese sentence with NER tags Best Exact-F1 — NLLB 15.98

The three fine-tuned settings. M1 generates the tagged Vietnamese sentence in one step; M2 and M3 split translation and tagging, differing in which comes first.

Mini-test results. Judge-F1 comes from an LLM-as-a-judge protocol that accepts a semantically equivalent Vietnamese entity translation (human–LLM agreement, Cohen's κ = 0.87); Exact-F1 requires an exact surface and label match. Inv. counts generated sentences with no valid entity tags, out of 450. The fine-tuned encoder-decoders are umT5-base and NLLB-200-distilled-600M.

Setting / ModelBLEUJudge-F1Exact-F1Inv.
Fine-tuned multilingual encoder-decoder models
NLLB-M1: Direct Trans.+NER48.6345.4415.770
NLLB-M2: Trans. → NER51.8340.8713.640
NLLB-M3: NER → Trans.46.9039.0215.980
umT5-M1: Direct Trans.+NER29.8526.886.010
umT5-M3: NER → Trans.31.1526.647.330
Prompt-based LLMs (3-shot)
gpt-5.453.3442.649.840
gemini-3.1-flash-lite-preview48.8041.5510.400
Qwen-SEA-LION-v4-8B-VL52.8918.439.85228
Qwen2.5-7B-Instruct43.6611.945.34115

The highest-BLEU rows are not the highest-F1 rows: good sentence-level translation does not by itself preserve entity boundaries and labels.

Translation quality against NER quality

Scatter plot of BLEU against F1 for the six encoder-decoder settings. NLLB-M1 sits highest in F1 at 45.4 without the highest BLEU; NLLB-M2 sits furthest right at 51.8 BLEU with a lower F1.

Circles are umT5-base, squares are NLLB-200-distilled-600M; each point is one setting. The axes do not rank the settings the same way — M2 sits furthest right in both model families, but M1 sits highest.

Limitations

En-ViMedNER is a silver-standard corpus: exhaustive independent verification and consensus adjudication cover the 450-sentence mini-test, not all 44,892 sentence pairs, and random sampling does not guarantee coverage of every rare semantic type or difficult translation phenomenon. Annotation and evaluation run at the sentence level, so translations that depend on abstract-level context (animals → động vật thí nghiệm, "laboratory animals") can read as unnatural in isolation, and the same surface form may be translated in one place and kept in English in another. The corpus is built from PubMed abstracts, so models trained on it should not be assumed reliable on clinical notes, patient-facing text, or safety-critical applications without further validation.

BibTeX

@inproceedings{vo-etal-2026-en-vimedner,
    title = "{E}n-{V}i{M}ed{NER}: An {E}nglish-{V}ietnamese Parallel Biomedical Corpus with {UMLS} Semantic Type Annotations",
    author = "Vo, Nhu  and
      Nguyen, Phuong  and
      Le, Nu Uyen Phuong  and
      Jauregi Unanue, Inigo  and
      Le, Dung D.  and
      Piccardi, Massimo  and
      Buntine, Wray",
    booktitle = "Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing ({EMNLP} 2026)",
    month = oct,
    year = "2026",
    address = "Budapest, Hungary",
    publisher = "Association for Computational Linguistics",
    note = "To appear"
}

Acknowledgments

This work is part of our wider medical machine translation project at VinUniversity. This research was undertaken within the framework of the Cross-College project Robust Vietnamese–English Clinical and Educational Medical Translation (Project ID: VUNI.2324.CC06), jointly undertaken by the College of Engineering & Computer Science and the College of Health Sciences at VinUniversity. Nhu Vo acknowledges financial support from the Vingroup Scholarship and the University of Technology Sydney. Dung D. Le acknowledges partial funding from the Center for AI Research at VinUniversity.

En-ViMedNER follows the licensing conditions of both MedMentions (CC0) and UMLS. We retain only UMLS semantic type codes (e.g. T047) as labels, and do not publish UMLS concept names, definitions, or hierarchy information.