ParaMed

English–Chinese biomedical parallel corpus construction from NEJM


ParaMed constructs an English–Chinese parallel corpus for biomedical machine translation using articles from the New England Journal of Medicine (NEJM). It processes domain-specific text to generate aligned sentence pairs for training and fine-tuning translation models.


Key Features:

  • Biomedical Parallel Corpus Generation: Produces approximately 100,000 aligned English–Chinese sentence pairs comprising ~3,000,000 tokens per language from NEJM articles.
  • Domain-Specific Data Processing Pipeline: Acquires and processes NEJM content to create a structured, in-domain bilingual corpus.
  • Machine Translation Fine-Tuning: Enhances translation performance by combining out-of-domain data with as few as 4,000 in-domain NEJM sentence pairs.
  • Quantified Translation Improvement: Achieves BLEU score improvements of 25.3 (English-to-Chinese) and 13.4 (Chinese-to-English), increasing to 33.0 and 24.3 respectively with larger datasets.

Scientific Applications:

  • Biomedical Machine Translation Model Development: Provides high-quality in-domain training data for improving English–Chinese translation of biomedical literature.

Methodology:

ParaMed implements a data acquisition and preprocessing pipeline to extract and align English and Chinese sentences from NEJM articles, generating a domain-specific parallel corpus used to fine-tune machine translation models and evaluate performance using BLEU metrics.

Topics

Details

License:
Not licensed
Cost:
Free of charge
Tool Type:
command-line tool
Operating Systems:
Mac, Linux, Windows
Programming Languages:
Python, Shell
Added:
2/12/2022
Last Updated:
2/12/2022

Operations

Data Inputs & Outputs

Text mining

Publications

Liu B, Huang L. ParaMed: a parallel corpus for English–Chinese translation in the biomedical domain. BMC Medical Informatics and Decision Making. 2021;21(1). doi:10.1186/s12911-021-01621-8. PMID:34488734. PMCID:PMC8422666.