ParaMed
ParaMed constructs an English–Chinese parallel corpus for biomedical machine translation using articles from the New England Journal of Medicine (NEJM). It processes domain-specific text to generate aligned sentence pairs for training and fine-tuning translation models.
Key Features:
- Biomedical Parallel Corpus Generation: Produces approximately 100,000 aligned English–Chinese sentence pairs comprising ~3,000,000 tokens per language from NEJM articles.
- Domain-Specific Data Processing Pipeline: Acquires and processes NEJM content to create a structured, in-domain bilingual corpus.
- Machine Translation Fine-Tuning: Enhances translation performance by combining out-of-domain data with as few as 4,000 in-domain NEJM sentence pairs.
- Quantified Translation Improvement: Achieves BLEU score improvements of 25.3 (English-to-Chinese) and 13.4 (Chinese-to-English), increasing to 33.0 and 24.3 respectively with larger datasets.
Scientific Applications:
- Biomedical Machine Translation Model Development: Provides high-quality in-domain training data for improving English–Chinese translation of biomedical literature.
Methodology:
ParaMed implements a data acquisition and preprocessing pipeline to extract and align English and Chinese sentences from NEJM articles, generating a domain-specific parallel corpus used to fine-tune machine translation models and evaluate performance using BLEU metrics.
Topics
Details
- License:
- Not licensed
- Cost:
- Free of charge
- Tool Type:
- command-line tool
- Operating Systems:
- Mac, Linux, Windows
- Programming Languages:
- Python, Shell
- Added:
- 2/12/2022
- Last Updated:
- 2/12/2022
Operations
Data Inputs & Outputs
Text mining
Inputs
Outputs
Publications
Liu B, Huang L. ParaMed: a parallel corpus for English–Chinese translation in the biomedical domain. BMC Medical Informatics and Decision Making. 2021;21(1). doi:10.1186/s12911-021-01621-8. PMID:34488734. PMCID:PMC8422666.