Tucuxi-BLAST
Tucuxi-BLAST applies DNA encoding and BLASTn alignment to perform scalable record linkage across large health-related administrative databases, improving linkage accuracy and tolerance to data-entry errors.
Key Features:
- DNA Encoding: Identification records are encoded into DNA sequences for representation and comparison.
- BLASTn Algorithm Integration: Uses the BLASTn algorithm to align DNA-encoded records across different databases.
- Error Tolerance: Accommodates misspellings and typographical errors in administrative records to improve linkage accuracy.
- Performance Efficiency: Demonstrated scalability in benchmarks with a simulated database of 300 million individuals and completed a 200,000-record linkage in 23 hours compared with approximately five days and seven hours for a state-of-the-art method.
- Superior Accuracy: Outperformed five existing record linkage tools on a gold-standard dataset derived from real health-related databases in both accuracy and speed.
Scientific Applications:
- Public Health Data Integration: Enables fast, large-scale linking of administrative health databases for integrated analyses.
- Medical and Epidemiological Research: Supports analyses that require merging records across disparate health-related datasets.
- Health Policy and Healthcare Delivery: Facilitates creation of comprehensive linked datasets to inform policy-making and healthcare delivery improvements.
Methodology:
Identification records are encoded into DNA sequences and aligned using the BLASTn algorithm.
Topics
Details
- License:
- GPL-3.0
- Cost:
- Free of charge
- Tool Type:
- command-line tool
- Operating Systems:
- Mac, Linux, Windows
- Programming Languages:
- Python
- Added:
- 9/27/2022
- Last Updated:
- 11/24/2024
Operations
Publications
Araujo JD, Santos-e-Silva JC, Costa-Martins AG, Sampaio V, de Castro DB, de Souza RF, Giddaluru J, Ramos PIP, Pita R, Barreto ML, Barral-Netto M, Nakaya HI. Tucuxi-BLAST: Enabling fast and accurate record linkage of large-scale health-related administrative databases through a DNA-encoded approach. PeerJ. 2022;10:e13507. doi:10.7717/peerj.13507. PMID:35846888. PMCID:PMC9281601.
DOI: 10.7717/peerj.13507
PMID: 35846888
PMCID: PMC9281601
Funding: - Brazilian agency Fundação de Amparo à Pesquisa do Estado de São Paulo-FAPESP: 2018/14933-2
- FAPESP: 2019/27139-5