Tucuxi-BLAST

Tucuxi-BLAST applies DNA encoding and BLASTn alignment to perform scalable record linkage across large health-related administrative databases, improving linkage accuracy and tolerance to data-entry errors.


Key Features:

  • DNA Encoding: Identification records are encoded into DNA sequences for representation and comparison.
  • BLASTn Algorithm Integration: Uses the BLASTn algorithm to align DNA-encoded records across different databases.
  • Error Tolerance: Accommodates misspellings and typographical errors in administrative records to improve linkage accuracy.
  • Performance Efficiency: Demonstrated scalability in benchmarks with a simulated database of 300 million individuals and completed a 200,000-record linkage in 23 hours compared with approximately five days and seven hours for a state-of-the-art method.
  • Superior Accuracy: Outperformed five existing record linkage tools on a gold-standard dataset derived from real health-related databases in both accuracy and speed.

Scientific Applications:

  • Public Health Data Integration: Enables fast, large-scale linking of administrative health databases for integrated analyses.
  • Medical and Epidemiological Research: Supports analyses that require merging records across disparate health-related datasets.
  • Health Policy and Healthcare Delivery: Facilitates creation of comprehensive linked datasets to inform policy-making and healthcare delivery improvements.

Methodology:

Identification records are encoded into DNA sequences and aligned using the BLASTn algorithm.

Topics

Details

License:
GPL-3.0
Cost:
Free of charge
Tool Type:
command-line tool
Operating Systems:
Mac, Linux, Windows
Programming Languages:
Python
Added:
9/27/2022
Last Updated:
11/24/2024

Operations

Publications

Araujo JD, Santos-e-Silva JC, Costa-Martins AG, Sampaio V, de Castro DB, de Souza RF, Giddaluru J, Ramos PIP, Pita R, Barreto ML, Barral-Netto M, Nakaya HI. Tucuxi-BLAST: Enabling fast and accurate record linkage of large-scale health-related administrative databases through a DNA-encoded approach. PeerJ. 2022;10:e13507. doi:10.7717/peerj.13507. PMID:35846888. PMCID:PMC9281601.

PMID: 35846888
PMCID: PMC9281601
Funding: - Brazilian agency Fundação de Amparo à Pesquisa do Estado de São Paulo-FAPESP: 2018/14933-2 - FAPESP: 2019/27139-5