QTL TableMiner++(QTM)

QTL TableMiner++ (QTM) extracts and semantically annotates quantitative trait locus (QTL) information from heterogeneous tables in plant science literature using Europe PMC as a primary data source and is implemented in Java to produce machine-readable, ontology-enriched QTL datasets.


Key Features:

  • Table mining capability: Specializes in extracting structured experimental details from tables in scientific publications, with emphasis on QTL mapping studies.
  • Keyword matching and ontology-based concept identification: Employs keyword matching to locate QTL-relevant tables and ontology-based concept identification to detect domain entities.
  • Normalization and classification: Normalizes tables using rules derived from captions, column headers, and footers, and classifies columns into descriptors, properties, and values based on headers and cell data types.
  • Abbreviation expansion: Expands abbreviations using the Schwartz and Hearst algorithm.
  • Semantic enrichment: Annotates extracted content with Crop Ontology, Plant Ontology, and Trait Ontology using the Apache Solr search platform.
  • Output formats: Stores processed information in a relational database (SQLite) and as text files (CSV).
  • Implementation and data source: Implemented in Java and uses the Europe PMC repository as the primary literature source.

Scientific Applications:

  • QTL mapping research: Produces machine-readable, ontology-enriched QTL datasets to support plant QTL mapping studies.
  • Meta-analysis and comparative studies: Facilitates meta-analyses and comparative studies by providing semantically interoperable QTL data.
  • Integrative genomics: Enables integrative genomics research through ontology-enriched, machine-readable QTL information.

Methodology:

QTM processes Europe PMC articles using keyword matching and ontology-based concept identification to locate QTL tables; normalizes tables by applying rules from captions, column headers, and footers; classifies columns into descriptors, properties, and values based on headers and cell data types; expands abbreviations via the Schwartz and Hearst algorithm; performs semantic enrichment with Crop Ontology, Plant Ontology, and Trait Ontology via Apache Solr; outputs results to SQLite and CSV; performance was evaluated with precision and recall against manually annotated corpora from open-access QTL mapping articles in tomato (Solanum lycopersicum) — 74.53% precision and 92.56% recall — and potato (S. tuberosum) — 82.82% precision and 98.94% recall.

Topics

Details

License:
Apache-2.0
Tool Type:
command-line tool
Operating Systems:
Linux
Programming Languages:
Java
Added:
7/31/2018
Last Updated:
11/25/2024

Operations

Data Inputs & Outputs

Other operations do not define inputs or outputs.

Publications

Singh G, Kuzniar A, van Mulligen EM, Gavai A, Bachem CW, Visser RG, Finkers R. QTLTableMiner++: semantic mining of QTL tables in scientific articles. BMC Bioinformatics. 2018;19(1). doi:10.1186/s12859-018-2165-7. PMID:29801439. PMCID:PMC5970438.

PMID: 29801439
PMCID: PMC5970438
Funding: - NWO via the Netherlands eScience Center: 27014204

Documentation