PDE
PDE extracts structured data from full-text PDF documents to support systematic literature reviews and meta-analyses.
Key Features:
- Table Extraction: The pdf2table function extracts tables directly from PDF files and exports them into Excel sheets.
- Keyword Search and Data Extraction: Performs high-throughput keyword searches across full-text PDF articles to identify occurrences of search terms.
- Adaptability and Customization: Allows customization of analysis parameters using regular expressions and employs machine learning algorithms to detect abbreviations related to search terms.
- Automated Categorization: Automatically categorizes full-text articles by relevance to predefined criteria, demonstrating up to 35% reduction in manual screening volume with 100% sensitivity under standard parameters.
- Metadata Export: Exports document metadata for downstream organization and referencing.
Scientific Applications:
- Systematic literature reviews and meta-analyses: Extracts tabular data and keyword-based evidence from multiple publications to support quantitative and qualitative synthesis.
- Pre-screening and prioritization: Prioritizes and filters articles for downstream manual review and captures minor results and details across large document sets.
Methodology:
Uses the pdf2table function for table extraction and Excel export, high-throughput keyword searching, regular expressions for pattern matching, and machine learning for abbreviation detection and article categorization.
Topics
Details
- License:
- GPL-3.0
- Cost:
- Free of charge
- Tool Type:
- library
- Operating Systems:
- Mac, Linux, Windows
- Programming Languages:
- R
- Added:
- 11/28/2021
- Last Updated:
- 11/28/2021
Operations
Publications
Stricker E, Scheurer ME. The PDF Data Extractor (PDE) Pre-screening Tool Reduced the Manual Review Burden for Systematic Literature Reviews by Over 35% Through Automated High-Throughput Assessment of Full-Text Articles. Unknown Journal. 2021. doi:10.1101/2021.07.13.452159.