nf-encyclopedia
nf-encyclopedia processes Data-Independent Acquisition (DIA) proteomics raw mass spectrometry data within a NextFlow pipeline to produce peptide- and protein-level quantification for downstream statistical analysis.
Key Features:
- NextFlow orchestration: Implements the pipeline in NextFlow to orchestrate modular steps and enable parallel execution.
- Integration with MSConvert, EncyclopeDIA, and MSstats: Combines MSConvert for raw data conversion, EncyclopeDIA for peptide and protein identification/quantification, and MSstats for protein-level statistical inference.
- GPF chromatogram library support: Leverages gas phase fractionated (GPF) chromatogram libraries to enhance peptide detection and quantification while remaining functional without such libraries.
- Reproducibility across environments: Demonstrated reproducible execution and consistent quantification results across different computational environments.
- Improved protein-level quantification: Incorporates MSstats to enhance protein-level quantitative performance beyond EncyclopeDIA alone.
- Cloud scalability: Benchmarked for scalable, parallelized execution in cloud environments to accommodate large-scale DIA experiments.
Scientific Applications:
- Systems biology: Enables comprehensive proteome quantification for network and pathway analyses.
- Clinical proteomics: Supports quantitative proteomic studies aimed at clinical sample characterization and comparison.
- Biomarker discovery: Facilitates identification and quantification of candidate protein biomarkers from DIA datasets.
- Disease mechanism elucidation: Enables comparative proteomics to reveal protein-level changes associated with disease states.
- Therapeutic target identification: Provides quantitative evidence for candidate protein targets and their regulation.
Methodology:
The NextFlow pipeline converts raw mass spectrometry data using MSConvert, processes the data with EncyclopeDIA (optionally using GPF chromatogram libraries) for peptide and protein identification and quantification, and applies MSstats for protein-level statistical refinement.
Topics
Details
- License:
- Apache-2.0
- Cost:
- Free of charge
- Tool Type:
- workflow
- Operating Systems:
- Mac, Linux, Windows
- Programming Languages:
- Python, R
- Added:
- 1/6/2024
- Last Updated:
- 11/24/2024
Operations
Publications
Allen C, Meinl R, Paez JS, Searle BC, Just S, Pino LK, Fondrie WE. nf-encyclopedia: A Cloud-Ready Pipeline for Chromatogram Library Data-Independent Acquisition Proteomics Workflows. Journal of Proteome Research. 2023;22(8):2743-2749. doi:10.1021/acs.jproteome.2c00613. PMID:37417926.