nf-encyclopedia

nf-encyclopedia processes Data-Independent Acquisition (DIA) proteomics raw mass spectrometry data within a NextFlow pipeline to produce peptide- and protein-level quantification for downstream statistical analysis.


Key Features:

  • NextFlow orchestration: Implements the pipeline in NextFlow to orchestrate modular steps and enable parallel execution.
  • Integration with MSConvert, EncyclopeDIA, and MSstats: Combines MSConvert for raw data conversion, EncyclopeDIA for peptide and protein identification/quantification, and MSstats for protein-level statistical inference.
  • GPF chromatogram library support: Leverages gas phase fractionated (GPF) chromatogram libraries to enhance peptide detection and quantification while remaining functional without such libraries.
  • Reproducibility across environments: Demonstrated reproducible execution and consistent quantification results across different computational environments.
  • Improved protein-level quantification: Incorporates MSstats to enhance protein-level quantitative performance beyond EncyclopeDIA alone.
  • Cloud scalability: Benchmarked for scalable, parallelized execution in cloud environments to accommodate large-scale DIA experiments.

Scientific Applications:

  • Systems biology: Enables comprehensive proteome quantification for network and pathway analyses.
  • Clinical proteomics: Supports quantitative proteomic studies aimed at clinical sample characterization and comparison.
  • Biomarker discovery: Facilitates identification and quantification of candidate protein biomarkers from DIA datasets.
  • Disease mechanism elucidation: Enables comparative proteomics to reveal protein-level changes associated with disease states.
  • Therapeutic target identification: Provides quantitative evidence for candidate protein targets and their regulation.

Methodology:

The NextFlow pipeline converts raw mass spectrometry data using MSConvert, processes the data with EncyclopeDIA (optionally using GPF chromatogram libraries) for peptide and protein identification and quantification, and applies MSstats for protein-level statistical refinement.

Topics

Details

License:
Apache-2.0
Cost:
Free of charge
Tool Type:
workflow
Operating Systems:
Mac, Linux, Windows
Programming Languages:
Python, R
Added:
1/6/2024
Last Updated:
11/24/2024

Operations

Publications

Allen C, Meinl R, Paez JS, Searle BC, Just S, Pino LK, Fondrie WE. nf-encyclopedia: A Cloud-Ready Pipeline for Chromatogram Library Data-Independent Acquisition Proteomics Workflows. Journal of Proteome Research. 2023;22(8):2743-2749. doi:10.1021/acs.jproteome.2c00613. PMID:37417926.

PMID: 37417926
Funding: - National Science Foundation: SBIR 2112191