PPfold

PPfold predicts consensus RNA secondary structure from multiple-sequence alignments by combining the pfold stochastic context-free grammar with phylogenetic analysis.


Key Features:

  • Parallelization and Multithreading: Distributes phylogenetic calculations and the inside-outside algorithm across multiple cores to reduce runtime for large RNA alignments.
  • Extended-Exponent Datatype: Employs an extended-exponent datatype to avoid floating-point underflow and enable accurate computation on large-scale alignments.
  • Pfold Model Implementation: Implements the pfold model that integrates stochastic context-free grammar with phylogenetic models for comparative secondary-structure prediction.
  • Integration of Structure-Probing Data: PPfold 3.0 incorporates flexible probabilistic models to integrate auxiliary data from structure probing experiments into single-sequence and alignment predictions.
  • Scalability to Large Datasets: Capable of predicting consensus structures for large alignments, including complete viral genomes and long genomic transcripts.
  • Comparable Accuracy: Reports improved single-sequence and alignment prediction accuracy competitive with RNAstructure.

Scientific Applications:

  • Comparative RNA secondary-structure prediction: Infers consensus secondary structures from multiple-sequence alignments for evolutionary and structural analyses.
  • Analysis of viral genomes and long transcripts: Predicts secondary-structure features across complete viral genomes and long genomic transcripts.
  • Integration with experimental probing data: Refines structure predictions by incorporating auxiliary data from structure probing experiments.
  • Identification of structural elements: Detects conserved known structural elements and novel RNA features within alignments.

Methodology:

Combines stochastic context-free grammar (pfold model) with phylogenetic analysis, applies the inside-outside algorithm, parallelizes phylogenetic calculations and the inside-outside algorithm across multiple cores, employs an extended-exponent datatype to prevent floating-point underflow, and uses flexible probabilistic models to integrate auxiliary data from structure probing experiments.

Topics

Details

Tool Type:
desktop application
Operating Systems:
Linux, Windows, Mac
Programming Languages:
Java
Added:
3/6/2015
Last Updated:
11/25/2024

Operations

Data Inputs & Outputs

RNA secondary structure prediction

Publications

Sükösd Z, Knudsen B, Værum M, Kjems J, Andersen ES. Multithreaded comparative RNA secondary structure prediction using stochastic context-free grammars. BMC Bioinformatics. 2011;12(1). doi:10.1186/1471-2105-12-103. PMID:21501497. PMCID:PMC3102635.

Sükösd Z, Knudsen B, Kjems J, Pedersen CN. PPfold 3.0: fast RNA secondary structure prediction using phylogeny and auxiliary data. Bioinformatics. 2012;28(20):2691-2692. doi:10.1093/bioinformatics/bts488. PMID:22877864.

Documentation