PhyloCSF

PhyloCSF identifies protein-coding regions in multi-species nucleotide sequence alignments by analyzing evolutionary signatures characteristic of conserved coding sequences.


Key Features:

  • Codon Substitution Frequencies (CSF): Quantifies codon substitution patterns to distinguish coding from non-coding sequence based on Codon Substitution Frequencies.
  • Substitution-pattern analysis: Evaluates high frequencies of synonymous codon substitutions and conservative amino acid changes and low frequencies of missense and nonsense substitutions.
  • Multi-species alignments: Operates on multi-species nucleotide sequence alignments to detect conserved coding signatures across taxa.
  • Whole-genome prediction tracks: Produces genome-wide PhyloCSF prediction tracks for species including human, mouse, chicken, fly, worm, and mosquito.
  • Machine learning integration: Integrates machine learning to predict and prioritize novel conserved protein-coding regions and high-scoring candidates for downstream curation.
  • Annotation evidence support: Supports interpretation with comparative annotation and evidence types such as single-nucleotide variant patterns indicative of purifying selection, exon splicing from GENCODE transcripts, and mass spectrometry translation evidence.

Scientific Applications:

  • Transcript classification: Differentiates protein-coding and non-coding RNAs among novel transcript models from high-throughput transcriptome sequencing.
  • Gene annotation expansion: Identified over 1,000 high-scoring human PhyloCSF regions, leading to the confident addition of 144 new conserved protein-coding genes to the GENCODE gene set and discovery of additional coding regions in 236 previously annotated protein-coding genes and 169 pseudogenes.
  • Comparative annotation of ORFs: Enables comparative annotation across vertebrate genomes to eliminate spurious open reading frames (ORFs) and distinguish coding regions from pseudogenes.
  • Refinement of disease-associated loci: Reclassified 118 GWAS variants previously annotated as non-coding as protein-altering, informing interpretation of genetic contributions to disease.

Methodology:

Analyzes evolutionary signatures in multi-species nucleotide sequence alignments using Codon Substitution Frequencies (CSF) by quantifying synonymous, conservative amino acid, missense, and nonsense substitution patterns, generates whole-genome PhyloCSF prediction tracks, and integrates machine learning to predict novel conserved protein-coding regions.

Topics

Details

License:
AGPL-3.0
Programming Languages:
OCaml
Added:
11/14/2019
Last Updated:
11/24/2024

Operations

Publications

Mudge JM, Jungreis I, Hunt T, Gonzalez JM, Wright JC, Kay M, Davidson C, Fitzgerald S, Seal R, Tweedie S, He L, Waterhouse RM, Li Y, Bruford E, Choudhary JS, Frankish A, Kellis M. Discovery of high-confidence human protein-coding genes and exons by whole-genome PhyloCSF helps elucidate 118 GWAS loci. Genome Research. 2019;29(12):2073-2087. doi:10.1101/gr.246462.118. PMID:31537640. PMCID:PMC6886504.

PMID: 31537640
PMCID: PMC6886504
Funding: - National Human Genome Research Institute of the National Institutes of Health under Award: U41HG007234 - Wellcome Trust: 208349/Z/17/Z, WT108749/Z/15/Z - National Institutes of Health: R01 HG004037 - Swiss National Science Foundation: PP00P3_170664 - National Human Genome Research Institute: U24HG003345

Links