aptardi

aptardi predicts sample-specific polyadenylation sites and refines transcript 3'-ends by integrating genome-aligned RNA-Seq, DNA nucleotide sequences, and an initial transcriptome.


Key Features:

  • Dual Data Integration: Integrates genome-aligned RNA-Seq data, DNA nucleotide sequences, and an initial transcriptome as inputs.
  • Machine Learning Paradigm: Employs a machine learning model to predict expressed polyadenylation sites, achieving approximately twofold higher precision and over threefold improved recall compared to standard transcriptome assemblers.
  • Transcript Refinement: Refines transcript 3'-ends based on predicted polyadenylation sites to reflect sample-specific expression.
  • Cross-Species Applicability: Trained on the Human Brain Reference RNA commercial standard and maintains high performance across diverse tissues and mammalian species.

Scientific Applications:

  • Gene regulation and post-transcriptional modification: Provides precise 3'-end annotations to support investigation of gene regulation and post-transcriptional modifications.
  • Alternative polyadenylation and transcript diversity: Enables analysis of alternative polyadenylation and resulting transcript diversity.
  • Tissue- and species-specific studies: Facilitates exploration of tissue-specific expression patterns and species-specific regulatory mechanisms.
  • Quantitation and differential expression: Produces refined transcript ends usable for quantitation and differential expression analyses.

Methodology:

Inputs comprise DNA nucleotide sequences, genome-aligned RNA-Seq data, and an initial transcriptome; a machine learning model is trained on these inputs to predict expressed polyadenylation sites; predicted sites are used to adjust transcript 3'-ends.

Topics

Details

License:
MIT
Tool Type:
workflow
Programming Languages:
Python
Added:
6/14/2021
Last Updated:
8/13/2021

Operations

Publications

Lusk R, Stene E, Banaei-Kashani F, Tabakoff B, Kechris K, Saba LM. Aptardi predicts polyadenylation sites in sample-specific transcriptomes using high-throughput RNA sequencing and DNA sequence. Nature Communications. 2021;12(1). doi:10.1038/s41467-021-21894-x. PMID:33712618. PMCID:PMC7955126.

PMID: 33712618
PMCID: PMC7955126
Funding: - U.S. Department of Health & Human Services | NIH | National Institute on Alcohol Abuse and Alcoholism: F31AA027430, R24AA013162 - U.S. Department of Health & Human Services | NIH | National Institute on Drug Abuse: P30DA044223

Links