SQANTI3
SQANTI3 performs quality control, curation, and structural and functional annotation of long-read transcript models from third-generation sequencing technologies such as Pacific Biosciences (PacBio) to support accurate isoform-level analyses.
Key Features:
- Integration with FIT pipeline: Functions as the first module of the Functional IsoTranscriptomics (FIT) pipeline to enable isoform-level analyses.
- Quality control and curation: Calculates quality descriptors for transcript models, splice junctions, and transcript ends to detect artifacts and replace erroneous sequences.
- Integrated functional annotation: Provides functional annotation of transcripts to support downstream functional iso-transcriptomics analyses.
- Comprehensive descriptor set: Implements 47 unique descriptors to evaluate transcript and preprocessing-pipeline quality and to inform filtering strategies that remove technical artifacts from sequencing approaches such as Pacific Biosciences (PacBio).
- Impact on quantification: Identifies and curates novel transcripts, including novel combinations of splice sites that produce new open reading frames (ORFs), thereby affecting short-read-based transcript quantification estimates.
- Functional insights and proteogenomics relevance: Characterizes enrichment of novel transcripts in metabolic and neural-specific functions and highlights challenges for proteogenomics detection of alternative isoforms in public proteomics databases.
Scientific Applications:
- Isoform discovery and validation: Enables discovery and validation of novel isoforms from long-read sequencing by identifying artifacts and curating transcript structures.
- Improved transcript quantification: Enhances the accuracy of short-read-based transcript quantification by specifying curated transcript models and filtering technical artifacts.
- Functional iso-transcriptomics analyses: Supports downstream functional analyses within the FIT pipeline through integrated functional annotation of isoforms.
- Proteogenomics target assessment: Assesses detectability of alternative isoforms for proteogenomics by characterizing novel ORFs and their representation in proteomics databases.
Methodology:
Calculates quality descriptors for transcript models, junctions, and transcript ends; applies a set of 47 descriptors to assess transcript and preprocessing-pipeline quality; develops filtering strategies to remove technical artifacts; and identifies and curates novel transcripts, replacing erroneous sequences.
Topics
Details
- License:
- GPL-3.0
- Maturity:
- Mature
- Cost:
- Free of charge
- Tool Type:
- command-line tool
- Operating Systems:
- Linux
- Programming Languages:
- Python
- Added:
- 7/7/2025
- Last Updated:
- 7/7/2025
Operations
Publications
Pardo-Palacios FJ, Arzalluz-Luque A, Kondratova L, Salguero P, Mestre-Tomás J, Amorín R, Estevan-Morió E, Liu T, Nanni A, McIntyre L, Tseng E, Conesa A. SQANTI3: curation of long-read transcriptomes for accurate identification of known and novel isoforms. Nature Methods. 2024;21(5):793-797. doi:10.1038/s41592-024-02229-2. PMID:38509328. PMCID:PMC11093726.
Tardaguila M, de la Fuente L, Marti C, Pereira C, Pardo-Palacios FJ, del Risco H, Ferrell M, Mellado M, Macchietto M, Verheggen K, Edelmann M, Ezkurdia I, Vazquez J, Tress M, Mortazavi A, Martens L, Rodriguez-Navarro S, Moreno-Manzano V, Conesa A. SQANTI: extensive characterization of long-read transcript sequences for quality control in full-length transcriptome identification and quantification. Genome Research. 2018;28(3):396-411. doi:10.1101/gr.222976.117. PMID:29440222. PMCID:PMC5848618.
Documentation
Downloads
- BinariesVersion: 5.5https://github.com/ConesaLab/SQANTI3/releases/download/v5.5/SQANTI3_v5.5.zipVersion 5.5 download link from github
- Downloads pagehttps://github.com/ConesaLab/SQANTI3/releases