TCW
TCW performs comprehensive single and comparative transcriptome analyses by integrating DIAMOND sequence similarity searches against UniProt taxonomic databases, Gene Ontology annotation, ORF finding with hit information and 5th-order Markov models, differential expression via R scripts, homolog identification and clustering, and Ka/Ks codon-based analyses for de novo and reference transcriptomes.
Key Features:
- Single-Transcriptome Analysis (singleTCW): Accepts sequence files and optional count data and performs DIAMOND similarity searches against UniProt taxonomic databases, Gene Ontology annotation, ORF finding using hit information, 5th-order Markov models and ORF length, and interfaces with the R environment to run provided R scripts for differential expression analysis.
- Comparative Transcriptome Analysis (multiTCW): Operates on multiple singleTCW databases to assign homologous gene pairs, perform pairwise analyses such as Ka/Ks from codon-based alignments, apply clustering methods, and provide cluster-level statistical analysis and annotation.
- Clustering Methods: Supports bidirectional best hit, Closure, Best Hit, OrthoMCL, and user-supplied clustering methods for grouping homologs and orthologs.
- Sequence Similarity Search: Integrates the DIAMOND program for rapid searches against UniProt taxonomic databases with the option to specify alternative databases.
- ORF Finding: Identifies open reading frames using external hit information, 5th-order Markov models, and ORF length criteria.
- Differential Expression Analysis: Interfaces with the R environment to execute supplied R scripts implementing popular differential expression methods.
- Pairwise and Evolutionary Analyses: Computes codon-based alignments and Ka/Ks estimates for pairwise comparisons.
- Database Management: Uses a MySQL-backed database implementation and a Java-based application framework for data storage and computational workflows.
- Support for de novo Transcriptomes: Handles de novo assembled transcriptome inputs for non-model organism analyses.
Scientific Applications:
- Non-model organism transcriptomics: Functional annotation and ORF prediction for de novo assembled transcriptomes from non-model species.
- Differential expression studies: Comparative expression analysis across conditions or samples using count data and R-based DE methods.
- Comparative transcriptomics: Identification of homologous pairs and clustering across species or conditions to study conserved and divergent transcripts.
- Evolutionary analyses: Codon-based alignments and Ka/Ks estimation to assess selection on coding sequences.
- Functional annotation: Assignment of Gene Ontology terms and integration of UniProt-based similarity for gene function inference.
Methodology:
Uses DIAMOND for sequence similarity searches against UniProt taxonomic databases; assigns GO annotations; finds ORFs using hit information, 5th-order Markov models, and ORF length criteria; performs differential expression via provided R scripts in the R environment; stores and manages data in MySQL databases within a Java-based application; assigns homologous pairs, computes codon-based alignments for Ka/Ks, and applies clustering methods including bidirectional best hit, Closure, Best Hit, OrthoMCL, or user-supplied approaches.
Topics
Details
- License:
- GPL-3.0
- Programming Languages:
- R, Java, SQL
- Added:
- 11/14/2019
- Last Updated:
- 12/27/2020
Operations
Publications
Soderlund CA. Transcriptome computational workbench (TCW): analysis of single and comparative transcriptomes. Unknown Journal. 2019. doi:10.1101/733311.