TCW

TCW performs comprehensive single and comparative transcriptome analyses by integrating DIAMOND sequence similarity searches against UniProt taxonomic databases, Gene Ontology annotation, ORF finding with hit information and 5th-order Markov models, differential expression via R scripts, homolog identification and clustering, and Ka/Ks codon-based analyses for de novo and reference transcriptomes.


Key Features:

  • Single-Transcriptome Analysis (singleTCW): Accepts sequence files and optional count data and performs DIAMOND similarity searches against UniProt taxonomic databases, Gene Ontology annotation, ORF finding using hit information, 5th-order Markov models and ORF length, and interfaces with the R environment to run provided R scripts for differential expression analysis.
  • Comparative Transcriptome Analysis (multiTCW): Operates on multiple singleTCW databases to assign homologous gene pairs, perform pairwise analyses such as Ka/Ks from codon-based alignments, apply clustering methods, and provide cluster-level statistical analysis and annotation.
  • Clustering Methods: Supports bidirectional best hit, Closure, Best Hit, OrthoMCL, and user-supplied clustering methods for grouping homologs and orthologs.
  • Sequence Similarity Search: Integrates the DIAMOND program for rapid searches against UniProt taxonomic databases with the option to specify alternative databases.
  • ORF Finding: Identifies open reading frames using external hit information, 5th-order Markov models, and ORF length criteria.
  • Differential Expression Analysis: Interfaces with the R environment to execute supplied R scripts implementing popular differential expression methods.
  • Pairwise and Evolutionary Analyses: Computes codon-based alignments and Ka/Ks estimates for pairwise comparisons.
  • Database Management: Uses a MySQL-backed database implementation and a Java-based application framework for data storage and computational workflows.
  • Support for de novo Transcriptomes: Handles de novo assembled transcriptome inputs for non-model organism analyses.

Scientific Applications:

  • Non-model organism transcriptomics: Functional annotation and ORF prediction for de novo assembled transcriptomes from non-model species.
  • Differential expression studies: Comparative expression analysis across conditions or samples using count data and R-based DE methods.
  • Comparative transcriptomics: Identification of homologous pairs and clustering across species or conditions to study conserved and divergent transcripts.
  • Evolutionary analyses: Codon-based alignments and Ka/Ks estimation to assess selection on coding sequences.
  • Functional annotation: Assignment of Gene Ontology terms and integration of UniProt-based similarity for gene function inference.

Methodology:

Uses DIAMOND for sequence similarity searches against UniProt taxonomic databases; assigns GO annotations; finds ORFs using hit information, 5th-order Markov models, and ORF length criteria; performs differential expression via provided R scripts in the R environment; stores and manages data in MySQL databases within a Java-based application; assigns homologous pairs, computes codon-based alignments for Ka/Ks, and applies clustering methods including bidirectional best hit, Closure, Best Hit, OrthoMCL, or user-supplied approaches.

Topics

Details

License:
GPL-3.0
Programming Languages:
R, Java, SQL
Added:
11/14/2019
Last Updated:
12/27/2020

Operations

Publications

Soderlund CA. Transcriptome computational workbench (TCW): analysis of single and comparative transcriptomes. Unknown Journal. 2019. doi:10.1101/733311.

Links