Seq2Logo

Seq2Logo generates sequence logos that visualize position-specific amino acid information content from multiple sequence alignments (MSAs) to characterize binding motifs and conserved or variable residues.


Key Features:

  • Sequence Weighting: Applies sequence weighting techniques to mitigate redundancy effects in MSAs so overrepresented sequences do not dominate the logo.
  • Pseudo Counts: Uses pseudo counts to compensate for low observation numbers and provide balanced amino acid frequency estimates.
  • Two-Sided Representation: Represents both amino acid enrichment and depletion at each alignment position to capture conservation and avoidance patterns.
  • Flexible Input Formats: Accepts peptide sequences, multiple sequence alignments (MSAs), and Blast sequence profiles as input.
  • Output Options: Produces graphical sequence logos and a position-specific scoring matrix (PSSM) as quantitative output.

Scientific Applications:

  • Binding Motif Analysis: Visualizes position-specific amino acid preferences to define and compare binding motifs.
  • Active Site Characterization: Identifies conserved and depleted residues relevant to catalytic or binding sites in proteins.
  • Evolutionary Conservation Analysis: Assesses position-specific conservation and variability across homologous sequences.
  • Drug Design and Protein Engineering: Informs selection of conserved or variable residues for targeting or modification in therapeutic and engineering contexts.

Methodology:

Sequence weighting and pseudo counts are applied to MSAs to correct for redundancy and sparse observations; a two-sided representation quantifies amino acid enrichment and depletion per position; outputs generated include sequence logos and a position-specific scoring matrix (PSSM).

Topics

Details

License:
Other
Tool Type:
web application
Operating Systems:
Linux, Windows, Mac
Programming Languages:
Python
Added:
3/25/2017
Last Updated:
11/25/2024

Operations

Data Inputs & Outputs

Publications

Thomsen MCF, Nielsen M. Seq2Logo: a method for construction and visualization of amino acid binding motifs and sequence profiles including sequence weighting, pseudo counts and two-sided representation of amino acid enrichment and depletion. Nucleic Acids Research. 2012;40(W1):W281-W287. doi:10.1093/nar/gks469. PMID:22638583. PMCID:PMC3394285.

Documentation

Links