Alphabet
Alphabet identifies groups of amino acids that co-occur in columns of aligned protein sequences stored in BLOCKS+ format to define conserved substitution groups, construct eMOTIF sequence motifs, and detect column correlations via MCSeq.
Key Features:
- Column-wise analysis: Examines aligned protein sequences column-by-column to detect groups of amino acids that co-occur more frequently than expected from overall composition.
- Reduced alphabet formation: Creates a reduced alphabet by grouping conserved amino acid substitutions to simplify representation of sequence motifs.
- Empirical conservation identification: Identifies empirically conserved amino acid substitution groups across alignment columns.
- Conditional distribution matrix: Uses a conditional distribution matrix that extends pairwise substitution matrices, with rows representing amino acid distributions conditioned on specific substitution groups.
- Database processing: Converts BLOCKS and HSSP aligned-sequence databases into conditional distribution matrices for systematic examination of substitution groups.
- MCSeq correlation detection: Detects correlations between alignment columns using the MCSeq method.
- eMOTIF construction: Constructs sequence motifs (eMOTIFs) from the reduced alphabet and conserved substitution groups.
- Empirical findings: Has identified twenty conserved amino acid substitution groups in the BLOCKS and HSSP databases.
Scientific Applications:
- Sequence motif construction (eMOTIFs): Supports construction of sequence motifs that aid interpretation of protein function and interaction.
- Column correlation analysis (MCSeq): Enables detection of correlated columns to elucidate structural and functional relationships within proteins.
- Evolutionary and biochemical inference: Provides evidence of conserved biochemical properties and substitution patterns relevant to studies of protein evolution, structure, and function.
Methodology:
Alphabet builds conditional distribution matrices that extend pairwise substitution matrices, represents each conditioning group as a row of amino acid distributions at aligned positions, converts BLOCKS and HSSP data into these matrices, and examines each possible substitution group for empirical conservation, identifying twenty conserved groups.
Topics
Details
- Tool Type:
- command-line tool
- Operating Systems:
- Linux
- Programming Languages:
- C
- Added:
- 12/18/2017
- Last Updated:
- 12/11/2018
Operations
Publications
Wu TD and Brutlag DL. Discovering empirically conserved amino acid substitution groups in databases of protein families. Proc Int Conf Intell Syst Mol Biol. 1996; 4:230-40.