ProBias
ProBias identifies and quantifies compositionally biased segments (CBS) in protein sequences to support analysis of amino-acid enrichment relevant to molecular function, evolution, and disease.
Key Features:
- User-defined bias detection: Allows specification of amino-acid sets for targeted detection of compositional bias.
- Statistical significance estimation: Applies discrete scan statistics to estimate the statistical significance of detected CBS and implements multiple-testing correction.
- Global compositional bias consideration: Accounts for global (sequence-wide) compositional bias when evaluating local clusters to distinguish local enrichment from background composition effects.
- Sensitivity and precision: Detects subtle deviations from a random-independence background model to identify potentially functional CBS in complex sequences.
Scientific Applications:
- Functional annotation: Highlights candidate regions or domains in uncharacterized proteins based on significant compositional enrichment patterns.
- Disease association studies: Supports linking protein groups and biological processes to characteristic CBS types to inform hypotheses about disease mechanisms.
- Evolutionary analysis: Enables large-scale assessment of CBS conservation across protein families and lineages.
Methodology:
Applies discrete scan statistics to detect CBS, quantifies statistical significance with multiple-testing correction, and incorporates global (sequence-wide) compositional bias when evaluating local enrichment.
Topics
Details
- Tool Type:
- command-line tool
- Operating Systems:
- Linux, Windows, Mac
- Programming Languages:
- Perl
- Added:
- 12/18/2017
- Last Updated:
- 12/10/2018
Operations
Publications
Kuznetsov IB, Hwang S. A novel sensitive method for the detection of user-defined compositional bias in biological sequences. Bioinformatics. 2006;22(9):1055-1063. doi:10.1093/bioinformatics/btl049.
Kuznetsov IB. ProBias: a web-server for the identification of user-specified types of compositionally biased segments in protein sequences. Bioinformatics. 2008;24(13):1534-1535. doi:10.1093/bioinformatics/btn233.