BoaG
BoaG analyzes NCBI's non-redundant (NR) protein sequence database to support large-scale genomic analyses of protein sequence diversity, taxonomic origin, and functional annotation.
Key Features:
- Domain-Specific Language: A domain-specific language tailored for genomics that enables precise queries and analyses against NCBI's NR protein database.
- Shared Data Science Infrastructure: A shared infrastructure that aggregates NR-derived datasets and supports collaborative computational analyses.
- CD-HIT Clustering: Incorporates CD-HIT clustering to organize protein sequences at various sequence similarity levels.
- Efficient Query Performance: Processes large NR datasets efficiently to compute metrics such as average protein sequence length and to retrieve common taxonomic assignments and functional annotations.
- Annotation Redundancy Analysis: Leverages clustering information to identify annotation redundancy in NR, including reported redundancy at the 95% sequence similarity level.
Scientific Applications:
- Taxonomic Origin Identification: Determine the taxonomic origin of protein sequences within NCBI's NR database.
- Functional Annotation Exploration: Explore and summarize functional annotations associated with NR protein sequences.
- Data Quality and Redundancy Assessment: Assess data quality by quantifying annotation redundancy in NR.
Methodology:
Integrates protein sequence data, functional annotations, and taxonomic assignments from the NR database and applies CD-HIT clustering to manage sequence similarity levels, enabling efficient querying and analysis of large datasets.
Topics
Details
- Added:
- 1/18/2021
- Last Updated:
- 2/10/2021
Operations
Publications
Bagheri H, Dyer R, Severin A, Rajan H. Comprehensive Analysis of Non Redundant Protein Database. Unknown Journal. 2020. doi:10.21203/rs.3.rs-54568/v1.