BoaG

BoaG analyzes NCBI's non-redundant (NR) protein sequence database to support large-scale genomic analyses of protein sequence diversity, taxonomic origin, and functional annotation.


Key Features:

  • Domain-Specific Language: A domain-specific language tailored for genomics that enables precise queries and analyses against NCBI's NR protein database.
  • Shared Data Science Infrastructure: A shared infrastructure that aggregates NR-derived datasets and supports collaborative computational analyses.
  • CD-HIT Clustering: Incorporates CD-HIT clustering to organize protein sequences at various sequence similarity levels.
  • Efficient Query Performance: Processes large NR datasets efficiently to compute metrics such as average protein sequence length and to retrieve common taxonomic assignments and functional annotations.
  • Annotation Redundancy Analysis: Leverages clustering information to identify annotation redundancy in NR, including reported redundancy at the 95% sequence similarity level.

Scientific Applications:

  • Taxonomic Origin Identification: Determine the taxonomic origin of protein sequences within NCBI's NR database.
  • Functional Annotation Exploration: Explore and summarize functional annotations associated with NR protein sequences.
  • Data Quality and Redundancy Assessment: Assess data quality by quantifying annotation redundancy in NR.

Methodology:

Integrates protein sequence data, functional annotations, and taxonomic assignments from the NR database and applies CD-HIT clustering to manage sequence similarity levels, enabling efficient querying and analysis of large datasets.

Topics

Details

Added:
1/18/2021
Last Updated:
2/10/2021

Operations

Publications

Bagheri H, Dyer R, Severin A, Rajan H. Comprehensive Analysis of Non Redundant Protein Database. Unknown Journal. 2020. doi:10.21203/rs.3.rs-54568/v1.

Links