Ensembl

Ensembl provides comprehensive, stable genome annotations and integrated sequence, gene, variation, regulatory, and comparative genomics data for analysis of chordate and other genomes.


Key Features:

  • Comprehensive genome annotation: Stable, automatic gene annotations integrating confirmed gene predictions with external data and functional genomics data, including genome-wide maps of protein–DNA interactions and regulatory builds.
  • Variation and orthology/paralogy annotations: Incorporates strain variation data and orthology/paralogy relationships inferred from gene trees.
  • Comparative genomics tools: Provides comparative genomics-based data mining and evolutionary analyses across supported species.
  • Genome resequencing support: Supports genome resequencing data and analysis for detailed genetic variation studies.
  • Software infrastructure: Maintains software infrastructure for handling large genomes and associated data analysis requirements.
  • Species coverage: Supports a wide array of chordate genomes including vertebrates, selected model organisms, and disease vectors, with additions such as marmoset, pig, zebra finch, lizard, gorilla, wallaby, platypus, horse, orangutan and various low-coverage mammalian species (release 56, September 2009, fully supported 51 species).
  • Integration with Ensembl Genomes Project: Collaborates with Ensembl Genomes to extend genomic resources to additional species.
  • Data distribution formats: Distributes data via downloadable flat files, an API, BioMart, and a genome browser.

Scientific Applications:

  • Genome annotation and gene function: Use integrated annotations and functional genomics data to infer gene structure and function across species.
  • Variation and population genetics: Analyze strain variation and resequencing data to study genetic variation and population-specific alleles.
  • Comparative genomics and phylogenomics: Infer orthology/paralogy relationships and perform evolutionary analyses using gene trees and comparative data mining.
  • Regulatory genomics: Map protein–DNA interactions and regulatory features to study transcriptional regulation and regulatory elements.
  • Cross-species studies: Compare genomes of vertebrates, model organisms and disease vectors to investigate conservation and divergence.

Methodology:

Stable, automatic annotation integrating confirmed gene predictions with external data sources; orthology/paralogy inference from gene trees; generation of genome-wide protein–DNA interaction maps and regulatory builds; comparative genomics-based data mining; and data distribution via downloadable flat files, an API and BioMart.

Topics

Collections

Details

Maturity:
Mature
Cost:
Free of charge
Tool Type:
web application
Operating Systems:
Linux, Windows, Mac
Programming Languages:
Java, SQL, Perl, C
Added:
6/11/2015
Last Updated:
4/17/2021

Operations

Publications

Birney E. Ensembl 2004. Nucleic Acids Research. 2004;32(90001):468D-470. doi:10.1093/nar/gkh038. PMID:14681459. PMCID:PMC308772.

Clamp M. Ensembl 2002: accommodating comparative genomics. Nucleic Acids Research. 2003;31(1):38-42. doi:10.1093/nar/gkg083. PMID:12519943. PMCID:PMC165530.

Hubbard T. Ensembl 2005. Nucleic Acids Research. 2004;33(Database issue):D447-D453. doi:10.1093/nar/gki138. PMID:15608235. PMCID:PMC540092.

Birney E. Ensembl 2006. Nucleic Acids Research. 2006;34(90001):D556-D561. doi:10.1093/nar/gkj133. PMID:16381931. PMCID:PMC1347495.

Cunningham F, Amode MR, Barrell D, Beal K, Billis K, Brent S, Carvalho-Silva D, Clapham P, Coates G, Fitzgerald S, Gil L, Girón CG, Gordon L, Hourlier T, Hunt SE, Janacek SH, Johnson N, Juettemann T, Kähäri AK, Keenan S, Martin FJ, Maurel T, McLaren W, Murphy DN, Nag R, Overduin B, Parker A, Patricio M, Perry E, Pignatelli M, Riat HS, Sheppard D, Taylor K, Thormann A, Vullo A, Wilder SP, Zadissa A, Aken BL, Birney E, Harrow J, Kinsella R, Muffato M, Ruffier M, Searle SM, Spudich G, Trevanion SJ, Yates A, Zerbino DR, Flicek P. Ensembl 2015. Nucleic Acids Research. 2014;43(D1):D662-D669. doi:10.1093/nar/gku1010. PMID:25352552. PMCID:PMC4383879.

Hubbard T. The Ensembl genome database project. Nucleic Acids Research. 2002;30(1):38-41. doi:10.1093/nar/30.1.38. PMID:11752248. PMCID:PMC99161.

Hubbard TJP, Aken BL, Beal K, Ballester B, Caccamo M, Chen Y, Clarke L, Coates G, Cunningham F, Cutts T, Down T, Dyer SC, Fitzgerald S, Fernandez-Banet J, Graf S, Haider S, Hammond M, Herrero J, Holland R, Howe K, Howe K, Johnson N, Kahari A, Keefe D, Kokocinski F, Kulesha E, Lawson D, Longden I, Melsopp C, Megy K, Meidl P, Ouverdin B, Parker A, Prlic A, Rice S, Rios D, Schuster M, Sealy I, Severin J, Slater G, Smedley D, Spudich G, Trevanion S, Vilella A, Vogel J, White S, Wood M, Cox T, Curwen V, Durbin R, Fernandez-Suarez XM, Flicek P, Kasprzyk A, Proctor G, Searle S, Smith J, Ureta-Vidal A, Birney E. Ensembl 2007. Nucleic Acids Research. 2007;35(Database):D610-D617. doi:10.1093/nar/gkl996. PMID:17148474. PMCID:PMC1761443.

Flicek P, Aken BL, Beal K, Ballester B, Caccamo M, Chen Y, Clarke L, Coates G, Cunningham F, Cutts T, Down T, Dyer SC, Eyre T, Fitzgerald S, Fernandez-Banet J, Graf S, Haider S, Hammond M, Holland R, Howe KL, Howe K, Johnson N, Jenkinson A, Kahari A, Keefe D, Kokocinski F, Kulesha E, Lawson D, Longden I, Megy K, Meidl P, Overduin B, Parker A, Pritchard B, Prlic A, Rice S, Rios D, Schuster M, Sealy I, Slater G, Smedley D, Spudich G, Trevanion S, Vilella AJ, Vogel J, White S, Wood M, Birney E, Cox T, Curwen V, Durbin R, Fernandez-Suarez XM, Herrero J, Hubbard TJP, Kasprzyk A, Proctor G, Smith J, Ureta-Vidal A, Searle S. Ensembl 2008. Nucleic Acids Research. 2007;36(Database):D707-D714. doi:10.1093/nar/gkm988. PMID:18000006. PMCID:PMC2238821.

Hubbard TJP, Aken BL, Ayling S, Ballester B, Beal K, Bragin E, Brent S, Chen Y, Clapham P, Clarke L, Coates G, Fairley S, Fitzgerald S, Fernandez-Banet J, Gordon L, Graf S, Haider S, Hammond M, Holland R, Howe K, Jenkinson A, Johnson N, Kahari A, Keefe D, Keenan S, Kinsella R, Kokocinski F, Kulesha E, Lawson D, Longden I, Megy K, Meidl P, Overduin B, Parker A, Pritchard B, Rios D, Schuster M, Slater G, Smedley D, Spooner W, Spudich G, Trevanion S, Vilella A, Vogel J, White S, Wilder S, Zadissa A, Birney E, Cunningham F, Curwen V, Durbin R, Fernandez-Suarez XM, Herrero J, Kasprzyk A, Proctor G, Smith J, Searle S, Flicek P. Ensembl 2009. Nucleic Acids Research. 2009;37(Database):D690-D697. doi:10.1093/nar/gkn828. PMID:19033362. PMCID:PMC2686571.

Flicek P, Aken BL, Ballester B, Beal K, Bragin E, Brent S, Chen Y, Clapham P, Coates G, Fairley S, Fitzgerald S, Fernandez-Banet J, Gordon L, Gräf S, Haider S, Hammond M, Howe K, Jenkinson A, Johnson N, Kähäri A, Keefe D, Keenan S, Kinsella R, Kokocinski F, Koscielny G, Kulesha E, Lawson D, Longden I, Massingham T, McLaren W, Megy K, Overduin B, Pritchard B, Rios D, Ruffier M, Schuster M, Slater G, Smedley D, Spudich G, Tang YA, Trevanion S, Vilella A, Vogel J, White S, Wilder SP, Zadissa A, Birney E, Cunningham F, Dunham I, Durbin R, Fernández-Suarez XM, Herrero J, Hubbard TJP, Parker A, Proctor G, Smith J, Searle SMJ. Ensembl's 10th year. Nucleic Acids Research. 2009;38(suppl_1):D557-D562. doi:10.1093/nar/gkp972. PMID:19906699. PMCID:PMC2808936.

Documentation

Downloads