28 July, 2010

Ensembl Events in August 2010

Although it's Summer vacation time, there are still a few Ensembl events scheduled the coming month:

18 Aug: Ensembl module in the EURASNET Workshop on Structural Bioinformatics of RNA & RNP at the Adam Mickiewicz University, Poznan, Poland
29 Aug: Ensembl module at the ESF European Meeting on Next Generation Sequencing at Leiden University Medical Center, The Netherlands

For details about these and other upcoming events, please have a look at the complete list of Ensembl training events.

27 July, 2010

Changes to Ensembl Mailing Lists

In the near future, we will be changing the way that the ensembl-dev and ensembl-announce mailing lists are managed.

The move to new hardware and list management software will increase responsiveness on the lists (ensembl-dev in particular has suffered from slow deliveries for some time), and in time will also allow us to provide a searchable archive of past posts.

The list addresses will change with the move:

  • ensembl-dev@ebi.ac.uk becomes dev@ensembl.org
  • ensembl-announce@ebi.ac.uk becomes announce@ensembl.org
The new addresses are active now.

All subscribers to the old lists have been moved to the new lists. Posts to the old list addresses will be automatically forwarded for a short time, but please update your email address books and any spam filters to reflect the new address list.

Information about the new lists, including details on how to subscribe and unsubscribe, can be found at http://lists.ensembl.org/.

As before, dev@ensembl.org is an open list which any subscriber can post to. Posting to announce@ensembl.org is restricted to Ensembl people.

07 July, 2010

All sites up and running

All Ensembl sites are up and running - the outage referred to in the previous post only lasted for about 20 minutes.

02 July, 2010

Site problems - further details and temporary alternatives

As Anne mentioned, at approximately 03:00 UK time on 2 July, the Sanger Institute Data Centre suffered a power failure. The effects of this have taken taken the main Ensembl web site and other resources off line. We are working to get the services restored.


In the mean time, we have two US mirror sites available:

http://useast.ensembl.org is recently launched and provides genome browsing and search facility.

http://uswest.ensembl.org provides genome browsing, but search relies on the main site in the UK.

Please note that BLAST services require the main site at Hinxton and will not be available until the systems have been restored.

Thank you for your patience and we apologise for any inconvenience caused.

More power failures

We are currently experiencing further system problems caused by a generator failure at approximately 3am UK time. Our IT staff are working to restore service, and we apologise for any inconvenience caused.

Our US mirror, uswest.ensembl.org, is still available, although features that depend on Sanger services (user accounts, BLAST) will not be working.

24 June, 2010

Ensembl Genomes browser workshops

The Ensembl project has focused on vertebrate organisms for the past ten years. Since Jan 2009, a sister project has formed, Ensembl Genomes, extending analysis to invertebrates, protists, plants, fungi and bacteria.

We do now offer separate workshops on both Ensembl and Ensembl Genomes. Though the browsing is similar, specifics can vary. For example, the genes in most non-vertebrate genomes are determined using different rules than in vertebrate genomes.

So, if you are thinking of hosting a browser workshop and you know your audience will mainly consist of researchers working on non-vertebrate genomes, the Ensembl Genomes browser is the right choice for you!

For more detailed information on the Ensembl Genomes browser workshops please have a look here.

22 June, 2010

Power failure

We are currently experiencing a power failure which has taken the main Ensembl site offline. Our mirror site uswest.ensembl.org is still running so please use this until we have the main site back online.

21 June, 2010

Ensembl Developers Workshop at Genome Informatics 2010

Ensembl announces a workshop for developers that will take place at the Wellcome Trust Genome Campus, Hinxton, UK. The workshop runs in the week of the Genome Informatics conference, from Monday 13 September until midday Wednesday 15 September 2010, and will cover the same content as in previous years.

Ensembl developers will present sessions on how to create your own core database, including the loading of a genome assembly into a database and the running of simple analyses using the Ensembl genebuild pipeline.

As this workshop is aimed at developers, we will be exploring the underlying code-base of the Ensembl Software System. The workshop comprises presentations around Ensembl as well as practical sessions. We would like to point out that this is not a 'learning how to program in Perl' workshop.

For the successful completion of the workshop exercises, participants will be expected to have comprehensive experience in PERL programming and a background in object oriented programming techniques. Familiarity with Perl, a Unix/Linux environment, databases and structured query languages (MySQL) are needed to follow the workshop and the programming examples. Knowledge of the Ensembl core API is essential. The workshop is suitable for people who have access to small- to large scale compute resources, and who plan to run genome-wide analyses on large-scale data sets.

The workshop will be given by:

• Bronwen Aken (Ensembl Genome Annotation team)
• Jan-Hinnerk Vogel (Ensembl Genome Annotation team)

Additional presenters will attend depending on the number of participants registering.

At the end of this course, attendees will

• have a good understanding of Ensembl's annotation pipeline
• have hands-on experience with loading a genome assembly
• have hands-on experience with the annotation pipeline

There is no registration fee to attend this course for non-industry participants, but we ask you to register in advance because the number of attendees is limited by the size of the training venue.

Please note that the workshop runs before the conference, so don't forget to arrange additional accommodation for the workshop days.

The aim of the Ensembl Developers Workshop is to promote the academic use of the Ensembl Analysis and Ensembl pipeline system used by the Ensembl Genome Annotation Group.

To register your interest to attend, please send an email to Bert Overduin (bert@ebi.ac.uk), together with a brief outline of the relevance of this workshop to your work and your details.

The registration deadline for this workshop is August, 15th. Notification of places will be as soon as possible.

We will circulate an agenda amongst participants in due time.

Ensembl Events in July 2010

As usually July is one of the quietest months of the year, but we still have a few workshops:

1-2 July: Ensembl module in the EBI Bioinformatics Roadshow at the Telethon Institute of Genetics and Medicine (TIGEM), Naples, Italy
13 July: Browser workshop at the University of Glasgow, Scotland

For details about these and other upcoming events, please have a look at the complete list of Ensembl training events.

11 June, 2010

SNPs in high LD


A new section has been added to the variation page in Ensembl 58 that allows you to find other variations in strong linkage disequilibrium with the SNP you are viewing.

Clicking on "Linked variations" from the menu on the left hand side of the variation page takes you to a view like this one for rs1333049. Linkage disequilibrium values are calcuated on the fly and presented in one table for each population.

The table shows both r2 and D' values, along with the distance between the linked and current variations, any overlapping genes and any phenotypes associated with the linked variations. The table can be sorted by any of these columns by clicking on the column header (see previous post). The view is extensively configurable - clicking on Configure this page allows you to select populations to be displayed, change the distance over which linked variations are looked for, and filter the variations returned.

This view is currently only available for Ensembl Human, and is limited to variations with enough associated genotypes to calculate linkage disequilibrium values.

08 June, 2010

New Features to Aid Browsing

With the recent release of version 58, we are pleased to announce a few features designed to make genome browsing simpler. Have you noticed the search function in the individual tables? For example, in the Gene tab: "Variation Table", search for a variation ID.

Have a look at our sortable tables. In this example, we can sort by ID, Type, location, allele, source, or validation status. Use the arrows next to the column title to choose a new column to sort by. Searchable, and sortable, tables can also be found for orthologues, variations across individuals or strains, protein motifs and domains, and more.

If you are browsing a genomic region, you may have noticed that the Location tab: "Region in detail" view has a sliding zoom bar. Zoom out to view neighbouring genes and features, or zoom in to your favourite exon.

02 June, 2010

Zebrafish RNA-seq gene models

We have been developing a pipeline to build gene models using only RNA-seq data. For release 58 we have added a preliminary set of Zebrafish RNA-seq gene models with an intention to integrate this new source of evidence into a full genebuild soon.

Zebrafish transcriptome data from 9 tissues were used to build a set of genes and splice variants. For each loci we chose the variant with the highest read support to display, further details on the process are available here.
To display the genes, go to the Region in Detail, or Region Overview. Use the "Configure this page" button and select "RNASeq Genes" from the "Genes" menu. The "Supporting DNA Alignments" menu contains supporting exon and intron features from each of the nine tissues. Clicking on these features in Ensembl location pages shows a simple read count for the intron features and RPKM values for transcripts and exons, (reads per kilobase of model per million mapped reads, from Mortazavi Nature Methods 2008).

This is a first attempt at visualising tissue specific read depth and alternative splicing, which we hope to develop further in the future.

01 June, 2010

Ensembl Genomes Release 5

The Ensembl Genomes Project is pleased to announce release 5 of Ensembl Genomes.

The main highlights of this release are:

  • Software migration to Ensembl 58
  • Total of 6 new bacterial genomes for Escherichia/Shigella and Staphylococcus collections for Ensembl Bacteria; pairwise alignments added for all collections.
  • 2 new genomes, Phaeodactylum tricornutum and Thalassiosira pseudonana, for Ensembl Protists; pairwise alignments added.
  • Pristionchus pacificus genome from Wormbase and updates to the Drosophilia variation database added to Ensembl Metazoa.
  • Updates to the variation databases for A.thaliana, O.sativa japonica and V.vinifera in Ensembl Plants.
For further details regarding this release please visit:
http://ensemblgenomes.org/releases/release5_notes


28 May, 2010

GERP constrained elements via DAS

It is now possible to get the GERP constrained elements via the DAS protocol. For instance the DAS command to get all the GERP elements on the BRCA2 gene (Human chr 13: 32889611-32973347) is:
http://www.ensembl.org/das/Homo_sapiens.GRCh37.constrained_element/features?segment=13:32889611,32973347

By default, you obtain both the constrained derived from our 16-way amniote alignments and the 33-way placental mammals ones (these include all the low-coverage genomes). You can filter the elements you want by using the argument type:
http://www.ensembl.org/das/Homo_sapiens.GRCh37.constrained_element/features?segment=13:32889611,32973347;type=33_eutherian_mammals

Read more on DAS or on the multiple alignments and constrained elements.

25 May, 2010

Ensembl Events in June 2010

In June we will have the following Ensembl events:

2-4 June: Ensembl module in the EBI Bioinformatics Roadshow at Charles University, Prague, Czech Republic
10 June: Browser workshop at Gothenburg University, Sweden
11 June: Browser workshop at Gothenburg University, Sweden (satellite of the European Human Genetics Conference 2010)
12-15 June: Presentation at the European Human Genetics Conference 2010
14-18 June: Ensembl module in the Joint EBI - Wellcome Trust Bioinformatics Summer School, Hinxton, UK
22-23 June: Browser workshops at the German Cancer Research Center (DKFZ), Heidelberg, Germany
23 June: Browser workshop at Trinity College Dublin, Ireland
25 June: Browser workshop at the Ludwig-Maximilians University, Munich, Germany
30 June - 1 July: Browser workshop at the University of Ljubljana, Slovenia

For details about these and other upcoming events, please have a look at the complete list of Ensembl training events.

24 May, 2010

Which whole-genome multiple aligner? Pecan!

Comparing and assessing the quality of whole-genome multiple alignments is a difficult task. In the protein world, many mathematical models are available. They are based on synonymous and non-synonymous substitutions and the physicochemical similarities among the aminoacids. None of this can be applied to non-coding sequences, the 99% of the human genome.

There are two main trends for whole-genome alignments. Authors have either used genomic features like ancestral repeats or have developed phylogenetic models to generate synthetic sequences for which the "real" alignment is known.

Two articles have been published recently, one proposing a new method based on artificial sequences (Kim & Sinha, BMC Bioinformatics 2010, 11:54) and the other one looking at the coverage, agreement and accuracy of the alignments in the ENCODE pilot regions (Chen & Tompa, Nature Biotechnology 2010, doi:10.1038/nbt.1637).

According to both studies, Pecan is the strongest contender, showing the clear advantage of using a consistency-based approach (see Paten et al., Genome Res. 2008, 18:1814-28) to align the sequences.

21 May, 2010

Musing about Fish Alignments.

Release 57 saw the release of a 5-way EPO alignment across the telost fish - Zebrafish, Stickleback, Medaka, Tetraodon and Fugu. Just recently I've spent some time browsing through them. They are very interesting, with the ancestral duplication in Fish showing more complex homology relationships than in mammals. Here's a nice, clean example

Simple Fish multiple alignments

Here is a far more complex region, when ENREDO has clearly picked up the ancestral duplication but is struggling to make it colinear across the entire region

Complex Fish multiple alignments

One thing which I don't think most people appreciate is the incredible phylogenetic depth in the telost linage. In terms of "millions of years of evolution" or "sequence divergence" actually the deepest splits in the telosts - such as ZebraFish to Stickleback - are almost as deep as telosts to mammals - certainly deeper than birds to mammals. So this is asking alot to find good, clearly co linear stretches, in particular when you think of the draft nature of these genomes.

Chatting to Javier, it might be much better to also look at a 4-way EPO on the "Stickleback" side of the telost linage, in other words, Medaka/Stickleback/Fugu/Tetraodon. This I think will come together better (in fact most of the "nice" regions in the 5-way EPO are actually regions without ZebraFish) and we might be able to look at taking that ancestral chromosome ordering and perhaps sequence in comparisons to Zebrafish.

14 May, 2010

Ensembl, Xfam and HMMs

This week, Albert Vilella and myself participated in the Xfam consortium meeting. The meeting focussed on protein, domains and ncRNA classification, and on the new developments of the HMMER package.

Although Ensembl is not part of Xfam, we share many interests. We are getting increasingly interested in the use HMMER models, especially since the release of HMMER3.0. Also, in the forthcoming release (version 58), Ensembl will provide gene trees for ncRNAs. Most of these ncRNA genes are annotated using Rfam models.

Stay tuned for more!

11 May, 2010

Neandertal Genome Browser

In collaboration with the Neandertal Genome Project, we have created an Ensembl-style browser of the Neandertal data available at http://projects.ensembl.org/neandertal. A draft sequence of the Neandertal genome was published in the May 7 issue of Science.

The Neandertal browser includes the ability to visualise the Neandertal data using the new Resembl code developed in collaboration with Illumina. The Resembl code will be introduced in the 1000 Genomes browser later this month and in Ensembl over the summer.

Data include:
- Neandertal sequencing reads from all 6 Neandertal fossils
- Neandertal contigs/consensus from all individuals combined
- Modern human sequencing reads to put the divergence of the Neandertal genomes into perspective
- Selective sweep scan to detect positive selection in early modern humans
- A catalog of changes consisting of Neandertal alleles for positions of non-synonymous difference between human and chimpanzee

Full details of the data types and instructions for using our new display tools are available on the data information page.

Links are also provided from the Neandertal Browser home page to the raw sequence data stored at the EBI for the Neandertal genome project and the modern human genome data.


Further information about the project is available from the project page at Max Planck Institute for Evolutionary Anthropology, from the genome paper and from other companion papers in the same issue of science.

We thank Janet Kelso, Ed Green and Udo Stenzel at the MPI for assistance and Eugene Kulesha at the EBI for work to create the Neandertal browser.

06 May, 2010

SNPedia in Ensembl


Ensembl is always extending the variation pages to include more information. Did you know that the latest data from SNPedia is now available?

SNPedia is a wiki-style resource for human genetics with public annotation of over 11,000 SNPs, released under a Creative Commons style license. We have integrated it into Ensembl, so you can view these SNP reports along with our other information including variations, genotype and allele frequencies from dbSNP, and SNPs from other sources including UniProt, Affymetrix and Illumina chipsets and phenotype annotations from several genome-wide association studies.

You need to configure the page to view SNPedia. From the variation page, e.g. rs1333049, click on "Configure this page" and then click on "External Data" to select SNPedia to appear in the left hand side menu of all variation pages via DAS. As this information comes directly from SNPedia via DAS it is always up-to-date.