Showing posts with label comparative genomics. Show all posts
Showing posts with label comparative genomics. Show all posts

21 September, 2010

Changing to LASTZ


People following the declarations of intentions for the next release (these are sent to announce@ensembl.org) may have noticed that we are releasing LASTZ pairwise alignments instead of BLASTZ ones. LASTZ is written by Bob Harris from the Penn State University as a replacement of BLASTZ, as BLASTZ is now considered obsolete (read the announcement).

This is the first release where we use LASTZ for the new alignments. We will update the previous alignments in the following releases.

24 May, 2010

Which whole-genome multiple aligner? Pecan!

Comparing and assessing the quality of whole-genome multiple alignments is a difficult task. In the protein world, many mathematical models are available. They are based on synonymous and non-synonymous substitutions and the physicochemical similarities among the aminoacids. None of this can be applied to non-coding sequences, the 99% of the human genome.

There are two main trends for whole-genome alignments. Authors have either used genomic features like ancestral repeats or have developed phylogenetic models to generate synthetic sequences for which the "real" alignment is known.

Two articles have been published recently, one proposing a new method based on artificial sequences (Kim & Sinha, BMC Bioinformatics 2010, 11:54) and the other one looking at the coverage, agreement and accuracy of the alignments in the ENCODE pilot regions (Chen & Tompa, Nature Biotechnology 2010, doi:10.1038/nbt.1637).

According to both studies, Pecan is the strongest contender, showing the clear advantage of using a consistency-based approach (see Paten et al., Genome Res. 2008, 18:1814-28) to align the sequences.

11 March, 2010

Orthologues on the trees


We have added a little trick for orthology-lovers. Starting from the orthologues page, you can choose to switch to the GeneTree. This will highlight the orthologue of interest, as well as the ancestral node that relates both genes.

Another useful feature added in Ensembl 57 is the possibility to display a set of genes (up to 10) using the Multi-Species view. Click on an internal node and select the "Jump to Multi-species view" option. This will show each of these genes in their respective genomic location, with genomic alignments when available.

26 June, 2009

Colourful gene trees


We have some news for the forthcoming ensembl release. We have added a few more display options for our gene trees. It will be possible to colour the background of the trees based on the taxonomy. It will be much easier to locate orthologues or paralogues in a given clade. For people who prefer more subtle colouring, they can choose to colour the branches instead.

It will also be possible to automatically collapse all the genes for a given clade. In the example shown in this figure, glires and diptera are collapsed. Moreover, the new version will also allow you to hide fish genes for instance or even all genes from low-coverage genomes.

All these options are configurable from the configure panel available through the 'configure page' link in the left panel.

21 January, 2009

How to get all the orthologous genes between two species

Many users ask us about how to download data from ensembl. Usually, the answer is using BioMart. Comparative genomics data are also available in the standard Mart for your favorite species. For instance to get all the human-mouse orthologs, one can select the human dataset, filter all the genes with no mouse orthologs and choose to output the mouse orthologs for all the resulting genes.

Here is how to get these data in 10 simple steps
1. Go to: http://www.ensembl.org/biomart/martview
2. Choose "Ensembl 52"
3. Choose "Homo sapiens genes (NCBI36)"
4. Click on "Filters" in the left menu
5. Unfold the "MULTI SPECIES COMPARISONS" box, tick the "Homolog filters" option and choose "Orthologous Mouse Genes" from the drop-down menu.
6. Click on "Attributes" in the left menu
7. Click on "Homologs"
8. Unfold the "MOUSE ORTHOLOGS" box and select the data you want to get (most probably the gene ID and maybe the orthology type as well).
9. Click on the "Results" button (top left)
10. Choose your favorite output

Here is the preview of the results:



Other people may prefer to use our Compara Perl API or get the data directly from the Compara DB. These options are also available.

11 December, 2008

GeneTrees: how do I read them? And can I view alignments using Jalview?

If you have clicked on the GeneTree link in Ensembl (for example, the gene tree for IL2), you may have noticed that we have a new way of displaying large GeneTrees. This time, if you have a large gene family with lots of genes that you want to look at, you won't need to ask the Miami Dolphins to let you plug your laptop into their huge screen...


This new feature in EnsemblCompara is called collapsible subtrees and allows for more compact, summarized views of interesting gene families like PAX2/PAX5/PAX8:

http://www.ensembl.org/Homo_sapiens/Gene/Compara_Tree?g=ENSG00000075891

If you check the legend at the bottom, you will see that "blue triangles" correspond to collapsed subtrees that have within-species paralogs of your gene. If you want to see all the within-species paralogs expanded, you can click on the option "View paralogs of current gene". You can even set that as a default if you want in the "Configure this page" options.

Jalview is a great way to view protein alignments in the tree. And were is my Jalview link now? Click on any internal node (square) in the tree, and be able to visualize the alignment (or subalignment) with the new Jalview applet by clicking on the Jalview link. You have to have Java installed though, or the link won't show. The two Jalview windows that pop up are one, the protein alignment and the other, the underlying TreeBeST tree. You can now use Jalview's sorting feature to sort your sequences according to the tree with: Calculate->Sort->By Tree Order->URL. Having the tree associated to the alignment allows for a more phylo-centric visualization of sequence conservation: if you click at a point in the tree, a red vertical line will appear that divides the alignment into different groups. If you choose Colour->Percentage Identity, the shades of blue will be relative to the subgroups in your tree (e.g., fish versus placental mammals). This is also useful to spot segments in the alignment that don't look that good, or gaps created in a subpart that can now be collapsed in the subalignment (Edit->Remove Empty Columns), or sequences that stand out as long branches in the alignment (View->Overview Window).


For even more tree funkiness, you can use PhyloWidget to visualize our NHX trees. Use our NHX tree ("Configure this page->Output for normal tree->NHX->Save and Close->Gene Tree(text)") to copy+paste the representation of the GeneTree into Phylowidget, with duplication/speciation events (red/blue), bootstrap values (greyscale) and taxonomy levels "View->Rendering->Show clade labels". Then use the "Zoom in/Zoom out" features, or clicking on an internal node, the "Tree Edit->collapse", and specially the "View->Branch lenghts [x]" and the "View->Layout->Options->Branch Scaling" options.


We hope these new features will help you in your research. We have some new ideas that we are currently testing to visualize even more phylogenetic information, and help make better judgement on the orthology and paralogy relationships in our EnsemblCompara GeneTrees. Stay tuned for more updates!

15 February, 2008

Orangutans colonise our gene trees

Orangutan on an Ensembl gene tree
(you wouldn't expect orangutans out of the trees, would you?)

Ensembl 49 will contain good news on the comparative genomics side. Apart from the new whole-genome multiple alignments for which we can now handle segmental duplications and infer ancestral sequences (see Ewan's post on the 20th of January), two new species will be available, namely horse and orangutan.

We are especially excited about the new orangutan genome as it is a key species in the primate lineage, in between the human, chimp and gorilla group and the Old World monkeys. Its inclusion in our gene trees will result in a better resolution of the phylogeny of the primate genes.