21 August, 2009

Ensembl events in September 2009

The Summer holiday time is almost over, so the Ensembl trainers are getting up to speed again in preparation for Autumn, which will be the busiest time of year for them.

The coming month we will have the following Ensembl events:

7 Sep: Demo for the National Genetics Reference Lab, Manchester, UK
7-10 Sep: Ensembl module in the EBI roadshow at the University of Tartu, Tartu, Estonia
16-17 Sep: Browser workshop at the Erasmus MC Molecular Medicine Postgraduate School, Rotterdam, The Netherlands

For details about these and other upcoming events, please have a look at the complete list of Ensembl training events.

Do you want to attend an Ensembl workshop, but are we never coming to your part of the world? Why not see if you can host a workshop at your own university / institute? For more information have a look here or mail us.

If you are in the Western US, we also may still be able to include your university / institute in our Western US tour coming January. Interested? Please mail me for more details.

19 August, 2009

EnsemblFungi (beta) now available...

In addition to Release 2 of Ensembl Genomes early this month, EBI-EMBL would also like to announce the new arrival of Ensembl Fungi (beta).

Release includes:

  • 2 yeast genomes: Saccharomyces cerevisiae and Schizosaccharomyces pombe.
  • 7 Aspergillus genomes A.clavatus, A.flavus, A.fumigatus, A.niger, A.oryzae, A.terreus and Neosartorya fischeri.
If you have any comments or feedback please do not hesitate to contact us at helpdesk@ensemblgenomes.org.

07 August, 2009

BLAST downtime now over

To our BLAST users,

BLAST is online again, the maintenance is finished.

Thank you for your patience.

The Ensembl team.

We have scheduled some downtime for our BLAST database servers on Monday 10th August *Postponed to Tuesday, 11 Aug*, in order to install some security patches. We will also be trying out an upgrade path and pushing that out to the main BLAST server if it proves compatible. There will be a downtime of between 30 minutes and 12 hours, depending on how smoothly the maintenance goes. The rest of the Ensembl website should be unaffected.

We suggest that you avoid running BLAST or BLAT searches in the 48 hours preceding the maintenance work as we cannot guarantee that data will not be lost. We will update this post as soon as the maintanance has been finished.

Thank you for your patience,
The Ensembl Team

06 August, 2009

Ensembl Genomes: Release 2

EMBL-EBI is pleased to announce release 2 of Ensembl Genomes; extending Ensembl further across the taxonomic space!

Highlights for this release include:

  • 9 new genomes of Escherichia species; 4 new Bacillus and Streptococcus genomes; and additional genomes of Mycobacterium and Staphylococcus genera added to Ensembl Bacteria; taking the total bacteria species/strain count to 134.
  • Bacillus subtilis now represented using the re-annotation by Barbe et al. (Microbiology 155 2009, 1758-1775).
  • Comparative genomic analyses (bacteria, protist, metazoan and pan-taxonomic Compara databases) updated.
  • Ensembl 54 software
Keep an eye out for Release 3 in September, which will include plants and fungi...

21 July, 2009

West Coast Mirror

This year we've invested in our own mirror - maintained by us - on the west coast of the US. This was mainly because assessing the web return time for our users showed a consistent additional 3 to 4 seconds if you were lucky enough to live out on the west coast (worse still if you are in Australia!). Although we did alot last year to improve the general response time of our web pages (for example, compressing our CSS and Javascript down to single files for the whole site, so these are only loaded once and then cach'ed locally), the Ensembl site delivers alot of dynamic content - and nothing but getting closer to the users can help this.

You can reach the site directly at uswest.ensembl.org or alternatively there is a little "world" icon on the top right of the page which switches to the star-and-stripes when you're on the west coast. Having the mirror not only helps our users who are on the west coast but also provides resilience when our main site goes down. As we're responsibile for provisioning it in-sync with our main site (its part of our release process) this mirror will stay current with the main site.

In some sense the mirror should be a low cost "per user" for us having the mirror - if users go to the mirror, it means less load on the main site, and so it's really how we distribute the "web farm" that sits behind Ensembl geographically. However, there are overheads from hiring rack space in the US to making our own release cycle more complex. This means we will need to assess whether running a US mirror makes sense in the long term. Our instinct is yes, but we need hard data on this.

These things need time to pick up, but already we'd be interested in feedback on this - for US users, is this site faster for you - in particular for East coast people who we think are probably still best off on the main site. Does it change with time of day? For Pacific rim users - Japan, Singapore, Korea, Australia - is the west coast site snappier for you? We'll be putting in place our own monitoring schemes, but user feedback is always good...

20 July, 2009

Ensembl browser workshops Western US January 2010

The Ensembl Outreach team is currently looking into the possibility of giving some 1-day Ensembl browser workshops in the Western US in the period directly after the Plant and Animal Genome Conference in San Diego, which takes place from 9-13 January 2010. Hosting institutions would only have to pay the instructor's expenses (accommodation and subsistence) and would share the domestic travel costs, but we would not otherwise charge for the workshops.
For more information about our browser workshops, please have a look here.
Interested? Please contact me for more details at bert@ebi.ac.uk.

15 July, 2009

Release 55

The Ensembl project is pleased to announce release 55 of Ensembl. Highlights of this release are:

* New GRCh37 human assembly;
* Wallaby 2x genome;
* Ability to display uploaded data on individual chromosomes or whole karyotype.



For more information visit our What's New page.

14 July, 2009

Ensembl Regulatory Features - now Martable

Release 55 has lots of goodies - not least the new, coordinated, GRCh37 assembly (more on that later), but one addition is the Martability of Ensembl Regulatory Features. Regulatory features are on by default on Human and Mouse, and each gene has a specific page for the regulatory features (for example http://www.ensembl.org/Homo_sapiens/Gene/Regulation?g=ENSG00000139618). Regulatory Features are developing fast, and the Martability is bringing out the richer information in the functional genomics database - for example, the classification of features into "promoter", "gene associated" and "unclassified". Next release we're hoping to release a more graphical view for each feature, but the present of the regulatory features in Mart allows the large scale users - from Perl, Java, R or just plain-only tab delimited text - to use them.

We're expecting alot of development in this area - the addition of Mouse DNaseI sites has allowed us to develop a Mouse build, and of course, the ENCODE project which is now on line in production mode will provide a far richer, deeper, dataset to work against.

So - watch this space.

08 July, 2009

More about the world of Ensembl

Who is using us?

During the long flight from London to Seoul, where I'm now giving workshops, I had time to do some analysis.

We have recently changed the way we analyse traffic on our website. Urchin Software allows us to pinpoint access to our site with more accuracy. (In a previous post, I showed data where some domains were excluded in the representation).

This month's data (June 2009) is more comprehensive and can be shown with the following heatmap. Again, dark-coloured countries show more use than lightly shaded ones. We can now detect some low level traffic in the African continent missed in previous analyses.

Following a recent thread, I normalised these data taking into account population. (Did I mention how long the flight from London to Seoul actually is?) This puts the Netherlands as the country with the highest ratio of Ensembl users normalised by population, followed by Sweden, Switzerland, the United Kingdom, Finland, Belgium, Iceland, Germany, Denmark and Singapore (the first non European country in the list) pushing the United States (excluding traffic from the US mirror) to the 15th position.

On another matter, I'm glad to see that our user base in this country is well established. Apart from the obvious Seoul, we can also see access from Taejon, Kwangju, Taegu, Pusan, Inchon, Pohang and up to 29 locations in the country, and following these workshops we hope this will increase further.

As they say over here 안녕

Xosé

02 July, 2009

Go, go, go .... the new Ensembl ontology database and API

Starting with release 55 of Ensembl we provide an ensembl_ontology database. It replaces the older ensembl_go database which used to be loaded straight from the public table dumps provided by the Gene Ontology group (and hence wasn't really an Ensembl database to start with). The associated API is now part of the Ensembl Core API, which should make working with GO terms in Ensembl more straightforward than it was in the past. Available methods include, amongst others, fetching all parent or child terms of a given GO term and fetching all genes, transcripts or translations annotated with a given GO term.

More detailed documentation on both database and API can be found at ensembl/misc-scripts/ontology/README.

Credit for developing the ensembl_ontology database and API goes to Andreas Kahari of the Ensembl Software team.

26 June, 2009

Colourful gene trees


We have some news for the forthcoming ensembl release. We have added a few more display options for our gene trees. It will be possible to colour the background of the trees based on the taxonomy. It will be much easier to locate orthologues or paralogues in a given clade. For people who prefer more subtle colouring, they can choose to colour the branches instead.

It will also be possible to automatically collapse all the genes for a given clade. In the example shown in this figure, glires and diptera are collapsed. Moreover, the new version will also allow you to hide fish genes for instance or even all genes from low-coverage genomes.

All these options are configurable from the configure panel available through the 'configure page' link in the left panel.

25 June, 2009

Ensembl Workshops in July

The summer is a time for research institutions to focus on research (or go on holiday). This causes a decrease in our Ensembl Outreach and Training (at least, workshop-wise!) However, we are giving a few workshops in July.

2 July: Internal Developer's Course (let's see what the developers themselves have to say about browsing Ensembl!)
6-7 July: Browser modules in EBI roadshow, Poznan Summerschool of Bioinformatics, Poland
9-10 July: Browser workshop in Seoul, Korea (Korea National Institute of Health)
13 July: Browser workshop in Daejeon, Korea (KRIBB)
13 July: Browser workshop at Wellcome Trust Genome Campus, for campus employees
14-15 July: Browser workshop in Pusan, Korea (Pusan National University)

Don't forget, we have loads of training videos on our tutorials page (and YouTube channel!)

15 June, 2009

US Mirror

Ensembl is pleased to announce the release of its West Coast US mirror (uswest.ensembl.org). This is a full mirror of the current Ensembl 54 release. We are providing this mirror to improve performance for users in the US, particularly on the West coast. It includes full search, BioMart and BLAST support (BLAST searching is actually run at Sanger with results passed back to the mirror).

This mirror is managed directly by the Ensembl web team, and we will aim to update it along with the main site, to keep it current. Credit for gettting this mirror up goes to James Smith and Eugene Bragin from the web team, with support from the Sanger systems team, particularly Peter Clapham, John Nicholson and Dave Holland.

Future plans: We will improve the mirror in the near future by allowing users to switch between the main and mirror site. Currently, we do not suggest logging in to the mirror. All user data must be retrieved by the main site at the Wellcome Trust Genome Campus. Speed is optimal if login is not used, however this will be improved in the future.

Many thanks,
The Ensembl Team

14 June, 2009

A view of the world from Ensembl

Following our recent post updating you with feedback from our recent user survey, I wanted to share with you the following data. I've been data mining our web logs to get a picture of worldwide access to Ensembl (this is page impressions) and I put them on a heat map representing access to our browser (this doesn't include either BioMart or direct access of our public databases). I filtered out commercial IP addresses (i.e. all those .com, and .net) to simplify the analysis. Countries are shaded according to how much they accessed the Ensembl browser in the month of May. Dark countries use the Ensembl browser more, light countries, less. And this is what you get:

Heat map May 2009

Can you see your country here? If not, consider running a browser workshop in order to join our wide community of Ensembl users!

This map should help us to monitor the success of our workshops (at this moment I'm writing this from Venezuela which hopefully should appear in future heat maps!) Although there are still some gaps in the map, we are happy to see that Ensembl is used globally and our efforts are recompensed. Training leads to more in depth understanding of how to use the browser, and we think our worldwide workshops are helpful in using Ensembl to guide and enhance research.

Greetings from Venezuela, or as you would say here... ¡Hasta la próxima

Xosé

10 June, 2009

Thanks for your feedback!

Dear all,

Thanks to all of you who participated in our Ensembl browser survey, which closed at well over 150 respondents. We appreciate hearing from users worldwide, and our responses came from Europe, the Americas, Asia, India, New Zealand and Australia.

We will be able to move forward based on the feedback, in order to make the browser smoother and easier to use.

Some of the things we learned.

How often do you use Ensembl?


The majority of our users (83%) access the Ensembl browser multiple times per week. (See pie chart above). A similar survey run by Ensembl in 2007 reflected a much lower figure: only 32% of users browsed as frequently, with the majority of respondents accessing the browser once a month or less. Genome browsers appear to have moved from being tentative, browse-for-interest curiousities to comprehensive resources that can be integrated into daily research.

We were pleased to see our users are accessing many types of Ensembl data, from transcripts to variations to comparative genomics. Most respondents found the tabs in our new interface to be intuitive, so we won't change those! Many users scroll through views using the buttons on each page, so again, we will keep those in place.

More of you are finding the Pre! and Archive! sites helpful. And 65% of users voted that the most important aspect was good information.

Things we will work on?

In addition to custom data and user upload advances, we are focusing on graphical display of our comparative data. These improvements are already in our pipeline. We did get some new ideas from the survey, for example, some of our respondents could not find the Pre! or Archive! pages. We will think about how to make these links more obvious.

Speed of the browser could be faster. We will keep you posted on the Ensembl mirror on the US west coast.

Finally, more users are finding our help pages and videos, including a new YouTube channel. and we will be constantly adding to the tutorials page. If you have ideas about videos or tutorials you would like to see, please send your comments to our helpdesk. We find it immensely helpful to hear feedback from our user community.

Thanks again for the view of Ensembl from our users! We will be polling again next year.

The Ensembl Team

28 May, 2009

Give us your feedback!

Hello all,

We are curious as to how our users are finding the current Ensembl web browser. We opened a survey in the current Ensembl (version 54). The survey will close in one week's time (Friday, 5 June). Thank you to everyone who has already entered their feedback.

For users who have not yet replied to the survey, we ask that you spare 10 minutes or so of your time to do so. Please give us your thoughts and feedback by clicking on the link below:

http://tinyurl.com/cv67vs

The feedback centers on the web browser, specifically the new interface launched in Nov, 2008.

Many thanks for your time.

Regards,
The Ensembl Team

21 May, 2009

Ensembl Events in June 2009

For June we have the following Ensembl events:

4-5 June : Browser workshop at the University of Cambridge, Cambridge, UK
11 June: Demo for the National Genetics Reference Lab, Manchester, UK
15-16 June: Browser workshop at the Facultad de Ciencias de la Universidad de Los Andes, Mérida, Venezuela
18-19 June: Developers workshop at the Facultad de Ciencias de la Universidad de Los Andes, Mérida, Venezuela
22 June: Browser workshop for NHS Molecular Genetics Laboratories at Liverpool Women's Hospital, Liverpool, UK
24-26 June: Ensembl module in the Bioinformatics for Vascular Biology course at the EBI, Hinxton, UK
29 June: Browser workshop at the Leibniz Institute for Zoo and Wildlife Research, Berlin, Germany

For details about these and other upcoming events, please have a look at the complete list of Ensembl training events.

15 May, 2009

Pre, Archive, and Vega downtime

Next week, 18-24 May, an upgrade is scheduled that will affect the following sites:

Archive sites versions 48-53 (Dec 2007-May 2009) (Downtime will start Monday 18 May)

BLAST and User Upload on Pre! sites (Wednesday 20 May)

BLAST and User Upload on the Vega site

We ask our users to plan to use these sites after the upgrade, if possible.

Regards,
The Ensembl Team

14 May, 2009

Ensembl 55

We are currently working on our next release which is due at the end of June 2009 and will contain the following:

Data

Human GRCh 37
We will be releasing a new genebuild for human based on the lastest assembly GRCh37 from the Genome Reference Consortium. A preliminary version of this assembly is available now in Ensembl Pre! Due to the new assembly we will have:
  • Updated repeat masking
  • New probeset mappings
  • cDNA update
  • A new ensembl-vega merge delivering a new gene set
Wallaby
Ensembl 55 includes the 2X genome for Tammar Wallaby (Macropus eugenii), this will be a projection build similar to our other 2X species.

C. elegans
We will also include an import of the WormBase release WS200 database for C. elegans.

Anole lizard - A gene patch incorporating the gene set provided by Chris Ponting at Oxford University means that we have a new gene set for the green anole lizard (Anolis Carolinensis).

Mouse - The mouse cDNA alignments have been updated.

Zebrafinch - There will be an updated gene set for the 6X zebra finch genome.

Zebrafish - Non-coding RNAs will be added to the Zv8 zebrafish assembly and there will also be some changes to protein coding gene models and new repeats and expression patterns.

Core

Schema Changes
  • Patch to update versions (patch_54_55_a.sql). * Add the missing types to go_xref (patch_54_55_b.sql).
  • Add new table dependent_xref (will hold the dependencys for the xrefs, i.e. if an EMBL entry come from a uniprot entry this relationship will be in the table)( patch_54_55_d.sql).
  • Add new tables for alternative splicing/transcript events (patch_54_55_c.sql).
  • Add new column 'is_constitutive' to the exon table (patch_54_55_e.sql)

Xrefs
Xrefs will be run for Human, Macacca, Opossum, Chimp, Chicken, Dog and Mouse (including Fantom Xrefs).

Ontology database schema and tools
The ensembl_go_NN databases are no longer being built. Instead we are replacing this with the ensembl_ontology_NN database which may be connected to using the core API.

Assembly mapping
Some of the databases will contain mapping coordinates between current and previous assemblies:
  • human: mapping from current GRCh37 to NCBI36, NCBI35 and NCBI34
  • mouse: mapping from current NCBIM37 to NCBIM36, NCBIM35 and NCBIM34
Other changes
  • API support for alternative transcripts/splicing events will be added
  • API support for constitutive exons will be added
  • Deprecated API modules will be removed
  • All slices will be created using the new_fast method from the SliceAdaptor to improve performance
  • seq_region seq edit support will be added. Seq_edits can already be stored and retrieved but these were not used in getting the sequence data. This will be changed so that "_rna_edit" attributes in the seq_region_attrib table will be used and the sequence changed.
  • MySQL and FASTA dumps will be copied to Amazon Public Datasets project
  • Gene name and xref projections

Mart
  • New functional genomics mart * A new Probe section added to Ensembl mart
  • New ontology mart
  • Constitutive exon information will be re-added to Ensembl mart

Variation
  • There will be a new human variation database generated by mapping NCBI36 coordinates to GRCh37 (using dbSNP 129)
  • Illumina array data for SNP/CNV is to be added
  • Transcript variations for Zebrafish and Zebrafinch will be reculated to include information from the new gene sets
  • Schema change - added a call to get consequence_type
Functional genomics
  • Human Regulatory Build will be updated using the GRCh37 assembly
  • Probe alignment and transcript annotation for all species will migrate from the core datbases to the functional genomics databases, this includes Affymetrix, Illumina, Codelink and Phalanx
  • Schema change, an is_current filed is to be added to the coord_system table
Comparative genomics

Alignments - The new human assembly means that the following alignments will be regenerated:
  • 9 eutherian mammals EPO multiple alignments
  • 31 eutherian mammals EPO multiple alignments
  • 12 amniota vertbrates Pecan multiple alignments
  • 4 catarrhini primate EPO multiple alignments
  • Pairwise BLASTZ-NET alignments of human against each of the other 9 and 31 eutherian mammals
  • Additional pairwise BLASTZ-NET alignments will be run for human-opossum, human-platypus, human- chicken and human-wallaby
  • Translated BLAT-NET will be regenerated for human against fugu, X.tropicalis, C.intestinalis, C.savignyi, stickleback, medaka, chicken, zebrafish, tetraodon, zebrafinch and anole lizard

Synteny will be recalculated for: rat vs. huamn, chicken vs. human and human vs. macaque, dog, chimpanzee, platypus, opossum, mouse, orangutan, horse and cow

Homologies amd families
  • 50 way GeneTrees and homologies with new/updated genebuilds and assemblies
  • Clustering using hcluster_sg
  • Multiple Sequence Alignments using consistency-based MCoffee meta-aligner (mafftgins + muscle + kalign + probcons) and new exon-skipping aware "skipper" algorithm.
  • New 'putative gene split' and 'distant paralog' homology types
  • Pairwise gene-based dN/dS calculations for high coverage species pairs
  • Updated MCL families including all Ensembl transcript isoforms and newest Uniprot Metazoa
  • Multiple sequence alignments with MAFFT
  • Stable IDs for GeneTrees (ENSGT00550NNNNNNNNN) and MCL Families (ENSFM00550NNNNNNNNN).





12 May, 2009

SNPs and ancestral alleles


The new Ensembl release includes a new view for SNPs and other genomic variations. It shows the alignment of the polymorphic position together with 10 base pairs of sequence up- and downstream. The user can choose among all available multiple alignments. Polymorphic positions in the other species are also shown.

This is very useful for looking at ancestral alleles, especially in combination with our EPO alignments as they include the inferred ancestral sequence. Although dbSNP provide predicted ancestral alleles for human SNPs, these are based on the chimp sequence only. In several cases, the ancestral sequence inferred from the multiple alignment is in disagreement with the chimp sequence like in this example. Using multiple alignments gives better results and more confidence to the calls.