Friday, April 25, 2014

Systematic Detection of Internal Symmetry in Proteins Using CE-Symm

In our latest paper, "Systematic Detection of Internal Symmetry in Proteins Using CE-Symm", we are taking a look at how internal symmetry in proteins is related to protein function. A large number of proteins have symmetry not only in their biological assemblies, but also within their tertiary structures. To investigate the question of how internal symmetry evolved, how symmetry and function are related, and the overall frequency of internal symmetry, we developed a new algorithm that can detect pseudo-symmetry within tertiary structures of proteins. Our results indicate that more domains are pseudo-symmetric than previously estimated. We establish a number of recurring types of symmetry–function relationships and describe several characteristic cases in detail.

Read more over at JMB.


Several protein domains with internal symmetry that CE-Symm detects. Coloring is by symmetry unit.




Tuesday, March 25, 2014

BioJava 3.0.8 released

 BioJava 3.0.8 was released on March 25th 2014 and is available from
BioJava maven repository at http://www.biojava.org/download/maven/

This release would not have been possible without contributions from
13 developers, thanks to all for their support!

BioJava 3.0.8 includes a lot of new features as well as numerous bug fixes and improvements.

New Features:
  •  new Genbank writer
  •  new parser for Karyotype file from UCSC
  •  new parser for Gene locations from UCSC 
  •  new parser for Gene names file from genenames.org
  •  new module for Cox regression code for survival analysis
  •  new calculation of accessible surface area (ASA)
  •  new module for parsing .OBO files (ontologies)
  •  improved representation of SCOP and Berkeley-SCOP classifications
 
For a detailed comparison see here:

For the next release we are planning some refactoring and removal of code that has been deprecated for a long time. As such the next release will be named 3.1.0.

About BioJava:

BioJava is a mature open-source project that provides a framework for
processing of biological data. BioJava contains powerful analysis and
statistical routines, tools for parsing common file formats, and
packages for manipulating sequences and 3D structures. It enables
rapid bioinformatics application development in the Java programming
language.

Happy BioJava-ing,

Andreas

Wednesday, September 4, 2013

RCSB PDB September 2013 release

This week we released the latest RCSB PDB web site update
 The main new features are:

- New tools to search for drugs and drug targets
- Improved interface for 3D visualisation using Jmol/JSmol
- An update to the representation of protein symmetry and stoichiometry.
- Improvements when performing sequence searches.

Jesse redesigned our what's new page and made it look really nice, take a look to see all the details!




Thursday, August 1, 2013

Workshop on Sustainable Software for Science: Practice and Experiences

Interested in sustainability of scientific software?

First Workshop on Sustainable Software for Science: Practice and
Experiences (WSSSPE)
(to held in conjunction with SC13, Sunday, 17 November 2013, Denver, CO, USA)
http://wssspe.researchcomputing.org.uk/


Tuesday, March 19, 2013

Trendspotting in the Protein Data Bank

Our most recent article has become publicly available. It describes some of the trends that we can observe in the Protein Data Bank:


Trendspotting in the Protein Data Bank. FEBS Letters, (0). Berman, H. M., Coimbatore Narayanan, B., Costanzo, L. D., Dutta, S., Ghosh, S., Hudson, B. P., Lawson, C. L., et al. (n.d.). doi:http://dx.doi.org/10.1016/j.febslet.2012.12.029


http://www.sciencedirect.com/science/article/pii/S0014579313000240

Friday, March 8, 2013

Spring 2013 release of RCSB PDB

This week we released the spring 2013 release of the RCSB PDB's website. This release was a lot of work and we added a ton of new features and improvements. Here my wrap-up of some of our highlights. For a full listing see the What's New page for more features and examples.

Protein Symmetry and Stoichiometry

The calculation of protein symmetry and stoichiometry is one of the major features of this release. The goal is to better describe biological assemblies of proteins according to some of their characteristics.

Symmetry refers to the point group symmetry of a protein complex. Protein complexes with quaternary structure can have rotational symmetry belonging to the point groups: cyclic (Cn), dihedral (Dn), tetrahedral (T), octahedral (O), or icosahedral (I). 


The stoichiometry of a protein complex represents the composition of its subunits. For example, the biological assembly of hemoglobin has two alpha and two beta subunits, represented by the formula A2B2.

Here a few examples:



The Crystal structure of the Clostridium perfringens NetB toxin in the membrane inserted form
has a cyclic symmetry of C7 and is formed by a homo-7-mer (A7)



The Crystal structure of phosphoserine phosphatase from T. onnurineus has Dihedral - D4 symmetry and is formed by a homo-8-mer (A8)
 
The Plasmodium falciparum malaria aminopeptidase has Tetrahedral symmetry and is a homomer composed of 12 chains.


Ferritin has Octahedral symmetry and is a Homo 24-mer


The Foot-and Mouth Disease Virus has Icosahedral symmetry and is formed by a Hetero-240 mer with the a A60B60C60D60 stoichiometry.



Tip: Enable the Axes and Polyhedron options on the Jmol page to get a better understanding of the composition of the protein.

Biologically Interesting Molecules


Various biologically interesting molecules, such as peptide-like antibiotic and inhibitor molecules are being annotated by the PDB. The latest RCSB PDB website provides better access to these data. These molecules are now search-able in the top-bar search, and we provide better reports and visualisation.

Access to Drug Targets in the PDB 

For this release we integrated drug and drug target data from DrugBank (www.drugbank.ca). For example: look up Ibuprofen .

We also provide access to the Anatomical Therapeutic Chemical (ATC) Classification System, which is used for drug classification.

Better site navigation


We have re-designed our page header and given it a cleaner and simpler feel.

Uniprot Gene Names

We added a new search function for Gene names. Proteins that can be linked to PDB via Uniprot can now be identified by their associated gene annotations.


SCOP track in Protein Feature View


 The Protein Feature View now has a new track - domain annotations from SCOP. Example: Obelin

This was just a quick summary of some of the new features. For a complete list, take a look at the What's new page.



Monday, December 10, 2012

Protein Feature View goes open source

The source code for the Protein Feature View has been released at github. This allows you to incorporate the dynamic SVG graphics visualizing UniProt and PDB relationships into your own web sites. You can either use the public JSON services provided by RCSB PDB to populate the view, or display your own data (after setting up your own services).






Saturday, December 8, 2012

RCSB PDB's NAR paper

The manuscript describing recent developments at the RCSB PDB has been released as part of the latest Nucleic Acid Research database issue.

View the manuscript at NAR



The Author Profile provides a graphical timeline on when a particular structure was released in the PDB.



Some of the highlights of this year are improvements in the following areas:
Our continuous efforts to provide a structural view of biology are also reflected by an increase in our user base.  The RCSB PDB web site currently hosts ∼240 000 unique visitors per month (based on the number of unique IP addresses), an increase from the 180 000 visitors last reported in 2011. 

Thursday, December 6, 2012

Ten Simple Rules For The Open Development of Scientific Software

Jim Procter and I drafted a new manuscript for PLOS Computational Biology describing ten simple rules for the open development of scientific software. Our motivation for writing this was to make the development of open scientific software more rewarding and the experience of using software more positive.  The ten rules are intended to serve as a guide for any computational scientist:


Ten Simple Rules For the Open Development of Scientific Software (article)

Tuesday, December 4, 2012

RCSB PDB is hiring

The RCSB PDB is one of the leading biological databases with more than 240,000 unique visitors per month. We have an open position in our team for a Lead Web Architect. The position is located in beautiful San Diego at UCSD .

A detailed job description and online application form can be found at:
http://jobs.ucsd.edu/bulletin/job.aspx?cat=information&sortby=post&jobnum_in=64091

Qualifications:

* MS Degree in Computer Science or comparable combination of education and experience with considerable focus in JavaEE software development.

* Established demonstrated work experience in the role of an architect and developer on medium to large size database-driven web applications using Java EE technology and standards.

* Advanced experience developing the presentation layer of a dynamic, database-driven web application using HTML, CSS, JavaScript, JavaScript Toolkits, Ajax, JSP, XML, Java. Experience resolving browser and cross-platform compatibility issues. Advanced experience with Struts2, Tiles, jQuery.

* Advanced experience with database design, Structured Query Language and RDBMS's such as MySQL. Expertise in web application server administration and configuration such as Tomcat.

* Established expertise in software life cycle methodologies. Experience with build tools such as Maven and Ant, and continuous integration systems such as Cruise Control. Experience with project tracking tools such as Jira.

Friday, November 30, 2012

BioJava 3.0.5 released


BioJava 3.0.5 has been released and is available from http://www.biojava.org/wiki/BioJava:Download as well as from the BioJava maven repository at http://www.biojava.org/download/maven/ .

New Features:

- New parser for CATH classification

- New parser for Stockholm file format

- Significantly improved representation of biological assemblies of protein structures. Now can re-create biological assembly from asymmetric unit

- Several bug fixes

Thanks to Daniel Asarnow for contributing the CATH parser and Amr Al Hossary and Marco Vaz for their contributions to the Stockholm parser.

Thursday, November 29, 2012

The PLOS Computational Biology Software Section

PLOS Computational Biology has been accepting and publishing so call Software Articles for about a year now. The articles that have been released under this category span a wide range of different topics in computational biology.  In today's editorial, Hilmar Lapp and I are providing a brief overview of what has been published  so far and describe some of the ideas behind the Software section.

http://www.ploscompbiol.org/article/info%3Adoi%2F10.1371%2Fjournal.pcbi.1002799

Sunday, October 28, 2012

RCSB PDB web site update Fall 2012

New Features at the RCSB PDB web site

 This week the  RCSB PDB released the latest major web site update. Here a quick description of some of the new features.

Protein Feature View

One of the main new features is the new Protein Feature View. It allows to compare the full length protein sequence, as defined by UniProt with the regions that have been determined in 3D and are available together with their coordinates from the Protein Data Bank.  Besides the visualization of the PDB and UniProt relationships, the  new view also adds additional annotations for a more comprehensive understanding of the protein. External data such as Pfam domains or regions for which Homology Models are available from the ProteinModelPortal are indicated. There are also some annotations that are being calculated on the fly: Protein disorder regions, as predicted by Peter Troshin's BioJava implementation of RONN are available as a histogram-style track. Finally, regions with increased hydrophobicity can be spotted by looking at the Hydropathy track.



The Protein Feature View is built using SVG graphics and extensively uses the jQuery-SVG library. Using SVG graphics for a prominently  feature on the site (it is on every protein-explorer page) has become possible since the majority of all modern browsers support these types of graphics nowadays. However, there is still a number of users who are stuck with old browser versions.  According to our web site traffic logs, this number is rapidly declining and we estimate that currently less than 15% of our users can't use the new view. These users won't see error messages on the protein-explorer page, thought.  The graphics will simply not be visible and provide a graceful fallback to the way the page used to look before the graphics were introduced.

 

Better Pfam integration


Another new feature of this release is a better integration with Pfam. Pfam family names are now searchable and one can quickly lookup all protein structures related to these families. Since Pfam is used in structural genomics projects to prioritize targets for crystallization, a possible use case is to look up domains of unknown function (DUFs) and whether 3D coordinates have already been determined for them. As already mentioned above, Pfam domains can be viewed as part of the new Protein Feature View. Weekly up-to-date Pfam-PDB mappings are being calculated by submitting newly released PDB entries to the HMMER3 web site. The details of this process are being described in more detail at the Pfam blog site.

Searching and Reporting



Other improvements of this RCSB PDB web site update include search and reporting improvements. RCSB searches have been improved for better supporting poly-proteins and their sub-components (see screenshot above). There is also better support for searching drug names (and more information about drugs on the Ligand Summary page (e.g. Lipitor), coming from DrugBank . Once a search has been performed, there are now four different types of reports available for investigating the results. Besides the "traditional" search results there is now a "condensed" view, which provides a compact summary of results. The "gallery" provides images for the proteins that have been found in the search. A "timeline" gives a historic overview when proteins were released in the PDB



A full description of all the new features is (as always) available on the What's New Page.

Friday, August 10, 2012

BioJava 2012 paper published

Today the latest BioJava paper was published, describing the BioJava version 3 series .

Thanks to all developers for their contributions, it would not have been possible without them!

Abstract:

http://bioinformatics.oxfordjournals.org/cgi/content/abstract/bts494?ijkey=BzJOy9GgM2XNw07&keytype=ref

PDF:

http://bioinformatics.oxfordjournals.org/cgi/reprint/bts494?ijkey=BzJOy9GgM2XNw07&keytype=ref

Citation:

BioJava: an open-source framework for bioinformatics in 2012

Andreas Prlic; Andrew Yates; Spencer E. Bliven; Peter W. Rose; Julius
Jacobsen; Peter V. Troshin; Mark Chapman; Jianjiong Gao; Chuan Hock
Koh; Sylvain Foisy; Richard Holland; Gediminas Rimsa; Michael L.
Heuer; H. Brandstatter-Muller; Philip E. Bourne; Scooter Willis

Bioinformatics 2012; doi: 10.1093/bioinformatics/bts494

Monday, May 21, 2012

BioJava 3.0.4 released

BioJava 3.0.4 just hit the servers. This is mainly a bug-fix release addressing a few issues with the protein structure and the disorder modules.

One new feature is that SCOP  can now be either accessed from the original SCOP site in the UK or the Berkeley version.

Wednesday, May 9, 2012

About Pfam and PDB mappings

Today's blog article is not going to get published here but at the Pfam blog site, where Rob Finn and I wrote about the Pfam and PDB mappings:


Does my family of interest have a determined 3D protein structure?

Friday, May 4, 2012

Systematic domain based structure alignments at the RCSB PDB

At the RCSB PDB web site we habe been providing pre-calculated and systematic protein structure alignments already for about two years. Every week we run systematic structure alignments for newly released proteins across a representative subset of the database and try to identify related proteins based on their 3D shape.

This week we have released a major upgrade to those efforts. Our pre-computed alignments are now using domain information to split protein chains into smaller subunits. We introduced this change because many proteins are built of more than just one domain. In that case our previous results were a bit unclear in the sense that results for any of the domains were displayed together and that made the data more difficult to compare and interpret.

The new domain based procedure is using the SCOP domain assignments where available to define how to break up protein chains. If the structures are too new to be annotated by SCOP (like all newly released proteins), then we use a software called ProteinDomainParser to define domains based on geometric criteria. Even if the algorithm sometimes defines a break point that might not be the same what SCOP would define, it is still interesting if you find structural neighbors with such fragments of proteins.



In addition, this release of the RCSB PDB web site also provides a new display of protein chains and how different sources annotate protein domains (see the image above). This domain summary shows SCOP domains, ProteinDomainParser domains and Pfam domains. Here an example for  a Cyclodextrin glycosyl transferase (3BMV) from Thermoanerobacterium thermosulfurigenes. It is composed of four domains, which are identified by all of the three data sources.



Friday, April 27, 2012

java.sun.com down?!

I just realized that java.sun.com is not reachable  and java.net displays this little message:


Thanks to google cache (all the sites containing the original info are unreachable as well) I found a notification in one of the Java User groups:


Notice: Java.net will be offline starting around noon pacific time (7pm UTC) on Friday, April 27, through noon pacific time (7pm UTC) on Monday, April 30 for scheduled maintenance.


A planned downtime of three days. If we would do that at work we would get into serious trouble! We go through all sorts of efforts to ensure 24/7 availability and even if we have major site updates we don't let that impact our availability.

I noticed this because one of my tomcat instances did now want to start up which was caused by a failing XML- dtd validation. 


Luckily this can be easily fixed by removing the dtd validation in the xml file. However I wonder how many tomcat servers world wide will have similar problems to start up until Monday! Such a long downtime of such "standard" URLs is rather unprofessional IMHO.



Thursday, March 29, 2012

PLoS comp biol goes Wikipedia

Wikipedia  has become a primary resource of information for many students when looking up basic information. However there is an interesting gap between the scientific community and the people who are regularly contributing to wikipedia articles. There are only few prominent scientists who are regulars, such as the Pfam authors who recently integrated wikipedia into the Xfam series of databases. Another major science related project on wikipedia with about ten thousand articles describing various genes is GeneWiki, lead by Andrew Su. A possible reason for this difference in communities might be the lack of acceptance as academic publishing for wikipedia articles. As of today PLoS comp biol tries to resolve this disparity by publishing a new type of manuscripts, Topic Pages.

Topic Pages are designed to provide review style articles. These articles serve as a copy of reference, that can be cited and will show up in Pubmed. It will also be released at wikipedia where a living copy of the document can be edited and updated by the wider public. This is done in collaboration with the wikiproject computational biology.

How does this work? In short, an article is first submitted to PLoS where it is peer reviewed and upon acceptance it will be published by PLoS comp biol as well as uploaded to wikipedia. While this sounds rather straightforward, one of the issues with this approach is around licensing.

PLoS is publishing all articles under a very liberal license, the Creative Commons Attribution License.  This means, you can do with the article what you want, even change the license, as long as you credit the original sources. This license is in fact more liberal than the wikipedia license, which is Creative Commons Attribution Share Alike. This means we can take a PLoS comp biol article and publish it on wikipedia, as long as we cite the original source of the text, but we can not do this in the opposite direction.

In order to avoid any licensing conflicts, Spencer Bliven set up a custom Mediawiki instance with the liberal PLoS style license. It can be found at http://topicpages.ploscompbiol.org/.  By drafting the manuscript there we are able to transfer the content easily over to both PLoS and wikipedia, once it has passed the PLoS review process. Besides this, also the review process is transparent and you can see what the referees commented on our article at the talk page of the article (both at wikipedia and the topicpages sites)

Our  latest paper is the first such Topic Page. It provides a review on Circular Permutations in proteins, a type of relationship in proteins, whereby the proteins have a changed order of amino acids in their protein sequence while their 3D shape remains very similar.



For more information read the full article at plos ( doi:10.1371/journal.pcbi.1002445 ) or take a look at the latest version of this at wikipedia. Also read the PLoS comp biol editorial, announcing the Topic Pages