Michael Barton over at Bioinformatics Zen is collecting responses from those working in the field of bioinformatics to survey the current climate (and projected future) of bioinformatics, with data to be made public and back analysis encouraged. The (fully functional) survey is replicated below, the original can be found here.
Showing posts with label bioinformatics. Show all posts
Showing posts with label bioinformatics. Show all posts
Tuesday, July 1, 2008
Monday, April 14, 2008
Envisioning the scientific community as One Big Lab
The blogosphere has been abuzz recently, or, at least, it seems that way if you've only been checking up on it sporadically the last few weeks. Jennifer Rohn's post about lab notebooks has spurred over 100 lively comments spanning electronic lab notebooks, peer-review, openness in science, and the reward system in science, making for an engrossing peek at the social science of science. Cameron's own musings on that discussion. Pawel Szczesny writes about what it means to be a freelancing scientist. All of this is fascinating and it is exciting to contemplate both what the future of science holds and the obstacles we will need to overcome; the fact that there are indeed stubborn obstacles (technological as well as cultural) and potentially tremendous rewards makes the anticipation of that future all the more heightened.
Emboldened by the collective fervor, I would like to propose an idea - an idea with the same name as this blog. But first, the back story.
About 8 months ago, one of my lab mates was writing up a short paper for submission to a translational bioinformatics conference. The work she was submitting revolved around a powerful literature-search tool tailored for pharmacogenomics called Pharmspresso. Although Pharmspresso had features lacking in existing search methods and was thus useful, the intent was for it to recognize genes, drugs and polymorphisms in free text, and so she needed a way to evaluate its performance. The evaluation task would be straightforward: given a set of pharmacogenomics papers, what percentage of the mentions of genes, drugs, and polymorphisms does Pharmspresso capture? Getting the list of recognized entities from Pharmspresso would be easy, just give it the documents and set it running. But what would be the gold standard?
Typically, gold standards are created by humans. In this case, it would be the list of entities recognized by human readers with the appropriate knowledge to make the distinctions, in the same set of papers. To get her gold standard then, she essentially asked favors of her colleagues in the lab and the department, which translated to a number of them reading papers and doing data entry during free time (or during faculty talks) at a departmental retreat in early fall - not exactly fun, but done out of a sense of duty to science and the goodness of their hearts.
Afterwards, while socializing during one of the poster sessions, this task came up, and the discussion (in which Samuel Flores, Magda Jonikas, Yael Garten, Alain Laederach, and Bernie Daigle all participated) quickly turned to alternative solutions for tackling this and similar problems in science - those requiring knowledge and resources external to your own. As another example, many bioinformaticians work on problems that produce predictions of functions which would benefit from experimental tests of their validity. Conversely, a wet lab may benefit greatly from someone with computational expertise guiding or leading the data analysis, or even providing the hypotheses for experimental studies (in the form of predictions). This is the stuff from which many collaborations are born, but it may be difficult to find the right people in the first place, or the task at hand might seem not quite collaboration-worthy.
In essence, the problem boils down to this: you or your lab possesses a certain collection of skills, knowledge, and resources (hereafter referred to as simply resources), but your needs may not be fully addressed by what you possess. The solution lies in this simple proposition: some other person or lab has what you're looking for.
While it makes sense for a lab or individual to grow their resources and be mostly self-sufficient, at some point it becomes more economical to outsource certain tasks - to companies for antibody development, software for data analysis, supercomputers for high-throughput computing, etc. In some cases, the exchange takes place directly at the academic level, for example, with some labs maintaining and sharing specific cell lines or mouse strains for use by other researchers, or less directly through the use of published and available tools for all sorts of tasks in bioinformatics. So it would seem that outsourcing is common and accepted. But aside from these sorts of established avenues, what other needs do scientists have in conducting their research that are not easily solved? How often is a line of inquiry abandoned or slowed because of a lack of necessary skills, knowledge, or material resources?
The idea behind One Big Lab is that the scientific community should act as, well, one big lab, sharing resources when it makes sense, and everyone, especially the community as a whole, benefits.
During that discussion at the departmental retreat, the solution boiled down to some form of online transaction service built around a credit system. Scientist X would like 5 gold standard outputs for a certain task, so she posts a description of the task along with some credit attached. Other users can then sign up to complete the task, after which they receive the stated number of credits. Of course, in order to post tasks, you need to have a balance of credits you can draw from - which you earn by doing other people's tasks. Getting credits into the system to start needs to be figured out (give everyone N credits? Money for credits?), but assuming there's some baseline of credit floating around amongst the various users, an equilibrium should eventually be reached (at least, that's the hope).
Variations on this theme are natural - have a peer rating system, have the final credit payment be subject to a bidding system (based somehow on user ratings, e.g. highly rated users can ask for more credits to complete a task and the task-poster may select which user to "hire" based on the user ratings as well as how much each user is asking), have some kind of mechanism for taking transactions "offline" into serious collaborations, etc. Tasks may run the gamut from routine and rote to intellectually stimulating and scientifically rewarding. Obviously, guidelines will have to be set for what transactions may be appropriate for this forum and which ones might be more suited for formal, collaborative relationships - but even here, a forum such as this could be very useful for finding collaborators.
In addition to the scientific transaction system, there could be other features that build on the community aspect, such as journal clubs, informal manuscript review, resources for students, and discussion forums. There could be repositories for knowledge or links to existing ones, informal or formal consulting, and casual exchange of ideas which could stimulate research or professional development. All of this should reinforce the idea that science is strengthened by community and the scientific community should not be held back by insufficient allocation of resources.
Although there are a number of websites out there that tackle some of these aspects, especially the community-building ones, I haven't really seen much resembling the transaction system, which is really the core of the idea. Pawel's freelance science comes close, and what I'd like to see is a formalized community-wide online service for essentially that. Maybe this is technically infeasible right now right the way grants work (it may be difficult to justify spending time or resources on other people's research) or with the way scientists work, but I would like to think that the basic premise - bringing together people with complementary skills and resources - makes sense and balances out in everyone's favor. (Whether this premise actually pans out in practice is up for debate - if we offered credits for cash, would anyone ever do someone else's tasks, or would demand outpace supply? By the same token, there could be "freelance" scientists like Pawel who primarily complete tasks, and could then have the option of "cashing out".) I'm sure there are a ton of tricky legal, IP, financial, organizational, etc not to mention social and cultural issues (would you trust someone you don't know to do work for you?), but I think the idea of having One Big Lab is worth exploring.
If I had the time, skills, and business acumen I would throw together a prototype and work out a business plan, but at the moment the most I can do is outsource it to the closest thing we have to One Big Lab - the blogosphere. ;)
Incidentally, Alain Laederach had come up with a similar idea about a year earlier and we thought about naming it "Experitrade" - an online system for trading experiments, essentially, but the name sounded too corporate and the grant he wrote never got off the ground. But the idea has persisted and inspired One Big Lab.
So, I'd welcome any thoughts, logical extensions, deal-makers or deal-breakers, important issues to consider, "prior art"... does anyone think this idea has legs? Will it work if it is completely altruistic? Does adding money into the equation detract from its mission or the science? What sorts of technical and organizational roadblocks are there? Clearly it makes the most sense, if any prototype is developed, to start small - with a couple participating labs or within a school or university, which helps with the trust issue as well. But I'd like to make sure I'm not completely missing the picture!
Emboldened by the collective fervor, I would like to propose an idea - an idea with the same name as this blog. But first, the back story.
About 8 months ago, one of my lab mates was writing up a short paper for submission to a translational bioinformatics conference. The work she was submitting revolved around a powerful literature-search tool tailored for pharmacogenomics called Pharmspresso. Although Pharmspresso had features lacking in existing search methods and was thus useful, the intent was for it to recognize genes, drugs and polymorphisms in free text, and so she needed a way to evaluate its performance. The evaluation task would be straightforward: given a set of pharmacogenomics papers, what percentage of the mentions of genes, drugs, and polymorphisms does Pharmspresso capture? Getting the list of recognized entities from Pharmspresso would be easy, just give it the documents and set it running. But what would be the gold standard?
Typically, gold standards are created by humans. In this case, it would be the list of entities recognized by human readers with the appropriate knowledge to make the distinctions, in the same set of papers. To get her gold standard then, she essentially asked favors of her colleagues in the lab and the department, which translated to a number of them reading papers and doing data entry during free time (or during faculty talks) at a departmental retreat in early fall - not exactly fun, but done out of a sense of duty to science and the goodness of their hearts.
Afterwards, while socializing during one of the poster sessions, this task came up, and the discussion (in which Samuel Flores, Magda Jonikas, Yael Garten, Alain Laederach, and Bernie Daigle all participated) quickly turned to alternative solutions for tackling this and similar problems in science - those requiring knowledge and resources external to your own. As another example, many bioinformaticians work on problems that produce predictions of functions which would benefit from experimental tests of their validity. Conversely, a wet lab may benefit greatly from someone with computational expertise guiding or leading the data analysis, or even providing the hypotheses for experimental studies (in the form of predictions). This is the stuff from which many collaborations are born, but it may be difficult to find the right people in the first place, or the task at hand might seem not quite collaboration-worthy.
In essence, the problem boils down to this: you or your lab possesses a certain collection of skills, knowledge, and resources (hereafter referred to as simply resources), but your needs may not be fully addressed by what you possess. The solution lies in this simple proposition: some other person or lab has what you're looking for.
While it makes sense for a lab or individual to grow their resources and be mostly self-sufficient, at some point it becomes more economical to outsource certain tasks - to companies for antibody development, software for data analysis, supercomputers for high-throughput computing, etc. In some cases, the exchange takes place directly at the academic level, for example, with some labs maintaining and sharing specific cell lines or mouse strains for use by other researchers, or less directly through the use of published and available tools for all sorts of tasks in bioinformatics. So it would seem that outsourcing is common and accepted. But aside from these sorts of established avenues, what other needs do scientists have in conducting their research that are not easily solved? How often is a line of inquiry abandoned or slowed because of a lack of necessary skills, knowledge, or material resources?
The idea behind One Big Lab is that the scientific community should act as, well, one big lab, sharing resources when it makes sense, and everyone, especially the community as a whole, benefits.
During that discussion at the departmental retreat, the solution boiled down to some form of online transaction service built around a credit system. Scientist X would like 5 gold standard outputs for a certain task, so she posts a description of the task along with some credit attached. Other users can then sign up to complete the task, after which they receive the stated number of credits. Of course, in order to post tasks, you need to have a balance of credits you can draw from - which you earn by doing other people's tasks. Getting credits into the system to start needs to be figured out (give everyone N credits? Money for credits?), but assuming there's some baseline of credit floating around amongst the various users, an equilibrium should eventually be reached (at least, that's the hope).
Variations on this theme are natural - have a peer rating system, have the final credit payment be subject to a bidding system (based somehow on user ratings, e.g. highly rated users can ask for more credits to complete a task and the task-poster may select which user to "hire" based on the user ratings as well as how much each user is asking), have some kind of mechanism for taking transactions "offline" into serious collaborations, etc. Tasks may run the gamut from routine and rote to intellectually stimulating and scientifically rewarding. Obviously, guidelines will have to be set for what transactions may be appropriate for this forum and which ones might be more suited for formal, collaborative relationships - but even here, a forum such as this could be very useful for finding collaborators.
In addition to the scientific transaction system, there could be other features that build on the community aspect, such as journal clubs, informal manuscript review, resources for students, and discussion forums. There could be repositories for knowledge or links to existing ones, informal or formal consulting, and casual exchange of ideas which could stimulate research or professional development. All of this should reinforce the idea that science is strengthened by community and the scientific community should not be held back by insufficient allocation of resources.
Although there are a number of websites out there that tackle some of these aspects, especially the community-building ones, I haven't really seen much resembling the transaction system, which is really the core of the idea. Pawel's freelance science comes close, and what I'd like to see is a formalized community-wide online service for essentially that. Maybe this is technically infeasible right now right the way grants work (it may be difficult to justify spending time or resources on other people's research) or with the way scientists work, but I would like to think that the basic premise - bringing together people with complementary skills and resources - makes sense and balances out in everyone's favor. (Whether this premise actually pans out in practice is up for debate - if we offered credits for cash, would anyone ever do someone else's tasks, or would demand outpace supply? By the same token, there could be "freelance" scientists like Pawel who primarily complete tasks, and could then have the option of "cashing out".) I'm sure there are a ton of tricky legal, IP, financial, organizational, etc not to mention social and cultural issues (would you trust someone you don't know to do work for you?), but I think the idea of having One Big Lab is worth exploring.
If I had the time, skills, and business acumen I would throw together a prototype and work out a business plan, but at the moment the most I can do is outsource it to the closest thing we have to One Big Lab - the blogosphere. ;)
Incidentally, Alain Laederach had come up with a similar idea about a year earlier and we thought about naming it "Experitrade" - an online system for trading experiments, essentially, but the name sounded too corporate and the grant he wrote never got off the ground. But the idea has persisted and inspired One Big Lab.
So, I'd welcome any thoughts, logical extensions, deal-makers or deal-breakers, important issues to consider, "prior art"... does anyone think this idea has legs? Will it work if it is completely altruistic? Does adding money into the equation detract from its mission or the science? What sorts of technical and organizational roadblocks are there? Clearly it makes the most sense, if any prototype is developed, to start small - with a couple participating labs or within a school or university, which helps with the trust issue as well. But I'd like to make sure I'm not completely missing the picture!
Friday, February 22, 2008
Review on open source CMS for bioinformatics
An alum of my lab recently published a review on open source content management systems and their uses in bioinformatics.
Wednesday, January 30, 2008
OpenMM, Google Cell, and a marriage between informatics and simulation
Vijay Pande gave a talk today at the SimBIOS weekly seminar series here at Stanford. You may know him from such hits as "Folding@home", which has gone almost triple platinum since it was first released. He is a major figure in the protein folding and molecular simulation world, so the talk was definitely well attended.
Vijay's talk centered around a few major projects and themes, many of which I thought might be interesting to those in the Open Science world. The following are my summaries of those themes.
OpenMM - Open Molecular Mechanics
The molecular dynamics community is fragmented, with many different codes and software packages existing to do MD with overlapping functionality. A result of this is that different labs do things their own way, and new advances are adopted slowly because they must be ported into each set of codes. To address this, they are developing OpenMM, an extensible API for molecular mechanics that will, in principle, unify the MD community the way OpenGL did for graphics. If OpenMM is used as the back end to software applications, advances in theory or hardware will immediately translate to those applications.
Google Cell
Ok, this isn't really what he called it, but it's what I immediately thought of when he talked about their hopes to build a structural picture of an entire cell in atomic detail. I think he may have referred to the project under the name "AMOEBA", since that is what they're shooting for first. The idea is to use structural data from x-ray crystallography, cryoEM, and tomography, from which more and more high-quality data is being produced every day, and turn to physics-based simulation to fill in the details. This is a very high-level idea and I love it, if they can do it. And if they do... well, that made me think of a potential interface for it - Google Cell. Like Google Maps but for the cell instead of the Earth. Pan and zoom and click on interesting features, maybe even do searches! We're definitely years away from it, but the possibilities are endless.
Simulation-aided drug design
Many approaches towards drug design involve docking of potential ligands to rigid crystal structures. But Jim Wells at UCSF has shown that proteins can undergo allosteric changes upon binding to different ligands. Simulation can allow docking with "induced fit", providing a more realistic prediction of binding and possibly even affinity. There was a really cool animation that went along with that part of the talk, which will hopefully be posted soon.
Physics-based simulation is transferable
One of the good things about simulation is that it is transferable across many disciplines. Informatics, too, actually. But when data is scarce, it sometimes pays to borrow tools from other fields. The example Vijay used was that of the protein folding problem. We know relatively little about how proteins fold, but we do know a lot about chemistry and physics, which ultimately govern how proteins fold. So why not exploit our knowledge of chemistry and physics to try and learn more about protein folding? And that is precisely what molecular dynamics simulations do. More generally, transferability is a useful concept to keep in mind. In fact, one of the founding ideas behind SimBIOS is the idea that many diverse biological problems can be solved using the same tools and basic principles, one of them being physics-based simulation.
Informatics & Simulation, happily ever after?
An interesting observation Vijay made was that informaticians and simulation people have traditionally formed two separate, and sometimes antagonistic camps. Something like "physics-based simulation can't teach us anything!" vs. "informatics is sloppy". But in Vijay's work, it is becoming evident that while both approaches have their strengths and weakness, much more can be accomplished when the two are combined. He used the analogy of trying to find a specific person in a large city. One could use a device that beeps when close to the target, but you could search all day without ever getting anywhere near the person. Instead, if you looked the person up in a phone book - used information, in other words - you could find the person's house quite easily, and then the device would be extremely useful. Informatics is very good at getting in the ballpark, but sometimes suffers from being lower resolution. Simulation, on the other hand, can be very precise, but sometimes needs guidance or it will never find the solution. A partnership between informatics and simulation therefore seems not only powerful, but natural.
Vijay's talk centered around a few major projects and themes, many of which I thought might be interesting to those in the Open Science world. The following are my summaries of those themes.
OpenMM - Open Molecular Mechanics
The molecular dynamics community is fragmented, with many different codes and software packages existing to do MD with overlapping functionality. A result of this is that different labs do things their own way, and new advances are adopted slowly because they must be ported into each set of codes. To address this, they are developing OpenMM, an extensible API for molecular mechanics that will, in principle, unify the MD community the way OpenGL did for graphics. If OpenMM is used as the back end to software applications, advances in theory or hardware will immediately translate to those applications.
Google Cell
Ok, this isn't really what he called it, but it's what I immediately thought of when he talked about their hopes to build a structural picture of an entire cell in atomic detail. I think he may have referred to the project under the name "AMOEBA", since that is what they're shooting for first. The idea is to use structural data from x-ray crystallography, cryoEM, and tomography, from which more and more high-quality data is being produced every day, and turn to physics-based simulation to fill in the details. This is a very high-level idea and I love it, if they can do it. And if they do... well, that made me think of a potential interface for it - Google Cell. Like Google Maps but for the cell instead of the Earth. Pan and zoom and click on interesting features, maybe even do searches! We're definitely years away from it, but the possibilities are endless.
Simulation-aided drug design
Many approaches towards drug design involve docking of potential ligands to rigid crystal structures. But Jim Wells at UCSF has shown that proteins can undergo allosteric changes upon binding to different ligands. Simulation can allow docking with "induced fit", providing a more realistic prediction of binding and possibly even affinity. There was a really cool animation that went along with that part of the talk, which will hopefully be posted soon.
Physics-based simulation is transferable
One of the good things about simulation is that it is transferable across many disciplines. Informatics, too, actually. But when data is scarce, it sometimes pays to borrow tools from other fields. The example Vijay used was that of the protein folding problem. We know relatively little about how proteins fold, but we do know a lot about chemistry and physics, which ultimately govern how proteins fold. So why not exploit our knowledge of chemistry and physics to try and learn more about protein folding? And that is precisely what molecular dynamics simulations do. More generally, transferability is a useful concept to keep in mind. In fact, one of the founding ideas behind SimBIOS is the idea that many diverse biological problems can be solved using the same tools and basic principles, one of them being physics-based simulation.
Informatics & Simulation, happily ever after?
An interesting observation Vijay made was that informaticians and simulation people have traditionally formed two separate, and sometimes antagonistic camps. Something like "physics-based simulation can't teach us anything!" vs. "informatics is sloppy". But in Vijay's work, it is becoming evident that while both approaches have their strengths and weakness, much more can be accomplished when the two are combined. He used the analogy of trying to find a specific person in a large city. One could use a device that beeps when close to the target, but you could search all day without ever getting anywhere near the person. Instead, if you looked the person up in a phone book - used information, in other words - you could find the person's house quite easily, and then the device would be extremely useful. Informatics is very good at getting in the ballpark, but sometimes suffers from being lower resolution. Simulation, on the other hand, can be very precise, but sometimes needs guidance or it will never find the solution. A partnership between informatics and simulation therefore seems not only powerful, but natural.
Subscribe to:
Posts (Atom)