Showing posts with label open science. Show all posts
Showing posts with label open science. Show all posts
Wednesday, June 11, 2008
Would open science profit from a non-profit?
In my first foray organizing a formal meeting (the PSB workshop), I've learned that pretty much everything comes down to money. Having a successful meeting requires getting people to attend, and getting people to attend often involves money. Getting money to allow people to attend can even require money (for example, publishing costs for conference proceedings).
One idea Cameron's mentioned a few times during our fundraising discussions is the Open Science Collective, and though he hasn't fully described what this is to me, I get the impression that it would be some abstract entity comprising individuals, organizations, resources, and activities - some larger body that would provide support for open science-related endeavors. At the very least, I think it would provide a means to fund raise separate from any particular event such that we (in the collective sense) could act independently. By this I mean that the OSC could support individuals to go to meetings, sponsor meetings like the PSB workshop, or even organize its own meetings.
The idea hasn't really been taken further yet, with whimsical t-shirt ideas pretty much our only tangible revenue strategy so far. But to give the OSC a step towards reality, perhaps it's time to start talking about building a non-profit. From a quick glance through of one website's guide to starting a non-profit, it appears that we need to create a mission statement, obtain a board of directors, and eventually file for either incorporation (plus maybe other things, like tax-exempt or tax-deductible status). This is about where I start getting fuzzy on the details, so if anyone has experience with non-profits or foundations (I don't even know what the difference is) please feel free to enlighten me!
At the most basic level, it would be nice to have some external entity with a bank account into which we could funnel funds that we (again, the collective we) could apply towards open science activities. If that t-shirt shop ever sells anything, the proceeds should go into that account. What the best way to accomplish that is, I'm not quite sure. If it is to start a non-profit, then perhaps the first step is to see who else is interested and start drafting a mission statement together?
According to idealist.org, a mission statement should cover the purpose, the business, the values, and the beneficiaries of the organization; i.e. the what, how, why, and who/where. This can be accomplished in one line but can also be expanded into paragraphs. But let's start small - here's a stab at a one liner mission statement:
"The Open Science Collective is an international and interdisciplinary [non-profit] organization that promotes open exchange and collaboration in science, and provides resources and support for the advancement of open science."
So some concluding questions: Does it make sense to have something like the Open Science Collective? If so, what should it be and how do we get there? What would be the mission statement of the OSC? If not the OSC, what do we need?
One idea Cameron's mentioned a few times during our fundraising discussions is the Open Science Collective, and though he hasn't fully described what this is to me, I get the impression that it would be some abstract entity comprising individuals, organizations, resources, and activities - some larger body that would provide support for open science-related endeavors. At the very least, I think it would provide a means to fund raise separate from any particular event such that we (in the collective sense) could act independently. By this I mean that the OSC could support individuals to go to meetings, sponsor meetings like the PSB workshop, or even organize its own meetings.
The idea hasn't really been taken further yet, with whimsical t-shirt ideas pretty much our only tangible revenue strategy so far. But to give the OSC a step towards reality, perhaps it's time to start talking about building a non-profit. From a quick glance through of one website's guide to starting a non-profit, it appears that we need to create a mission statement, obtain a board of directors, and eventually file for either incorporation (plus maybe other things, like tax-exempt or tax-deductible status). This is about where I start getting fuzzy on the details, so if anyone has experience with non-profits or foundations (I don't even know what the difference is) please feel free to enlighten me!
At the most basic level, it would be nice to have some external entity with a bank account into which we could funnel funds that we (again, the collective we) could apply towards open science activities. If that t-shirt shop ever sells anything, the proceeds should go into that account. What the best way to accomplish that is, I'm not quite sure. If it is to start a non-profit, then perhaps the first step is to see who else is interested and start drafting a mission statement together?
According to idealist.org, a mission statement should cover the purpose, the business, the values, and the beneficiaries of the organization; i.e. the what, how, why, and who/where. This can be accomplished in one line but can also be expanded into paragraphs. But let's start small - here's a stab at a one liner mission statement:
"The Open Science Collective is an international and interdisciplinary [non-profit] organization that promotes open exchange and collaboration in science, and provides resources and support for the advancement of open science."
So some concluding questions: Does it make sense to have something like the Open Science Collective? If so, what should it be and how do we get there? What would be the mission statement of the OSC? If not the OSC, what do we need?
Sunday, May 25, 2008
Science protocols - Recipe for success?
I enjoy cooking and baking, and while I own my share of well-thumbed cookbooks, on a day to day basis I am likely to find my recipes on my favorite cooking websites. There are a number of good ones out there, including Epicurious and Food Network, but the one I go to 99% of the time is AllRecipes.
Why do I prefer this website? For one, the look and feel is inviting, intuitive, and informative. There is no barrier to entry and novice and expert cooks alike will find what they need easily without intimidation or pandering. A nice perk is the ability to search by ingredients, helping you find recipes that will use what you have on hand. But the most important feature is content - the community at Allrecipes is substantial and helpful, not only providing the recipes themselves, but also feedback on the recipes that is often corroborated multiple times. Tips like decreasing the number of eggs or doubling the sauce, roasting at a lower temperature for longer, cutting out the salt, or adding more lime juice can truly be the difference between a successful dish and not.
There has been a lot of discussion on science social networking sites and on whether the promise of "web 2.0" is being delivered in science yet (see David Crotty's post at CSH, Bora's question, and musings at The Scientist over on NN). Some reasons why so-called Science 2.0 hasn't been catching on include the fact that scientists are extremely busy and don't have time to invest in familiarizing themselves with new online networking sites or web tools that have no immediately obvious benefit to them (though some disagree with the claim that scientists are busier than those in other fields). As David points out, much of it boils down to inertia: if we already have a way of doing things, the only way we'll change is if the new way is obviously advantageous and it doesn't take too much effort to adopt it.
It was while reading these related discussions that I started thinking about scientific protocols and how much added benefit could be derived from community content. There are protocol websites out there (OpenWetWare, CSH Protocols, etc) which are a great start, but for the most part these are put up by the original user or published by a journal and rarely generate feedback that could be useful to those looking for a particular protocol - such as slight temperature changes, buffer modifications, or other tweaks that either led to better results or fixed problems. Although it's been a while since I've worked in a wet lab, it seems that a lot of fine optimization goes into a protocol before it produces what is eventually published, and this can often take months to refine.
Given how similar protocols are to recipes, is it that far of a stretch to imagine a protocol version of AllRecipes giving similar benefits? Just as you save time, money, and ingredients by learning from other cooks, you would save time, money, and resources in the lab from other researchers. Granted, this assumes that scientists aren't the type who would say, "What - give other labs a head start by learning from my mistakes? Are you crazy!?" but instead would say, "Think of how much this could help science in general if we all helped each other do experiments more efficiently!" Imagine going to a protocol website, searching by your requirements (protein name, species, type of assay, perhaps), going to the highest rated protocol, and reading a number of reviews that unanimously suggest tweaking one particular step. Or imagine finding the quickest (30-min meals)/most efficient (10 dinners for under $10!)/most popular (95% of people choose this recipe)/best (rated 5 stars by 500 users) protocols for doing X Y or Z as reviewed by scientists like you.
Some might find this kind of crowdsourcing offputting for the scientific domain, others might say it's about time. I know the picture is not so simple, but it just seems silly that we're not benefiting from what other fields (like cooking) have already embraced. Funding is scarce and time is a precious enough resource as it is - why waste both by banging our heads against the same wall others have banged on when we can move forward by finding the door?
I'd be interested to learn if there are actually any protocol websites out there that more fully resemble the types of recipe websites I mention. The solution isn't to create an AllRecipes for protocols (as David mentions in his post) but to provide a service that is useful to scientists and that encourages them to participate. Since the application area is more focused than general science collaboration/networking sites, are the benefits more obvious and will it gain traction more easily?
And for something that is neither here nor there, what is it about science that keeps it from exploiting and embracing the web the way practically everything else has done? (I have inklings but would enjoy hearing others' opinions.)
Why do I prefer this website? For one, the look and feel is inviting, intuitive, and informative. There is no barrier to entry and novice and expert cooks alike will find what they need easily without intimidation or pandering. A nice perk is the ability to search by ingredients, helping you find recipes that will use what you have on hand. But the most important feature is content - the community at Allrecipes is substantial and helpful, not only providing the recipes themselves, but also feedback on the recipes that is often corroborated multiple times. Tips like decreasing the number of eggs or doubling the sauce, roasting at a lower temperature for longer, cutting out the salt, or adding more lime juice can truly be the difference between a successful dish and not.
There has been a lot of discussion on science social networking sites and on whether the promise of "web 2.0" is being delivered in science yet (see David Crotty's post at CSH, Bora's question, and musings at The Scientist over on NN). Some reasons why so-called Science 2.0 hasn't been catching on include the fact that scientists are extremely busy and don't have time to invest in familiarizing themselves with new online networking sites or web tools that have no immediately obvious benefit to them (though some disagree with the claim that scientists are busier than those in other fields). As David points out, much of it boils down to inertia: if we already have a way of doing things, the only way we'll change is if the new way is obviously advantageous and it doesn't take too much effort to adopt it.
It was while reading these related discussions that I started thinking about scientific protocols and how much added benefit could be derived from community content. There are protocol websites out there (OpenWetWare, CSH Protocols, etc) which are a great start, but for the most part these are put up by the original user or published by a journal and rarely generate feedback that could be useful to those looking for a particular protocol - such as slight temperature changes, buffer modifications, or other tweaks that either led to better results or fixed problems. Although it's been a while since I've worked in a wet lab, it seems that a lot of fine optimization goes into a protocol before it produces what is eventually published, and this can often take months to refine.
Given how similar protocols are to recipes, is it that far of a stretch to imagine a protocol version of AllRecipes giving similar benefits? Just as you save time, money, and ingredients by learning from other cooks, you would save time, money, and resources in the lab from other researchers. Granted, this assumes that scientists aren't the type who would say, "What - give other labs a head start by learning from my mistakes? Are you crazy!?" but instead would say, "Think of how much this could help science in general if we all helped each other do experiments more efficiently!" Imagine going to a protocol website, searching by your requirements (protein name, species, type of assay, perhaps), going to the highest rated protocol, and reading a number of reviews that unanimously suggest tweaking one particular step. Or imagine finding the quickest (30-min meals)/most efficient (10 dinners for under $10!)/most popular (95% of people choose this recipe)/best (rated 5 stars by 500 users) protocols for doing X Y or Z as reviewed by scientists like you.
Some might find this kind of crowdsourcing offputting for the scientific domain, others might say it's about time. I know the picture is not so simple, but it just seems silly that we're not benefiting from what other fields (like cooking) have already embraced. Funding is scarce and time is a precious enough resource as it is - why waste both by banging our heads against the same wall others have banged on when we can move forward by finding the door?
I'd be interested to learn if there are actually any protocol websites out there that more fully resemble the types of recipe websites I mention. The solution isn't to create an AllRecipes for protocols (as David mentions in his post) but to provide a service that is useful to scientists and that encourages them to participate. Since the application area is more focused than general science collaboration/networking sites, are the benefits more obvious and will it gain traction more easily?
And for something that is neither here nor there, what is it about science that keeps it from exploiting and embracing the web the way practically everything else has done? (I have inklings but would enjoy hearing others' opinions.)
Monday, April 14, 2008
Envisioning the scientific community as One Big Lab
The blogosphere has been abuzz recently, or, at least, it seems that way if you've only been checking up on it sporadically the last few weeks. Jennifer Rohn's post about lab notebooks has spurred over 100 lively comments spanning electronic lab notebooks, peer-review, openness in science, and the reward system in science, making for an engrossing peek at the social science of science. Cameron's own musings on that discussion. Pawel Szczesny writes about what it means to be a freelancing scientist. All of this is fascinating and it is exciting to contemplate both what the future of science holds and the obstacles we will need to overcome; the fact that there are indeed stubborn obstacles (technological as well as cultural) and potentially tremendous rewards makes the anticipation of that future all the more heightened.
Emboldened by the collective fervor, I would like to propose an idea - an idea with the same name as this blog. But first, the back story.
About 8 months ago, one of my lab mates was writing up a short paper for submission to a translational bioinformatics conference. The work she was submitting revolved around a powerful literature-search tool tailored for pharmacogenomics called Pharmspresso. Although Pharmspresso had features lacking in existing search methods and was thus useful, the intent was for it to recognize genes, drugs and polymorphisms in free text, and so she needed a way to evaluate its performance. The evaluation task would be straightforward: given a set of pharmacogenomics papers, what percentage of the mentions of genes, drugs, and polymorphisms does Pharmspresso capture? Getting the list of recognized entities from Pharmspresso would be easy, just give it the documents and set it running. But what would be the gold standard?
Typically, gold standards are created by humans. In this case, it would be the list of entities recognized by human readers with the appropriate knowledge to make the distinctions, in the same set of papers. To get her gold standard then, she essentially asked favors of her colleagues in the lab and the department, which translated to a number of them reading papers and doing data entry during free time (or during faculty talks) at a departmental retreat in early fall - not exactly fun, but done out of a sense of duty to science and the goodness of their hearts.
Afterwards, while socializing during one of the poster sessions, this task came up, and the discussion (in which Samuel Flores, Magda Jonikas, Yael Garten, Alain Laederach, and Bernie Daigle all participated) quickly turned to alternative solutions for tackling this and similar problems in science - those requiring knowledge and resources external to your own. As another example, many bioinformaticians work on problems that produce predictions of functions which would benefit from experimental tests of their validity. Conversely, a wet lab may benefit greatly from someone with computational expertise guiding or leading the data analysis, or even providing the hypotheses for experimental studies (in the form of predictions). This is the stuff from which many collaborations are born, but it may be difficult to find the right people in the first place, or the task at hand might seem not quite collaboration-worthy.
In essence, the problem boils down to this: you or your lab possesses a certain collection of skills, knowledge, and resources (hereafter referred to as simply resources), but your needs may not be fully addressed by what you possess. The solution lies in this simple proposition: some other person or lab has what you're looking for.
While it makes sense for a lab or individual to grow their resources and be mostly self-sufficient, at some point it becomes more economical to outsource certain tasks - to companies for antibody development, software for data analysis, supercomputers for high-throughput computing, etc. In some cases, the exchange takes place directly at the academic level, for example, with some labs maintaining and sharing specific cell lines or mouse strains for use by other researchers, or less directly through the use of published and available tools for all sorts of tasks in bioinformatics. So it would seem that outsourcing is common and accepted. But aside from these sorts of established avenues, what other needs do scientists have in conducting their research that are not easily solved? How often is a line of inquiry abandoned or slowed because of a lack of necessary skills, knowledge, or material resources?
The idea behind One Big Lab is that the scientific community should act as, well, one big lab, sharing resources when it makes sense, and everyone, especially the community as a whole, benefits.
During that discussion at the departmental retreat, the solution boiled down to some form of online transaction service built around a credit system. Scientist X would like 5 gold standard outputs for a certain task, so she posts a description of the task along with some credit attached. Other users can then sign up to complete the task, after which they receive the stated number of credits. Of course, in order to post tasks, you need to have a balance of credits you can draw from - which you earn by doing other people's tasks. Getting credits into the system to start needs to be figured out (give everyone N credits? Money for credits?), but assuming there's some baseline of credit floating around amongst the various users, an equilibrium should eventually be reached (at least, that's the hope).
Variations on this theme are natural - have a peer rating system, have the final credit payment be subject to a bidding system (based somehow on user ratings, e.g. highly rated users can ask for more credits to complete a task and the task-poster may select which user to "hire" based on the user ratings as well as how much each user is asking), have some kind of mechanism for taking transactions "offline" into serious collaborations, etc. Tasks may run the gamut from routine and rote to intellectually stimulating and scientifically rewarding. Obviously, guidelines will have to be set for what transactions may be appropriate for this forum and which ones might be more suited for formal, collaborative relationships - but even here, a forum such as this could be very useful for finding collaborators.
In addition to the scientific transaction system, there could be other features that build on the community aspect, such as journal clubs, informal manuscript review, resources for students, and discussion forums. There could be repositories for knowledge or links to existing ones, informal or formal consulting, and casual exchange of ideas which could stimulate research or professional development. All of this should reinforce the idea that science is strengthened by community and the scientific community should not be held back by insufficient allocation of resources.
Although there are a number of websites out there that tackle some of these aspects, especially the community-building ones, I haven't really seen much resembling the transaction system, which is really the core of the idea. Pawel's freelance science comes close, and what I'd like to see is a formalized community-wide online service for essentially that. Maybe this is technically infeasible right now right the way grants work (it may be difficult to justify spending time or resources on other people's research) or with the way scientists work, but I would like to think that the basic premise - bringing together people with complementary skills and resources - makes sense and balances out in everyone's favor. (Whether this premise actually pans out in practice is up for debate - if we offered credits for cash, would anyone ever do someone else's tasks, or would demand outpace supply? By the same token, there could be "freelance" scientists like Pawel who primarily complete tasks, and could then have the option of "cashing out".) I'm sure there are a ton of tricky legal, IP, financial, organizational, etc not to mention social and cultural issues (would you trust someone you don't know to do work for you?), but I think the idea of having One Big Lab is worth exploring.
If I had the time, skills, and business acumen I would throw together a prototype and work out a business plan, but at the moment the most I can do is outsource it to the closest thing we have to One Big Lab - the blogosphere. ;)
Incidentally, Alain Laederach had come up with a similar idea about a year earlier and we thought about naming it "Experitrade" - an online system for trading experiments, essentially, but the name sounded too corporate and the grant he wrote never got off the ground. But the idea has persisted and inspired One Big Lab.
So, I'd welcome any thoughts, logical extensions, deal-makers or deal-breakers, important issues to consider, "prior art"... does anyone think this idea has legs? Will it work if it is completely altruistic? Does adding money into the equation detract from its mission or the science? What sorts of technical and organizational roadblocks are there? Clearly it makes the most sense, if any prototype is developed, to start small - with a couple participating labs or within a school or university, which helps with the trust issue as well. But I'd like to make sure I'm not completely missing the picture!
Emboldened by the collective fervor, I would like to propose an idea - an idea with the same name as this blog. But first, the back story.
About 8 months ago, one of my lab mates was writing up a short paper for submission to a translational bioinformatics conference. The work she was submitting revolved around a powerful literature-search tool tailored for pharmacogenomics called Pharmspresso. Although Pharmspresso had features lacking in existing search methods and was thus useful, the intent was for it to recognize genes, drugs and polymorphisms in free text, and so she needed a way to evaluate its performance. The evaluation task would be straightforward: given a set of pharmacogenomics papers, what percentage of the mentions of genes, drugs, and polymorphisms does Pharmspresso capture? Getting the list of recognized entities from Pharmspresso would be easy, just give it the documents and set it running. But what would be the gold standard?
Typically, gold standards are created by humans. In this case, it would be the list of entities recognized by human readers with the appropriate knowledge to make the distinctions, in the same set of papers. To get her gold standard then, she essentially asked favors of her colleagues in the lab and the department, which translated to a number of them reading papers and doing data entry during free time (or during faculty talks) at a departmental retreat in early fall - not exactly fun, but done out of a sense of duty to science and the goodness of their hearts.
Afterwards, while socializing during one of the poster sessions, this task came up, and the discussion (in which Samuel Flores, Magda Jonikas, Yael Garten, Alain Laederach, and Bernie Daigle all participated) quickly turned to alternative solutions for tackling this and similar problems in science - those requiring knowledge and resources external to your own. As another example, many bioinformaticians work on problems that produce predictions of functions which would benefit from experimental tests of their validity. Conversely, a wet lab may benefit greatly from someone with computational expertise guiding or leading the data analysis, or even providing the hypotheses for experimental studies (in the form of predictions). This is the stuff from which many collaborations are born, but it may be difficult to find the right people in the first place, or the task at hand might seem not quite collaboration-worthy.
In essence, the problem boils down to this: you or your lab possesses a certain collection of skills, knowledge, and resources (hereafter referred to as simply resources), but your needs may not be fully addressed by what you possess. The solution lies in this simple proposition: some other person or lab has what you're looking for.
While it makes sense for a lab or individual to grow their resources and be mostly self-sufficient, at some point it becomes more economical to outsource certain tasks - to companies for antibody development, software for data analysis, supercomputers for high-throughput computing, etc. In some cases, the exchange takes place directly at the academic level, for example, with some labs maintaining and sharing specific cell lines or mouse strains for use by other researchers, or less directly through the use of published and available tools for all sorts of tasks in bioinformatics. So it would seem that outsourcing is common and accepted. But aside from these sorts of established avenues, what other needs do scientists have in conducting their research that are not easily solved? How often is a line of inquiry abandoned or slowed because of a lack of necessary skills, knowledge, or material resources?
The idea behind One Big Lab is that the scientific community should act as, well, one big lab, sharing resources when it makes sense, and everyone, especially the community as a whole, benefits.
During that discussion at the departmental retreat, the solution boiled down to some form of online transaction service built around a credit system. Scientist X would like 5 gold standard outputs for a certain task, so she posts a description of the task along with some credit attached. Other users can then sign up to complete the task, after which they receive the stated number of credits. Of course, in order to post tasks, you need to have a balance of credits you can draw from - which you earn by doing other people's tasks. Getting credits into the system to start needs to be figured out (give everyone N credits? Money for credits?), but assuming there's some baseline of credit floating around amongst the various users, an equilibrium should eventually be reached (at least, that's the hope).
Variations on this theme are natural - have a peer rating system, have the final credit payment be subject to a bidding system (based somehow on user ratings, e.g. highly rated users can ask for more credits to complete a task and the task-poster may select which user to "hire" based on the user ratings as well as how much each user is asking), have some kind of mechanism for taking transactions "offline" into serious collaborations, etc. Tasks may run the gamut from routine and rote to intellectually stimulating and scientifically rewarding. Obviously, guidelines will have to be set for what transactions may be appropriate for this forum and which ones might be more suited for formal, collaborative relationships - but even here, a forum such as this could be very useful for finding collaborators.
In addition to the scientific transaction system, there could be other features that build on the community aspect, such as journal clubs, informal manuscript review, resources for students, and discussion forums. There could be repositories for knowledge or links to existing ones, informal or formal consulting, and casual exchange of ideas which could stimulate research or professional development. All of this should reinforce the idea that science is strengthened by community and the scientific community should not be held back by insufficient allocation of resources.
Although there are a number of websites out there that tackle some of these aspects, especially the community-building ones, I haven't really seen much resembling the transaction system, which is really the core of the idea. Pawel's freelance science comes close, and what I'd like to see is a formalized community-wide online service for essentially that. Maybe this is technically infeasible right now right the way grants work (it may be difficult to justify spending time or resources on other people's research) or with the way scientists work, but I would like to think that the basic premise - bringing together people with complementary skills and resources - makes sense and balances out in everyone's favor. (Whether this premise actually pans out in practice is up for debate - if we offered credits for cash, would anyone ever do someone else's tasks, or would demand outpace supply? By the same token, there could be "freelance" scientists like Pawel who primarily complete tasks, and could then have the option of "cashing out".) I'm sure there are a ton of tricky legal, IP, financial, organizational, etc not to mention social and cultural issues (would you trust someone you don't know to do work for you?), but I think the idea of having One Big Lab is worth exploring.
If I had the time, skills, and business acumen I would throw together a prototype and work out a business plan, but at the moment the most I can do is outsource it to the closest thing we have to One Big Lab - the blogosphere. ;)
Incidentally, Alain Laederach had come up with a similar idea about a year earlier and we thought about naming it "Experitrade" - an online system for trading experiments, essentially, but the name sounded too corporate and the grant he wrote never got off the ground. But the idea has persisted and inspired One Big Lab.
So, I'd welcome any thoughts, logical extensions, deal-makers or deal-breakers, important issues to consider, "prior art"... does anyone think this idea has legs? Will it work if it is completely altruistic? Does adding money into the equation detract from its mission or the science? What sorts of technical and organizational roadblocks are there? Clearly it makes the most sense, if any prototype is developed, to start small - with a couple participating labs or within a school or university, which helps with the trust issue as well. But I'd like to make sure I'm not completely missing the picture!
Wednesday, March 26, 2008
March 26 is Document Freedom Day!
Today marks the first observation of Document Freedom Day, from here on out an annual celebration held on the last Wednesday of March.
From the official website:
Given the work on open data standards, structured data, and open repositories being done by Cameron, Peter MR, and others in the open science community, this is definitely cause for celebration! Unfortunately, the United States seems to be lagging behind other countries in its observance of this holiday (but maybe we'll give it another year).
Thanks to Alain Laederach for the tip!
From the official website:
Document Freedom Day (DFD) is a global day for document liberation. It will be a day of grassroots effort to educate the public about the importance of Free Document Formats and Open Standards in general.
Complementary to Software Freedom Day, we aim to have local teams all over the world organise events on the last Wednesday of March. 2008 is the first year that Document Freedom Day is being called for, and we are looking for people around the world who are willing to join the effort.
DFD's main goals are:
- promotion and adoption of free document formats
- forming a global network
- coordination of activities that happen on 26th of March, Document Freedom Day
Once a year, we will celebrate Document Freedom Day as a global community. Between those days, DFD will be focused on facilitating community action and building awareness for issues of Document Freedom and Open Standards.
Given the work on open data standards, structured data, and open repositories being done by Cameron, Peter MR, and others in the open science community, this is definitely cause for celebration! Unfortunately, the United States seems to be lagging behind other countries in its observance of this holiday (but maybe we'll give it another year).
Thanks to Alain Laederach for the tip!
Sunday, February 3, 2008
Sharing in the news
A news feature in the latest issue of Nature ("Genetics by Numbers") discusses the recent proliferation of genome-wide association studies. The story brings up a couple interesting points relevant to Open Science.
- Data sharing is essential. Genome-wide association studies rely on having a lot of data to work with, and collecting the data (through SNP-chips, for example) is still too expensive for most researchers. Even one database of samples may not be enough - only by pooling many such independent sample collections together will conclusive evidence be gathered for variants with modest effects.
- Data sharing has its share of problems. The article cited a study that found that researchers sometimes abuse shared data, either by going outside the bounds of the original agreements, or by not accounting for certain aspects of the data relevant to their research question. Data shared does not necessarily mean better or faster science if it is analyzed poorly. This highlights the importance of thinking through your analysis, finding out how the data was collected, and ensuring that the data is appropriately "cleaned up" for your purposes before you analyze it. While there is increased benefit from sharing data, there is also increased responsibility to use the data properly.
Another potential problem mentioned was that scientists would have fewer incentives to collect new data - and thus science could stagnate a little. I'm not sure how big of a concern this should be - after all, citation is a big incentive, and publishing (and sharing) new data would lead to more citations - but it is a concern I hadn't heard much about before. - Data sharing made "Soviet"? Some US researchers may resent the mandatory policies set by the NIH, and think collaboration and sharing was progressing fine without them. Said one, "I don't want to share my data with anyone because the NIH decides I should, I want to do it because I decide to do it." Perhaps there has been some narcissistic pleasure associated with data sharing (not that there's anything wrong with that), and the NIH mandate is now raining on the parade. It's like doing chores - it can be fun if you don't know you're supposed to do it or take to it naturally (maybe you even get an altruistic kick out of it), but it's a lot less fun when you're told you have to do it or else.
Friday, February 1, 2008
Ignorance of the masses - should we worry?
This is the second somewhat negative post I've written on Open Science. You might say I've moved from the Honeymoon phase to the phase where all you can do is judge and criticize and focus on the bad. Let's hope I move on to the mature, balanced, productive phase soon! In any case, I still highly support the concept of Open Science and want to see it grow, but right now I am using this blog to explore both sides.
There are many issues and questions surrounding Open Science which I have been slowly familiarizing myself with over the last few weeks. Some things, like intellectual property rights, privacy, and scooping, are obvious and comprise the bulk of the debate. I started thinking about a different issue related to Open Science recently, mostly inspired by the escalating battle between evolution and Creationism/ID and the comments of former presidential hopeful Mike Huckabee. The following may be more politically charged than appropriate for a blog like this, so consider yourself warned.
It boils down to this: the public is essentially ignorant. What I mean is that most people know a lot about very few topics, and very little about everything else. Most of what they learn about everything else comes from the media. I won't even go into the problems with our education system or the fact that most Americans have a very strange idea of what science is. The problem is that it doesn't take much for a study to be misinterpreted, or science to be misrepresented. Mainstream media will go for the most sensational spin. Think about all those "health" and "wellness" magazines that immediately latch on to and exaggerate the latest studies on coffee, supplements, and compounds in food, regardless of where they were published.
If Open Science is fully realized, bleeding edge scientific research will be at everyone's fingertips. Preliminary results, perhaps before appropriate controls are performed, will be available to people who don't have the training (or desire) to distinguish between rigorously obtained findings and works in progress. Prior to this, the only science accessible to the world outside went through the filter of a peer-review journal (and presumably is already summarized and interpreted in the way that best describes all the data and findings in the entire study). Without a filter, is there more risk for misinterpretation and misrepresentation by those outside the scientific sphere? If so, what precautions can we take to mitigate it?
There are many issues and questions surrounding Open Science which I have been slowly familiarizing myself with over the last few weeks. Some things, like intellectual property rights, privacy, and scooping, are obvious and comprise the bulk of the debate. I started thinking about a different issue related to Open Science recently, mostly inspired by the escalating battle between evolution and Creationism/ID and the comments of former presidential hopeful Mike Huckabee. The following may be more politically charged than appropriate for a blog like this, so consider yourself warned.
It boils down to this: the public is essentially ignorant. What I mean is that most people know a lot about very few topics, and very little about everything else. Most of what they learn about everything else comes from the media. I won't even go into the problems with our education system or the fact that most Americans have a very strange idea of what science is. The problem is that it doesn't take much for a study to be misinterpreted, or science to be misrepresented. Mainstream media will go for the most sensational spin. Think about all those "health" and "wellness" magazines that immediately latch on to and exaggerate the latest studies on coffee, supplements, and compounds in food, regardless of where they were published.
If Open Science is fully realized, bleeding edge scientific research will be at everyone's fingertips. Preliminary results, perhaps before appropriate controls are performed, will be available to people who don't have the training (or desire) to distinguish between rigorously obtained findings and works in progress. Prior to this, the only science accessible to the world outside went through the filter of a peer-review journal (and presumably is already summarized and interpreted in the way that best describes all the data and findings in the entire study). Without a filter, is there more risk for misinterpretation and misrepresentation by those outside the scientific sphere? If so, what precautions can we take to mitigate it?
Thursday, January 31, 2008
The scooping debate continues
Bora Zivkovic over at ScienceBlogs posted about a scooping story that was just published in Nature. The story itself is quite scandalous, since the individual in question doesn't have a great reputation as far as I could tell from reading various comments on blogs posting on the subject. Go see some of them at ScienceBlogs to get an idea. What I want to address in this post has to do with Bora's commentary on the story, since he suggests that in an Open Science world, scooping would be much more difficult to pull off, since everything is documented and associated with time stamps, and the community can rally behind the "first to blog". I'm not sure the picture is that simple.
Something that I did not fully appreciate before about scooping is that the "scooper" will claim that the discovery was made independently, which is difficult to disprove in many fields. Paleontology and archeology may have very slight advantage here in that some types of discoveries are singular and tangible - a fossil in the desert, artifacts under a dirt mound - physical things at physical locations. Someone would be hard-pressed to say they were at the same place someone else was, digging up the same object (though, as the aetosaur controversy makes clear, scooping of a related sort can still happen). In biology, you might slave away for years to discover that protein A regulates protein B, but if someone else publishes it first, you're out of luck. Sure, you have the reagents and the cell lines and the protein products - but so do your competitors, as long as they have the basic equipment and resources in place to reproduce what you did. So the fear that making your research public would increase your risk of getting scooped maybe is not that unfounded. Closed scientists could easily leech off the hard work of Open scientists with no one being able to prove anything.
Of course, that is an extremely pessimistic view, but unfortunately, it's a kneejerk reaction from many people in biomedical fields. Until Open Science becomes so widespread that anything Closed is viewed with suspicion, there will be the possibility of exploitation. And that statement itself smacks highly of Big Brother. I don't think we want a "tryanny of Openness" any more than we want scooping to happen. My conclusions from this rather sobering train of thought are that yes, scooping is a moral outrage and the fact that it is an issue is frustrating, but because it does happen, we need to think carefully about how to prevent it from escalating as more people go Open. Can we prevent it? Is collective disapproval enough?
Something that I did not fully appreciate before about scooping is that the "scooper" will claim that the discovery was made independently, which is difficult to disprove in many fields. Paleontology and archeology may have very slight advantage here in that some types of discoveries are singular and tangible - a fossil in the desert, artifacts under a dirt mound - physical things at physical locations. Someone would be hard-pressed to say they were at the same place someone else was, digging up the same object (though, as the aetosaur controversy makes clear, scooping of a related sort can still happen). In biology, you might slave away for years to discover that protein A regulates protein B, but if someone else publishes it first, you're out of luck. Sure, you have the reagents and the cell lines and the protein products - but so do your competitors, as long as they have the basic equipment and resources in place to reproduce what you did. So the fear that making your research public would increase your risk of getting scooped maybe is not that unfounded. Closed scientists could easily leech off the hard work of Open scientists with no one being able to prove anything.
Of course, that is an extremely pessimistic view, but unfortunately, it's a kneejerk reaction from many people in biomedical fields. Until Open Science becomes so widespread that anything Closed is viewed with suspicion, there will be the possibility of exploitation. And that statement itself smacks highly of Big Brother. I don't think we want a "tryanny of Openness" any more than we want scooping to happen. My conclusions from this rather sobering train of thought are that yes, scooping is a moral outrage and the fact that it is an issue is frustrating, but because it does happen, we need to think carefully about how to prevent it from escalating as more people go Open. Can we prevent it? Is collective disapproval enough?
Sunday, January 27, 2008
Science and sharing
Following up on my last post on this subject, it appears there has been a recent spate of posts and articles about sharing in science. Taken together, they're a great summary of the challenges Open Science is facing, but also of the benefits and the steps some groups are taking to enhance openness.
Both Neil Saunders and Cameron Neylon point out this article in the NY Times, and a post on the 23andme blog.
In response to the 23andme post, Jasper A. Bovenberg refers to his paper in Genomics, Society, and Policy from 2005.
There is also an article from BusinessWeek in 2007 about some big pharma companies embracing Science 2.0, and an article in Scientific American from a few weeks ago debating the pros and cons.
Both Neil Saunders and Cameron Neylon point out this article in the NY Times, and a post on the 23andme blog.
In response to the 23andme post, Jasper A. Bovenberg refers to his paper in Genomics, Society, and Policy from 2005.
There is also an article from BusinessWeek in 2007 about some big pharma companies embracing Science 2.0, and an article in Scientific American from a few weeks ago debating the pros and cons.
Friday, January 25, 2008
Is the danger of being scooped field-dependent?
The students in my program get together once in a while for what we call "Researchome", also known fondly as "dinner-ome", originally conceived as a casual forum in which students could present their research or other topics of interest to other students, while getting dinner for free. But without someone to present, there is no justification to have Researchome, so rather than deprive a dozen grad students of free food, I threw together a quick presentation of Open Science and my proposal for PSB for our Researchome last night.
Biomedical informatics students are a smart bunch, so there was some great discussion. Naturally, the concern over getting scooped came up, and while I was quick to pooh-pooh it as a naive/narcissistic fear, the others were fairly adamant that it was a valid concern. Several gave personal anecdotes. And the picture that started to emerge was one where the danger of being scooped was highly dependent on the field you were in - theoretical vs. applied, basic vs. translational, science vs. medicine, all of which may put different emphasis on the idea vs. the implementation.
According to one of the students at the Researchome, in theoretical disciplines such as math or physics, credit is given as soon as an idea is recorded. But in fields like cell biology, just having the idea for an experiment or a hypothesis is not enough; instead, you must conduct the experiment and demonstrate successful results through a peer-reviewed publication before credit is given. Because of this, people in these fields are more reluctant to be open about their research before it has been published, and getting scooped can have real consequences for someone's career and funding. In my limited explorations of the world wide open science web, it seems as if a significant portion of those participating are chemists. Is getting scooped less of a concern in chemistry than it is in, say, molecular biology, and, if so, why? If there is a discrepancy between fields in the danger of being scooped, how should this be addressed as the open science community moves forward? Is it possible to change the standards by which success and intellectual credit are determined?
Aside from this interesting issue, some valid points were brought up concerning the proposal for an open science session at PSB. One is that the audience at PSB is by and large composed of scientists who don't generate their own data, but use the data generated by others. A lot of high-throughput, -omics, and bioinformatics-minded people. Therefore, open data and open source will probably be of greater interest to them than the more overarching idea of open notebook science. Focusing on standards, exchange formats, and tools and methodologies for conducting open science may be a good approach.
The other consensus that the students came to was that open science is so broad and important a topic that it should be featured as a session at a much larger conference, such as ISMB or AMIA. Many were bemused as to why I chose what is arguably a niche conference as a venue for open science. My rationale at this point is that it is the soonest we could possibly organize a meeting on open science jointly with an established conference, and I think the audience is relevant enough for it to be productive. Being smaller, it may also be a good stepping stone towards a bigger meeting, and I wouldn't be surprised if some efforts began for that before PSB 2009 comes around.
Many thanks to the students who attended the Researchome for their feedback. I'll be working on a draft of the proposal over the next couple days.
Biomedical informatics students are a smart bunch, so there was some great discussion. Naturally, the concern over getting scooped came up, and while I was quick to pooh-pooh it as a naive/narcissistic fear, the others were fairly adamant that it was a valid concern. Several gave personal anecdotes. And the picture that started to emerge was one where the danger of being scooped was highly dependent on the field you were in - theoretical vs. applied, basic vs. translational, science vs. medicine, all of which may put different emphasis on the idea vs. the implementation.
According to one of the students at the Researchome, in theoretical disciplines such as math or physics, credit is given as soon as an idea is recorded. But in fields like cell biology, just having the idea for an experiment or a hypothesis is not enough; instead, you must conduct the experiment and demonstrate successful results through a peer-reviewed publication before credit is given. Because of this, people in these fields are more reluctant to be open about their research before it has been published, and getting scooped can have real consequences for someone's career and funding. In my limited explorations of the world wide open science web, it seems as if a significant portion of those participating are chemists. Is getting scooped less of a concern in chemistry than it is in, say, molecular biology, and, if so, why? If there is a discrepancy between fields in the danger of being scooped, how should this be addressed as the open science community moves forward? Is it possible to change the standards by which success and intellectual credit are determined?
Aside from this interesting issue, some valid points were brought up concerning the proposal for an open science session at PSB. One is that the audience at PSB is by and large composed of scientists who don't generate their own data, but use the data generated by others. A lot of high-throughput, -omics, and bioinformatics-minded people. Therefore, open data and open source will probably be of greater interest to them than the more overarching idea of open notebook science. Focusing on standards, exchange formats, and tools and methodologies for conducting open science may be a good approach.
The other consensus that the students came to was that open science is so broad and important a topic that it should be featured as a session at a much larger conference, such as ISMB or AMIA. Many were bemused as to why I chose what is arguably a niche conference as a venue for open science. My rationale at this point is that it is the soonest we could possibly organize a meeting on open science jointly with an established conference, and I think the audience is relevant enough for it to be productive. Being smaller, it may also be a good stepping stone towards a bigger meeting, and I wouldn't be surprised if some efforts began for that before PSB 2009 comes around.
Many thanks to the students who attended the Researchome for their feedback. I'll be working on a draft of the proposal over the next couple days.
Friday, January 18, 2008
New meaning to "publish or perish" - an opening for Open Science?
The saying "publish or perish" is well-known in academia, and typically both actions refer to the same subject - you, the aspiring/struggling grad student/post-doc/fellow/assistant professor. A recent correspondence in Nature puts a new and bracing spin on the phrase.
I think at some point most academic researchers have experienced the conflict that can arise when it is time to write a paper. On the one hand, you're getting a chance to reward those months or years of hard work with some exposure and a line on your CV, and invest in the potential for future collaborations. On the other, maybe you've just gotten started on a really promising or exciting research direction, are in a groove, work-wise, and to have something like writing suddenly vying for your attention just means that both activities suffer. You feel that you can't drop what you're working on to write the paper, but the paper writing is distracted and unfocused because you're still trying to conduct research half the time (and thinking about it more than that). But we march on to these two seemingly competing drummers, fueled somewhat by the vague hope that our work, once it is in the public domain, will also contribute a drop in the bucket that is scientific advancement of our species.
But what about other species? In conservation biology, "publish or perish" can take on new, and frighteningly literal, meaning. Time spent working on publications is time taken away from research on ecosystems and endangered wildlife. In the meantime, earth's natural resources and diversity suffer. To prevent this from happening, the authors of the letter suggest (only slightly ironically) the adoption of a new impact factor:
Even if this proposal was made half in jest, it does highlight some important questions. The first sentence essentially asks: how much faster could research be conducted (and, by translation, medical or scientific advances be developed) if there was less emphasis on publication? The second sentence is quite a bit more complex, since it seems it would bring in value judgments on the worth of specific research questions - something that would be hard to define objectively and is easily influenced by prevailing trends, funding, and big talk.
So let's talk about the first idea - that the pace of scientific advancement suffers from the emphases placed on publication. Obviously, research needs to be disseminated if it is going to contribute. But here is where Open Science comes in. Suppose Open Science and Open Notebook Science became the norm rather than the burgeoning, but still fringe, movement that it is now. Two big questions immediately come to my mind: Would publication matter as much as it does now? Would research proceed faster? I say no and yes.
With most, if not all, of your methods, data, and results made public, formal publication would not be necessary for others to learn of and benefit from your work. Peer-review may become an intrinsic part of the entire research process. Of course, a formal summary of your work adds great value and would be indispensable for someone searching for information on your field of study, but much of the pressure to publish could be alleviated. Add to that the increased exposure to the entire community and you get enormous potential to speed up your research in addition to research in general. You can learn what is working and not working in your experiments, get useful feedback and suggestions, and meet people who may be able to help you, all on a much faster timescale. At the same time, new ideas may be spawned, collaborations fostered, and interesting connections made between concepts.
Obviously, the future of Open Science is not going to be as rosy as that, at least in the early stages of its evolution (issues like patenting and privacy are valid and worth lengthy discussion in their own right, but are beyond the scope of this post). In fields like conservation biology, however, the shadow cast by "publish or perish" has terribly real implications, and the move towards Open Science will help to lift it. Can anyone really argue that Open Science is a bad thing?
I think at some point most academic researchers have experienced the conflict that can arise when it is time to write a paper. On the one hand, you're getting a chance to reward those months or years of hard work with some exposure and a line on your CV, and invest in the potential for future collaborations. On the other, maybe you've just gotten started on a really promising or exciting research direction, are in a groove, work-wise, and to have something like writing suddenly vying for your attention just means that both activities suffer. You feel that you can't drop what you're working on to write the paper, but the paper writing is distracted and unfocused because you're still trying to conduct research half the time (and thinking about it more than that). But we march on to these two seemingly competing drummers, fueled somewhat by the vague hope that our work, once it is in the public domain, will also contribute a drop in the bucket that is scientific advancement of our species.
But what about other species? In conservation biology, "publish or perish" can take on new, and frighteningly literal, meaning. Time spent working on publications is time taken away from research on ecosystems and endangered wildlife. In the meantime, earth's natural resources and diversity suffer. To prevent this from happening, the authors of the letter suggest (only slightly ironically) the adoption of a new impact factor:
This impact factor would be based on an estimation of how much worse the conservation status of an endangered species or ecosystem might be in the absence of the candidate's research. It would select for targeted investigation that should help to fill in 'the great divide', and would exclude opportunistic ecology papers claiming to be of conservation significance.
Even if this proposal was made half in jest, it does highlight some important questions. The first sentence essentially asks: how much faster could research be conducted (and, by translation, medical or scientific advances be developed) if there was less emphasis on publication? The second sentence is quite a bit more complex, since it seems it would bring in value judgments on the worth of specific research questions - something that would be hard to define objectively and is easily influenced by prevailing trends, funding, and big talk.
So let's talk about the first idea - that the pace of scientific advancement suffers from the emphases placed on publication. Obviously, research needs to be disseminated if it is going to contribute. But here is where Open Science comes in. Suppose Open Science and Open Notebook Science became the norm rather than the burgeoning, but still fringe, movement that it is now. Two big questions immediately come to my mind: Would publication matter as much as it does now? Would research proceed faster? I say no and yes.
With most, if not all, of your methods, data, and results made public, formal publication would not be necessary for others to learn of and benefit from your work. Peer-review may become an intrinsic part of the entire research process. Of course, a formal summary of your work adds great value and would be indispensable for someone searching for information on your field of study, but much of the pressure to publish could be alleviated. Add to that the increased exposure to the entire community and you get enormous potential to speed up your research in addition to research in general. You can learn what is working and not working in your experiments, get useful feedback and suggestions, and meet people who may be able to help you, all on a much faster timescale. At the same time, new ideas may be spawned, collaborations fostered, and interesting connections made between concepts.
Obviously, the future of Open Science is not going to be as rosy as that, at least in the early stages of its evolution (issues like patenting and privacy are valid and worth lengthy discussion in their own right, but are beyond the scope of this post). In fields like conservation biology, however, the shadow cast by "publish or perish" has terribly real implications, and the move towards Open Science will help to lift it. Can anyone really argue that Open Science is a bad thing?
Friday, January 11, 2008
It's a small world after all
One Big Lab is inspired by the idea of connectedness between scientists. Research shouldn't have to be hindered by a lack of time, resources, or knowledge; if there exists someone who possesses any of those three things who is willing to help someone else, then there should be a way to connect those individuals. In the end, everybody wins - especially society, and especially science.
Many things will have to happen before this idea can become tangible, and we have some thoughts on what One Big Lab could eventually look like, but for now, this blog will be a place for reflection and discussion on all things Open Science. Hopefully, we'll learn a lot, meet some people, and bring the ideas behind One Big Lab to life.
Many things will have to happen before this idea can become tangible, and we have some thoughts on what One Big Lab could eventually look like, but for now, this blog will be a place for reflection and discussion on all things Open Science. Hopefully, we'll learn a lot, meet some people, and bring the ideas behind One Big Lab to life.
Subscribe to:
Posts (Atom)