My latest research is focused on cluster/list annotation in biology. Given a cluster or list of genes or proteins that were grouped together using some metric (expression profile, sequence or structure similarity, interactions, etc), how can you discover descriptive terms or labels for that cluster? This seems to be a common question, and yet I've had trouble finding tools that help you do what I am specifically trying to do (investigation of a list of biological entities). I've found many that can give you tons of information for single genes or proteins, which I don't consider that helpful, and a few that can give you information for a group, but these are either organism specific or limited to one or two types of data (e.g. GO terms).
Since I am developing a method to do this based on text, I'd like to be able to compare my method to existing ones that solve the same problem. What I am looking for is two or three available methods that give you information relevant to a list of biological entities from multiple species, at least one of which uses literature or text-mining. Does anyone know of such methods, or have ideas of where to look? Various PubMed and Google searches have failed me!
Unrelated, but also done today: Submitted the PSB proposal to Nature Precedings as per several of your requests. Will update once word is back from their review process.
Showing posts with label text mining. Show all posts
Showing posts with label text mining. Show all posts
Tuesday, February 26, 2008
Wednesday, January 16, 2008
A win for Open Access
NIH's public access policy is now mandatory for all NIH-funded investigators. The policy requires submission of a full, electronic version of each published manuscript to PubMed Central, where the full text or PDF is freely available for viewing or download. The BioMed Central blog has a good post about it.
This might be the biggest push towards open access to date, with a government mandate covering the majority of US researchers. Not surprisingly, many publishers are not happy with this development, but I don't think there's much they can do about it, since they will lose a significant number of authors/manuscripts to Open Access journals if they try to deny them, and with that will go a significant amount of their prestige.
Now, what would be even more helpful (and perhaps PubMed Central is doing this) is to provide free access to the full text of each article. Even better, free access to structured text of each article. Natural language processing of biomedical text is a growing field that would benefit hugely if sections, figures, tables, captions, references, etc, were all labeled as such in a computer-readable way. Maybe it's not such a pipe dream, even, to imagine all articles structured this way and indexed with biomedical terms as an automatic pre-publication step.
It will be interesting to follow the drama between Open Access, traditional publishers, and the NIH policy in the next few months!
This might be the biggest push towards open access to date, with a government mandate covering the majority of US researchers. Not surprisingly, many publishers are not happy with this development, but I don't think there's much they can do about it, since they will lose a significant number of authors/manuscripts to Open Access journals if they try to deny them, and with that will go a significant amount of their prestige.
Now, what would be even more helpful (and perhaps PubMed Central is doing this) is to provide free access to the full text of each article. Even better, free access to structured text of each article. Natural language processing of biomedical text is a growing field that would benefit hugely if sections, figures, tables, captions, references, etc, were all labeled as such in a computer-readable way. Maybe it's not such a pipe dream, even, to imagine all articles structured this way and indexed with biomedical terms as an automatic pre-publication step.
It will be interesting to follow the drama between Open Access, traditional publishers, and the NIH policy in the next few months!
Subscribe to:
Posts (Atom)