Mining relations in the GENIA corpus

Mining relations in the GENIA corpus

inproceedings
Fabio Rinaldi, Gerold Schneider, Kaarel Kaljurand, James Dowdall, Christos Andronis, Andreas Persidis, Ourania Konstanti
Discovering the interactions between genes and proteins is seen as one of the core tasks in molecular biology. The quantity of research results in this area is growing at such a rate that it is very difficult for individual researchers to keep track of them. As such results appear mainly in the form of scientific articles, it is necessary to process them in an efficient manner in order to be able to extract the relevant results. Many databases exist that aim at consolidating the newly gained knowledge in a format that is easily accessible and searchable, however the creators of such databases normally make use of human readers who manually ‘curate’ the rel- evant papers. This is an expensive and time consuming process, besides, there might be a significant time lag be- tween the publication of a result and its introduction into such databases. In this paper we propose a method for discovery of inter- actions between genes and proteins from the scientific liter- ature, based on a complete syntactic analysis of the corpus. We report on preliminary results.
Mining relations in the GENIA corpus
2004
Pisa, Italy
September
61-68
Proc. of the Second European Workshop on Data Mining and Text Mining for Bioinformatics
cl