Linguistics Research - Human Language, Phonetics, Syntax, Phonology

Linguistics Research Today is a free monthly online journal that collates and summarizes the latest research about Linguistics, including details on human language, phonetics, syntax, phonology.


Linguistics Research Today

Home

View Latest Issue

Information About Linguistics

Books on Linguistics

Advertising in Research Today

View Other Research Today Publications



Discovering semantic features in the literature: a foundation for building functional associations.

Chagoyen M, Carmona-Saez P, Shatkay H, Carazo JM, Pascual-Montano A

Biocomputing Unit, Centro Nacional de Biotecnologia-CSIC, Madrid, Spain. monica.chagoyen@cnb.uam.es

BACKGROUND: Experimental techniques such as DNA microarray, serial analysis of gene expression (SAGE) and mass spectrometry proteomics, among others, are generating large amounts of data related to genes and proteins at different levels. As in any other experimental approach, it is necessary to analyze these data in the context of previously known information about the biological entities under study. The literature is a particularly valuable source of information for experiment validation and interpretation. Therefore, the development of automated text mining tools to assist in such interpretation is one of the main challenges in current bioinformatics research. RESULTS: We present a method to create literature profiles for large sets of genes or proteins based on common semantic features extracted from a corpus of relevant documents. These profiles can be used to establish pair-wise similarities among genes, utilized in gene/protein classification or can be even combined with experimental measurements. Semantic features can be used by researchers to facilitate the understanding of the commonalities indicated by experimental results. Our approach is based on non-negative matrix factorization (NMF), a machine-learning algorithm for data analysis, capable of identifying local patterns that characterize a subset of the data. The literature is thus used to establish putative relationships among subsets of genes or proteins and to provide coherent justification for this clustering into subsets. We demonstrate the utility of the method by applying it to two independent and vastly different sets of genes. CONCLUSION: The presented method can create literature profiles from documents relevant to sets of genes. The representation of genes as additive linear combinations of semantic features allows for the exploration of functional associations as well as for clustering, suggesting a valuable methodology for the validation and interpretation of high-throughput experimental data.

Published 2 March 2006 in BMC Bioinformatics, 7: 41.
Full-text of this article is available online (may require subscription).

Place a permanent text-link or advertisement here for just US$15.

© 2005-2008 Linguistics Research Today. All Rights Reserved.



Linguistics Research Today Archive:

Volume 1 (2005)
  Issue 1 (August)
  Issue 2 (September)
  Issue 3 (October)
  Issue 4 (November)
  Issue 5 (December)

Volume 2 (2006)
  Issue 1 (January)
  Issue 2 (February)
  Issue 3 (March)
  Issue 4 (April)
  Issue 5 (May)
  Issue 6 (June)
  Issue 7 (July)
  Issue 8 (August)
  Issue 9 (September)
  Issue 10 (October)
  Issue 11 (November)
  Issue 12 (December)

Volume 3 (2007)
  Issue 1 (January)
  Issue 2 (February)
  Issue 3 (March)
  Issue 4 (April)
  Issue 5 (May)
  Issue 6 (June)
  Issue 7 (July)
  Issue 8 (August)
  Issue 9 (September)
  Issue 10 (October)
  Issue 11 (November)
  Issue 12 (December)

Volume 4 (2008)
  Issue 1 (January)
  Issue 2 (February)
  Issue 3 (March)
  Issue 4 (April)
  Issue 5 (May)



Linguistics Books

2008 Novel & Short Story Writer's Market (Novel and Short Story Writer's Market)

2008 Novel & Short Story Writer's Market (Novel and Short Story Writer's Market)