Chunk Parsing and Entity Relation Extracting to Chinese Text by Using Conditional Random Fields Model
- 1
- 2
Abstract
Currently, large amounts of information exist in Web sites and various digital media. Most of them are in natural lan-guage. They are easy to be browsed, but difficult to be understood by computer. Chunk parsing and entity relation extracting is important work to understanding information semantic in natural language processing. Chunk analysis is a shallow parsing method, and entity relation extraction is used in establishing relationship between entities. Because full syntax parsing is complexity in Chinese text understanding, many researchers is more interesting in chunk analysis and relation extraction. Conditional random fields (CRFs) model is the valid probabilistic model to segment and label sequence data. This paper models chunk and entity relation problems in Chinese text. By transforming them into label solution we can use CRFs to realize the chunk analysis and entities relation extraction.
- E. C. Mary and J. M. Raymond, “Relational Learning of Pattern-match Rules for Information Extraction,” Ph.D. Thesis, University of Texas, Austin, 1998.
- S. Stephen, “Learning Information Extraction Rules for Semi-Structured and Free Text,” Machine Learning, Vol. 34, No. 13, 1999, pp. 233-272.
- D. Freitag and A. McCallum, “Information Extraction with HMM Structures Learned by Stochastic Optimization,” Proceedings of 18th Conference on Artificial Intelligence, AAAI Press, Edmonton, 2002, pp. 584-589.
- R. Souyma and C. Mark, “Representing Sentence Structure in Hidden Markov Models for Information Extraction,” Proceedings of the Seventeenth International Joint Conference on Artificial Intelligence, Morgan Kaufmann, Washington, 2001, pp. 1273-1279.
- T. Scheffer, C. Decomain and S. Wrobel, “Active Hidden Markov Models for Information Extraction,” Proceedings of the Fourth International Symposium on Intelligent Data Analysis, Springer, Lisbon, 2001, pp. 301-109.
- D. Freitag, A. McCallum and F. Pereira, “Maximum En-tropy Markov Models for Information Extraction and Segmentation,” Proceedings of the Seventeenth Interna-tional Conference on Machine Learning, Morgan Kauf-mann, San Francisco, 2000, pp. 591-598.
- H. L. Sun and S. W. Yu, “Shallow Parsing: An Over-view,” Contemporary Linguistics, 2000.
- S. Miller, M. Crystal, H. Fox, L. Ramshaw, R. Schwartz, R. Stone and R. Weischedel, “Algorithms that Learn to Extract Information-BBN: Description Of The SIFT Sys-tem as Used for MUC-7, Proceedings of MUC-7, Fairfax, 1998.
- J. Lafferty, A. McCallum and F. Pereira, “Conditional Random Fields: Probabilistic Models for Segmenting and Labeling Sequence Data,” Proceedings of the International Conference on Machine Learning (ICML), 2001, pp. 282-289.
- Y. Y. Luo and D. G. Huang, “Chinese Word Segmentation Based on the Marginal Probabilities Generated by CRFs,” Journal of Chinese Information Processing, Vol. 23, No. 5, 2009, pp. 3-8.
- M.-C. Hong, K. Zhang, J. Tang and J.-Z. Li “A Chinese Part-of-Speech Tagging Approach Using Conditional Random Fields,” Computer Science, Vol. 33, No. 10, 2006, pp. 148-152.
- S. P. Abney and C. Tenny, “Parsing by Chunks. Principle based Parsing: Computation and Psycholinguistics,” Kluwer Academic Publishers, Dordrecht, 1991, pp. 257-278.
- F. Erik, “Tjong Kim Sang and Sabine Buch holz. Intro-duction to the Conll-2000 Shared Task: Chunking,” Pro-ceedings of CoNLL-2000 and LLL2000, Lisbin, 2000, pp. 127-132.