SEMI-AFFINE MODES AND MODALS
-
- ROMANOWSKA A. B.
- FACULTY OF MATHEMATICS AND INFORMATION SCIENCES, WARSAW UNIVERSITY OF TECHNOLOGY
Search this article
Abstract
Data sparseness or overfitting is a serious problem in natural language processing employing machine learning methods. This is still true even for the maximum entropy (ME) method, whose flexible modeling capability has alleviated data sparseness more successfully than the other probabilistic models in many NLP tasks. Although we usually estimate the model so that it completely satisfies the equality constraints on feature expectations with the ME method, complete satisfaction leads to undesirable overfitting, especially for sparse features, since the constraints derived from a limited amount of training data are always uncertain. To control overfitting in ME estimation, we propose the use of box-type inequality constraints, where equality can be violated up to certain predefined levels that reflect this uncertainty. The derived models, inequality ME models, in effect have regularized estimation with L_1 norm penalties of bounded parameters. Most importantly, this regularized estimation enables the model parameters to become sparse. This can be thought of as automatic feature selection, which is expected to improve generalization performance further. We evaluate the inequality ME models on text categorization datasets, and demonstrate their advantages over standard ME estimation, similarly motivated Gaussian MAP estimation of ME models, and support vector machines (SVMs), which are one of the state-of-the-art methods for text categorization.
Journal
-
- Machine Learning
-
Machine Learning 61 (1), 159-194, 2005-01-01
Springer Science + Business Media
- Tweet
Keywords
- affine spaces
- barycentric algebras
- modes
- cancellative modes
- modes embeddable into (semi)modules
- semimodules over commutative semirings
- idempotent (sub)reducts of semimodules over commutative semirings
- modes of submodes
- the structure of modes of submodes
- identities true in modes of submodes
- semilattice modes
- modals
- distributive and entropic modals
- duality for modes and modals
Details 詳細情報について
-
- CRID
- 1573668927285051136
-
- NII Article ID
- 120000861188
-
- NII Book ID
- AA1150654X
-
- ISSN
- 13460862
-
- Web Site
- http://hdl.handle.net/10119/3305
-
- Text Lang
- en
-
- Data Source
-
- CiNii Articles