CEFR-based Lexical Simplification Dataset
抄録
This study creates a language dataset for lexical simplification based on Common European Framework of References for Languages (CEFR) levels (CEFR-LS). Lexical simplification has continued to be one of the important tasks for language learning and education.There are several language resources for lexical simplification that are available for generating rules and creating simplifiers using machine learning. However, these resources are not tailored to language education with word levels and lists of candidates tending to be subjective. Different from these, the present study constructs a CEFR-LS whose target and candidate words are assigned CEFR levels using CEFR-J wordlists and English Vocabulary Profile, and candidates are selected using an online thesaurus. Since CEFR is widely used around the world, using CEFR levels makes it possible to apply a simplification method based on our dataset to language education directly. CEFR-LS currently includes 406 targets and 4912 candidates. To evaluate the validity of CEFR-LS for machine learning, two basic models are employed for selecting candidates and the results are presented as a reference for future users of the dataset.
収録刊行物
-
- Proceedings of International Conference on Language Resources and Evaluation
-
Proceedings of International Conference on Language Resources and Evaluation 11 3254-3258, 2018-05
European Language Resources Association
- Tweet
詳細情報 詳細情報について
-
- CRID
- 1050298532705123456
-
- NII論文ID
- 120006477471
-
- HANDLE
- 2324/1932611
-
- 本文言語コード
- en
-
- 資料種別
- journal article
-
- データソース種別
-
- IRDB
- CiNii Articles