Team Ai
Datasetpublic

star092304/CEFR-Annotated-WordNet

CEFR-Annotated WordNet: LLM-Based Proficiency-Guided Semantic Database for Language Learning Paper: CEFR-Annotated WordNet: LLM-Based Proficiency-Guided Semantic Database for Language Learning Authors: Masato Kikuchi, Masatsugu Ono, Toshioki Soga, Tetsu Tanabe, Tadachika Ozono Overview CEFR-Annotated WordNet is a comprehensive semantic database that maps WordNet 3.0 senses and SemCor 3.0 instances to CEFR (Common European Framework of Reference for Languages)… See the full description on the dataset page: https://huggingface.co/datasets/star092304/CEFR-Annotated-WordNet.

sourceHugging Faceupdated 4mo agoView on Hugging Face
1likes384downloads
Dataset Card

CEFR-Annotated WordNet: LLM-Based Proficiency-Guided Semantic Database for Language Learning

Paper: CEFR-Annotated WordNet: LLM-Based Proficiency-Guided Semantic Database for Language Learning

Authors: Masato Kikuchi, Masatsugu Ono, Toshioki Soga, Tetsu Tanabe, Tadachika Ozono


Overview

CEFR-Annotated WordNet is a comprehensive semantic database that maps WordNet 3.0 senses and SemCor 3.0 instances to CEFR (Common European Framework of Reference for Languages) proficiency levels. This dataset enables LLM-based analysis of word complexity and is specifically designed for language learning applications.

The dataset provides multiple data formats:

  • —Lookup tables for sense-level and instance-level CEFR annotations
  • —Fine-tuning datasets (CSV and JSONL) for training models on CEFR proficiency assessment
  • —Two specialized model fine-tuning splits: DeBERTa and LLM formats

Directory Structure

CEFR-Annotated-WordNet/
├── cefr-level_annotation/        # Core annotation files
│   ├── semcor_instance_cefr.tsv  # SemCor instance → CEFR mapping
│   └── wordnet_sensekey_cefr.tsv # WordNet sense key → CEFR mapping
├── fine_tuning_debert/            # DeBERTa model fine-tuning data
│   └── semcor_mapped.csv         # CSV format with context and labels
├── fine_tuning_llm/               # LLM model fine-tuning data
│   ├── train.jsonl               # Training split (chat format)
│   └── valid.jsonl               # Validation split (chat format)
├── LICENSES/                      # License files
│   ├── CC-BY-4.0.txt
│   ├── WORDNET-3.0-LICENSE.txt
│   └── SEMCOR-3.0-LICENSE.txt
├── NOTICE                         # Attribution and copyright notice
└── README.md                      # This file

Dataset Files

1. CEFR-Level Annotation Files (cefr-level_annotation/)

These are core lookup tables mapping linguistic units to CEFR proficiency levels.

semcor_instance_cefr.tsv
  • —Format: Tab-separated values (TSV)
  • —Columns:
  • —instance_id: SemCor instance identifier (format: d###.s###.t###)
  • —CEFR_level: CEFR proficiency level (A1, A2, B1, B2, C1, C2)
  • —Example entries:
  d000.s000.t000	A1
  d000.s000.t002	C1
  d000.s000.t003	B2
  d000.s001.t004	B2
  • —Use case: Map SemCor instances to proficiency levels for corpus-based analysis or instance-level sense disambiguation.
  • —Size: Comprehensive coverage of SemCor 3.0 instances
wordnet_sensekey_cefr.tsv
  • —Format: Tab-separated values (TSV)
  • —Columns:
  • —sensekey: WordNet sense key (format: lemma%pos:sense:lexfile:sense_number)
  • —CEFR_level: CEFR proficiency level (A1, A2, B1, B2, C1, C2)
  • —Example entries:
  a.m.%4:02:00::	A1
  abandon%2:31:00::	B2
  abandon%2:31:01::	C1
  ability%1:07:00::	B1
  able%5:00:00:competent:00	C2
  • —Use case: Direct lookup for sense-level proficiency assessment in WordNet-based applications, lexical complexity estimation, and vocabulary ranking.
  • —Size: Complete WordNet 3.0 sense coverage

2. DeBERTa Fine-Tuning Data (fine_tuning_debert/)

This directory contains data formatted for training masked language models like DeBERTa.

semcor_mapped.csv
  • —Format: Comma-separated values (CSV) with headers
  • —Columns:
  • —instance_id: WordNet sense key identifier
  • —sentence: Full sentence context from SemCor
  • —target_word: The word to be classified
  • —lemma: Lemmatized form of the target word
  • —cefr_level: Target CEFR proficiency level (A1-C2)
  • —Example entry:
  say%2:32:00::,The Fulton County Grand Jury said Friday an investigation of Atlanta 's recent primary election produced `` no evidence '' that any irregularities took place .,said,said,B1
  investigation%1:09:00::,The Fulton County Grand Jury said Friday an investigation of Atlanta 's recent primary election produced `` no evidence '' that any irregularities took place .,investigation,investigation,B2
  • —Use case: Direct training data for sequence classification models like DeBERTa v3 for CEFR level prediction.
  • —Data characteristics:
  • —Sentences from real SemCor corpus
  • —Single target word per row with full context
  • —Balanced CEFR level distribution across training samples

3. LLM Fine-Tuning Data (fine_tuning_llm/)

This directory contains data formatted for instruction-following models and large language models.

train.jsonl
  • —Format: JSON Lines (one JSON object per line)
  • —Structure: Chat-style messages format compatible with GPT, Claude, and other LLMs
  • —Fields per entry:
  • —messages: Array containing conversational turns with roles and content
  • —role: One of system, user, assistant
  • —content: Text content for each role
  • —Example structure:
json
  {
    "messages": [
      {
        "role": "system",
        "content": "You are an expert in assessing English vocabulary proficiency levels according to the CEFR framework..."
      },
      {
        "role": "user",
        "content": "Assess the CEFR level of the word 'investigate' in this sentence: \"The Fulton County Grand Jury said Friday an investigation of Atlanta's recent primary election produced no evidence...\"\n\nProvide only the CEFR level (A1, A2, B1, B2, C1, C2)."
      },
      { "role": "assistant", "content": "B2" }
    ]
  }
  • —Use case: Training LLMs for CEFR proficiency assessment through instruction fine-tuning (LoRA, QLoRA, etc.).
  • —Size: Training split with thousands of examples
  • —Data characteristics:
  • —System prompt defining the CEFR assessment task
  • —Contextual sentences from SemCor corpus
  • —Target words with clear proficiency labels
valid.jsonl
  • —Format: Same JSON Lines format as train.jsonl
  • —Purpose: Validation/evaluation split for monitoring model performance during training
  • —Use case: Evaluate CEFR classification accuracy during fine-tuning iterations
  • —Size: Validation subset (typically 10-20% of training data)

CEFR Proficiency Levels

The dataset uses the Common European Framework of Reference for Languages (CEFR) scale with 6 levels:

LevelDescriptionExample
A1BeginnerBasic words: "is", "the", "be", "end"
A2ElementaryCommon everyday vocabulary: "have", "say", "day"
B1IntermediateGeneral vocabulary with some complexity: "ability", "investigate", "evidence"
B2Upper IntermediateAdvanced general vocabulary: "abandon", "jury", "produce"
C1AdvancedSophisticated vocabulary: "abnormal", "aboard", "abandon" (specialized senses)
C2ProficiencyRare, highly specialized vocabulary: "able" (competent sense), "oblige", "aberrant"

Use Cases

1. Lexical Complexity Assessment

Use the sense-level mappings (wordnet_sensekey_cefr.tsv) to analyze vocabulary complexity in texts.

2. CEFR-Based Word Sense Disambiguation

Combine with WordNet to disambiguate word senses while assessing proficiency levels.

3. Language Learning Applications

  • —Generate vocabulary recommendations based on CEFR levels
  • —Create proficiency-matched reading materials
  • —Build adaptive learning systems

4. Model Training

  • —Fine-tune DeBERTa using semcor_mapped.csv for sequence classification
  • —Instruction-tune LLMs using train.jsonl and valid.jsonl for CEFR assessment

5. Academic Research

  • —Study vocabulary distribution across proficiency levels
  • —Analyze sense-level complexity patterns
  • —Build datasets for NLP research on complexity assessment

Data Statistics

FileEntriesCoverage
wordnet_sensekey_cefr.tsvComplete WordNet 3.0 sensesAll synsets in WordNet
semcor_instance_cefr.tsvAll SemCor instancesSemCor 3.0 corpus
semcor_mapped.csvThousands of sense instancesRepresentative SemCor sentences
train.jsonlLarge training splitDiverse vocabulary and contexts
valid.jsonlValidation samplesHeld-out evaluation set

Licenses

This dataset combines multiple licensed resources:

  • —Our CEFR Annotations & Splits: CC BY 4.0
  • —Applies to: CEFR labels, dataset splits, and novel annotations
  • —Files: All CEFR labels in .tsv, .csv, and .jsonl files
  • —WordNet 3.0 (Princeton University): WordNet License
  • —Applies to: Sense keys and WordNet identifiers
  • —SemCor 3.0 (Princeton University): SemCor License
  • —Applies to: SemCor instance IDs and sentence text

Important: When using this dataset:

  1. 1.Cite the original paper (see header above)
  2. 2.Comply with WordNet and SemCor licenses
  3. 3.Respect CC BY 4.0 requirements for attribution

Citation

If you use this dataset in your research, please cite the original paper:

bibtex
@article{kikuchi2024cefr,
  title={CEFR-Annotated WordNet: LLM-Based Proficiency-Guided Semantic Database for Language Learning},
  author={Kikuchi, Masato and Ono, Masatsugu and Soga, Toshioki and Tanabe, Tetsu and Ozono, Tadachika},
  journal={arXiv preprint arXiv:2510.18466},
  year={2024}
}

Getting Started

Quick Start with Python

python
import pandas as pd

# Load sense-level CEFR mappings
sensekey_df = pd.read_csv('cefr-level_annotation/wordnet_sensekey_cefr.tsv', sep='\t')
print(f"Total senses: {len(sensekey_df)}")

# Load instance-level mappings
instance_df = pd.read_csv('cefr-level_annotation/semcor_instance_cefr.tsv', sep='\t')
print(f"Total instances: {len(instance_df)}")

# Load fine-tuning data for DeBERTa
train_df = pd.read_csv('fine_tuning_debert/semcor_mapped.csv')
print(f"Training examples: {len(train_df)}")

# Load LLM fine-tuning data
import json
llm_train = [json.loads(line) for line in open('fine_tuning_llm/train.jsonl')]
print(f"LLM training examples: {len(llm_train)}")

Training a DeBERTa Model

python
from transformers import AutoTokenizer, AutoModelForSequenceClassification
from datasets import Dataset

# Load data
df = pd.read_csv('fine_tuning_debert/semcor_mapped.csv')
dataset = Dataset.from_pandas(df)

# Prepare for training
tokenizer = AutoTokenizer.from_pretrained('microsoft/deberta-v3-base')
model = AutoModelForSequenceClassification.from_pretrained('microsoft/deberta-v3-base', num_labels=6)

# Continue with your training pipeline...

FAQ

Q: Can I use this dataset commercially? A: Yes, as long as you comply with the CC BY 4.0 license (attribution required).

Q: How is the CEFR level assigned to each sense? A: Please refer to the original paper for the annotation methodology.

Q: Can I contribute improvements or corrections? A: We welcome feedback! Please refer to the NOTICE file for contact information.

Q: What's the difference between instance-level and sense-level annotations? A: Sense-level (wordnet_sensekey_cefr.tsv) provides a fixed mapping for each WordNet sense. Instance-level (semcor_instance_cefr.tsv) provides CEFR labels for specific word uses in context from the SemCor corpus.


Contact & References

For questions or feedback about this dataset, please contact the authors through the information provided in the original paper: CEFR-Annotated WordNet: LLM-Based Proficiency-Guided Semantic Database for Language Learning

Related Resources: