Team Ai
Datasetpublic

Wajinimi/Children_Intent_Classification

MAMA Communicative Intent Dataset (INCA-A Annotated) Overview The MAMA Communicative Intent Dataset is a linguistically annotated corpus of child utterances designed to support research in child-centred Natural Language Processing (NLP) and communicative intent recognition in early language development. The dataset contains 10,800 child utterances annotated using the INCA Communicative Coding System (Ninio et al., 1994), a developmental framework that identifies… See the full description on the dataset page: https://huggingface.co/datasets/Wajinimi/Children_Intent_Classification.

sourceHugging Facecc-by-4.0updated 7mo agoView on Hugging Face
0likes152downloads
README.md243 linesDownload Raw Back to root
1---2language:3- en4license: cc-by-4.05pretty_name: MAMA Communicative Intent Dataset (INCA-A Annotated)6tags:7- NLP8- child-language9- intent-classification10- conversational-ai11- developmental-linguistics12- dialogue13- child-speech14task_categories:15- text-classification16task_ids:17- intent-classification18---19 20# MAMA Communicative Intent Dataset (INCA-A Annotated)21 22## Overview23 24The **MAMA Communicative Intent Dataset** is a linguistically annotated corpus of child utterances designed to support research in **child-centred Natural Language Processing (NLP)** and **communicative intent recognition in early language development**.25 26The dataset contains **10,800 child utterances** annotated using the **INCA Communicative Coding System** (Ninio et al., 1994), a developmental framework that identifies the communicative functions underlying children's speech.27 28To make the dataset suitable for machine learning, the original INCA codes were mapped to **23 refined intent categories** representing distinct communicative behaviours in early child language.29 30This dataset was created as part of the **MAMA (Machine-Assisted Maternal Assistant)** research project, which investigates how artificial intelligence systems can better understand the communicative behaviour of young children.31 32Unlike many NLP corpora that normalise or correct non-standard language, this dataset **preserves authentic developmental linguistic features**, including telegraphic speech and missing grammatical markers.33 34---35 36# Dataset Summary37 38| Property | Value |39|--------|------|40| Total utterances | 10,800 |41| Total labelled instances | 11,410 |42| Intent categories | 23 |43| Annotation framework | INCA Communicative Coding System |44| Language | English |45| Task | Intent Classification |46 47The dataset reflects **naturalistic child language**, resulting in class imbalance typical of real-world conversational data.48 49Example distribution:50 51| Intent | Frequency |52|------|------|53| Observation / Reference | 2,964 |54| Narrative / Storytelling | 1,664 |55| Comfort | 4 |56 57---58 59# Annotation Framework60 61The dataset is based on the **INCA Communicative Coding System**, introduced in:62 63> Ninio, A., Snow, C., Pan, B., & Rollins, P. (1994).  64> *Classifying communicative acts in children's interactions.*65 66The INCA system categorises **communicative functions** in children's speech rather than grammatical structure alone.67 68In this dataset, INCA codes were mapped into **refined NLP intent categories** suitable for supervised machine learning.69 70---71 72# Mapping from INCA Codes to Refined Intent Categories73 74| INCA Category | INCA Code | Refined Intent Category |75|---------------|----------|-------------------------|76| Directing hearer’s attention | DHA / CL | Attention |77| Speech elicitation | EI, RT, EA | Imitation |78| Questions | QN, YQ, TQ | Question |79| Evaluation | ET | Excitement |80| Discussing related-to-present | DRP | Narrative or Storytelling |81| Discussing joint focus | DJF | Observation or Reference |82| Statements | WS | Desire or Action |83| Negotiating activity | NIA / DW | Disagreement or Correction |84| Marking | MRK | Gratitude |85| Comforting | CMO | Comfort |86| Directiveness | RP | Request |87| Directiveness | RD, CS | Refusal |88| Directiveness | GR | Explanation or Justification |89| Declaration | YD / AP | Agreement or Acknowledgment |90| Marking | MK | Greeting |91| Marking | EM | Distress or Pain |92| Marking | EN | Emotion |93| Fantasy discussion | DFW | Playtalk or Fantasy |94| Possession negotiation | PSS | Possession |95| Request / Suggest | RP | Need |96| Dare / Challenge | DR | Command |97| Disapprove / Protest | DS, ED, DW | Complaint |98 99---100 101# Annotation Protocol102 103Annotation followed a **two-stage validation procedure**.104 105### Stage 1 — Initial Annotation106 107All utterances were initially labelled by the **primary researcher** using the INCA communicative coding framework.108 109### Stage 2 — Expert Re-annotation110 111To strengthen validity, the dataset was independently reviewed by two domain experts:112 113- **Developmental Psychologist**114- **Experienced Early-Years Teacher**115 116This ensured both **developmental theoretical grounding** and **practical child-language expertise**.117 118---119 120# Inter-Annotator Reliability121 122Agreement between annotators was measured using **Cohen's Kappa (κ)**.123 124### Observed Agreement125 126\[127P_o = \frac{\sum C_{ii}}{N}128\]129 130Where:131 132- \(C_{ii}\) = number of rows where annotators assigned the same category  133- \(N\) = total number of annotated rows134 135### Cohen's Kappa136 137\[138\kappa = \frac{P_o - P_e}{1 - P_e}139\]140 141Where expected agreement is defined as:142 143\[144P_e = \sum \left(\frac{R_i}{N} \cdot \frac{C_i}{N}\right)145\]146 147Where:148 149- \(R_i\) = rows assigned to category \(i\) by annotator 1  150- \(C_i\) = rows assigned to category \(i\) by annotator 2151 152The resulting score was:153 154**κ = 0.81**155 156According to **Landis and Koch (1977)**, this represents **almost perfect agreement**, indicating strong reliability in intent categorisation.157 158---159 160# Linguistic Characteristics161 162## Average Utterance Length163 164Average utterance length was computed as:165 166\[167\text{Average utterance length} =168\frac{\sum_{i=1}^{N} (\text{token length of utterance}_i)}{N}169\]170 171Analysis revealed that:172 173- **Explanation or Justification**174- **Narrative or Storytelling**175- **Desire or Action**176 177tend to produce **longer utterances**, indicating more verbose communicative behaviour.178 179In contrast:180 181- **Observation or Reference**182 183typically contains **shorter utterances**, reflecting concise descriptions of objects or events in the shared environment.184 185---186 187## Lexical Diversity188 189Lexical diversity analysis showed variation across communicative intents.190 191Categories such as:192 193- **Observation**194- **Narrative or Storytelling**195 196exhibited **higher vocabulary diversity**, reflecting descriptive language use.197 198Conversely:199 200- **Agreement or Acknowledgement**201 202showed **low lexical diversity**, as these responses often rely on short, formulaic expressions such as: Yes, Okay, Yeah203 204 205---206 207# Developmental Linguistic Features208 209The dataset preserves several characteristics typical of **early child language**.210 211| Feature | Percentage |212|------|------|213| Telegraphic speech | 19.67% |214| Shortened forms | 8.32% |215| Missing function words | 9.52% |216 217Telegraphic speech refers to utterances dominated by **content words** while omitting grammatical elements.218 219Examples: Want Juice, Doggy Running, Baby Sleep220 221 222Preserving these patterns allows models trained on this dataset to better interpret **non-standard developmental speech**.223 224---225 226# Intended Use227 228The dataset is intended for research in:229 230- Child-centred NLP231- Intent classification232- Developmental linguistics233- Child-robot interaction234- Conversational AI for children235- Human–AI interaction in early childhood environments236 237---238 239 240 241 242 243