Team Ai
Modelpublic

SamAgnoli/deberta-v3-base-spatial-language-detection

sourceHugging Facemitupdated 1mo agoView on Hugging Face
1likes18downloads
Model Card

DeBERTa-v3-base — Spatial Language Detection

A fine-tuned `microsoft/deberta-v3-base` for word-level spatial-language detection: given an utterance and a target word, it decides whether that word is being used as spatial language (location, direction, or a spatial relationship) in that context — 1 = spatial, 0 = not. The same word can be spatial in one utterance ("go up the ramp") and not in another ("what's up?"), so the model always judges a word together with its sentence.

For the full pipeline (dictionary gating, calibrated confidence, evaluation) and example datasets, see the GitHub repo: https://github.com/SamAgnoli/spatial-language-classifier

Input format

This is a sentence-pair classifier: pass the utterance as the first segment and the target word as the second — tokenizer(utterance, target_word). Passing a whole sentence on its own is not how the model was trained and gives unreliable results.

Labels

idlabelmeaning
0not_spatialword is not spatial language
1spatialword is spatial language

Usage

python
from transformers import AutoModelForSequenceClassification, AutoTokenizer
import torch

model_id = "SamAgnoli/deberta-v3-base-spatial-language-detection"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id)

utterance   = "The cat is sitting on top of the bookshelf."
target_word = "top"                       # the word you want judged
inputs = tokenizer(utterance, target_word, return_tensors="pt", truncation=True)
with torch.no_grad():
    logits = model(**inputs).logits
pred = logits.argmax(-1).item()
print(model.config.id2label[pred])        # -> "spatial"

Or with a pipeline (note the text / text_pair keys):

python
from transformers import pipeline

clf = pipeline("text-classification",
               model="SamAgnoli/deberta-v3-base-spatial-language-detection")
print(clf({"text": "The cat is sitting on top of the bookshelf.", "text_pair": "top"}))

Training data

Parent–child conversational reflections recorded at a museum tinkering exhibit (Polinsky et al., 2023): 33,284 word tokens across 152 conversations. A bag-of-words spatial dictionary (Cannon et al., 2007; Polinsky et al., 2023) selected 3,516 candidate tokens (10.6% of all words, 176 unique word types) which trained human coders then judged in utterance context; 1,455 were coded spatial (41.4% of candidates, 4.4% of all words). Every remaining token is labelled non-spatial.

Training

  • —Base model: microsoft/deberta-v3-base
  • —Task: binary sentence-pair classification (a word, in its utterance, spatial vs. not)
  • —Split: group-aware 70/15/15 by speaker session (no session spans splits) — 23,100 / 5,714 / 4,470 tokens
  • —Hyperparameters: 2 epochs · lr 2e-5 · batch 16 · weight decay 0.01 · warmup ratio 0.1 (289 steps) · max_length 128 · fp16 · seed 42 · best checkpoint by F1
  • —Class imbalance: no class weighting or resampling. The dictionary gate already raises the positive rate from 4.4% of all words to 41.4% of candidates, and checkpoint selection used minority-class F1 rather than accuracy.
  • —Framework: 🤗 Transformers 5.16.1

Evaluation (held-out test set)

Reported for two views: dictionary candidates only (the meaningful view — words a spatial dictionary flags as plausibly spatial) and overall (every word, dominated by trivially non-spatial tokens).

Candidates only — 526 words (315 not-spatial, 211 spatial)

classprecisionrecallF1support
not_spatial0.9460.8980.922315
spatial0.8590.9240.890211

Spatial candidate accuracy: 0.909 (478 of 526 words correct) · macro-F1 0.906 · Cohen's κ 0.812

Reading it per class: the model catches 92.4% of truly-spatial words (recall) at 85.9% precision; for non-spatial words it's 89.8% recall at 94.6% precision. (sensitivity 0.924, specificity 0.898, PPV 0.859, NPV 0.946.)

Confusion matrix (candidates): TN 283 · FP 32 · FN 16 · TP 195.

Overall — every word, 4,470 tokens

accuracy 0.988 · spatial-F1 0.882 · Cohen's κ 0.876

The two views share the same spatial predictions (211 spatial words, same 195 caught). Only the non-spatial pool differs, which is why "overall" accuracy looks higher — it's padded with ~3,900 easy non-candidate words the model trivially gets right. Judge the model by the candidates-only view.

Calibration

Raw probabilities are over-confident, so a post-hoc temperature scaling factor (T = 1.461, fit on the 592 validation candidates) rescales them into a calibrated P(spatial) you can read literally. Temperature scaling is monotonic, so the hard 0/1 decision is unchanged — verified by assertion on every test row.

On the candidate subset this lowers expected calibration error from .059 to .050 (uniform binning; .044 → .041 quantile) and the Brier score from .074 to .069. See section 7.5 of the repo notebook.

Intended use & limitations

  • —Built for per-word spatial judgments within an English utterance. In production it is paired with a spatial-dictionary gate that selects candidate words; words the dictionary misses (e.g., misspellings) are never sent to the model.
  • —Trained on parent–child tinkering-reflection speech — performance on other domains, genres, or languages is not guaranteed.
  • —The data is strongly imbalanced (4.4% spatial overall); judge quality by the candidates-only view, not the overall numbers.
  • —The held-out test set is a single group-aware split of 526 candidate words. Metrics carry meaningful sampling uncertainty at that size; treat differences of a point or two as noise.
  • —Review predictions before relying on them in downstream systems.

Citation

Agnoli, S., Shi, Q., & Uttal, D. (2026). Fine-tuning DeBERTa-v3 to automate spatial language classification. Conference on Artificial Intelligence in Measurement and Education (AIME-Con).