Team Ai
Modelpublic

StanfordSCALE/assertion_sentence_expresses_certainty_or_emphasis

sourceHugging Faceupdated 22d agoView on Hugging Face
0likes42downloads
Model Card

Assertion: sentence expresses certainty or emphasis

This classifier was trained for EduBehaviors: Assertion-based schemas for auditable dialogue coding and is usable through the Python package EduBehaviors-kit. This classifier was trained on an LLM-annotated subset of teacher utterances from the TalkMoves Dataset. See the Datasets section below for more information.


Training Details

Datasets

This model's columns are assertion_sentence_expresses_certainty_or_emphasis and split_sentence_expresses_certainty_or_emphasis.

Base rate (share of rows labeled as True): 9.5% overall — 9.2% train, 10.0% dev, 9.8% test.

Labels and annotation

Labels were generated with LLM annotators. Krippendorff's alpha for this assertion is 0.195.

Hyperparameters

ParameterValue
Base model (body)sentence-transformers/paraphrase-mpnet-base-v2
HeadLogisticRegression
Body learning rate2e-05
Head learning rate0.01
Batch size16 (contrastive phase) / 32 (head)
Epochs10
Max steps5000 (contrastive phase)
Eval max steps100
Seed20260904
Mixed precisionenabled on GPU

Evaluation

Results

SplitnBase ratePrecisionRecallF1 (positive class)ROC-AUCAverage precision
dev85810.0%0.4390.4190.4290.7950.375
test2,1449.8%0.4550.4290.4410.8280.417

Limitations

  • —Labels come from LLM annotators, not human coders. Agreement with Krippendorff's Alpha is 0.195; this is poor.This model's predictions and the underlying data are unreliable.
  • —Trained on teacher utterances only. Behaviour on student speech is untested.

How to Use

Message Structure

The model was trained on text built as:

{utterance}

The utterance is passed through as-is.

Running instructions

bash
pip install setfit
python
from setfit import SetFitModel

model = SetFitModel.from_pretrained("StanfordSCALE/assertion_sentence_expresses_certainty_or_emphasis")

text = 'I mean this one makes the most circles and its the most colorful'
model.predict([text])        # -> array([1]) when the assertion holds
model.predict_proba([text])  # -> [[P(no), P(yes)]]

Citation

bibtex
@misc{assertion_sentence_expresses_certainty_or_emphasis,
  author = {Stanford SCALE Initiative},
  title  = {Assertion classifier: sentence expresses certainty or emphasis},
  year   = {2026},
  url    = {https://huggingface.co/StanfordSCALE/assertion_sentence_expresses_certainty_or_emphasis}
}