llm-annotated
assertions_llm_annotated_talkmoves
Bottom-Up Assertion Labels
This is a subset of the TalkMoves Dataset of K-12 mathematics lesson transcripts, created for EduBehaviors: Assertion-based schemas for auditable dialogue coding. This dataset contains teacher utterances labeled with assertions, distinct behaviors or attributes of utterances that may serve as features for the modeling of larger constructs. Labels in this dataset are LLM-generated, with models reported within the dataset itself. This dataset was used to… See the full description on the dataset page: https://huggingface.co/datasets/StanfordSCALE/assertions_llm_annotated_talkmoves.summeval-annotated-latest
SummEval-LLMEval Dataset
Overview
The original SummEval dataset (Fabbri et al., 2021) consists of 1,600 summaries annotated by human expert evaluators using a 5-point Likert scale across 4 criteria: coherence, consistency, fluency, and relevance. These 1,600 summaries are based on 100 source articles from the CNN/DailyMail dataset (Hermann et al., 2015). For each source article, SummEval collects 16 summaries generated by 16 different automatic summarization systems. Each… See the full description on the dataset page: https://huggingface.co/datasets/bay-calibration-llm-evaluators/summeval-annotated-latest.pandalm-annotated-fullllmbar-annotated-latest
LLMBar-Select Dataset
Introduction
The LLMBar-Select dataset is a curated subset of the original LLMBar dataset introduced by Zeng et al. (2024). The LLMBar dataset consists of 419 instances, each containing an instruction paired with two outputs: one that faithfully follows the instruction and another that deviates while presenting superficially appealing qualities. It is designed to evaluate LLM-based evaluators more rigorously and objectively than previous benchmarks.… See the full description on the dataset page: https://huggingface.co/datasets/bay-calibration-llm-evaluators/llmbar-annotated-latest.hanna-annotated-fullllmbar-annotated-latest
