datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
teaching_slidesThe dataset is currently publicly available. The calibration data is contained in the files below. Please do not click the dataset card, as it only contains images. Kindly cite our paper.
This dataset is a Chinese PPT dataset。Our paper is entitled Dual Prompt-Gate: A Training-Free Dual-Expert Framework for Educational PPT Digitization and Structured Content Extraction
HS_direct_teaching_260806_cam3_remaparc-agi-3-teaching-suite
ARC-AGI-3 teaching suite
100 grid-reasoning tasks across 16 puzzle families, with two splits:
split
rows
answers in this dataset
why
taught
80
yes
16 families x 5 worked examples, each documented in INSTRUCTIONS.md and reasoned through in CHAIN_OF_THOUGHT.md
heldout
20
no
16 unseen instances of the same families, plus 4 tasks that are provably not solvable by reasoning from their examples
The teaching material is what makes this a teaching suite rather than a… See the full description on the dataset page: https://huggingface.co/datasets/dnhkng/arc-agi-3-teaching-suite.cs3-ai-teaching-kit
CS3 AI Teaching Kit
38 lessons for the Jetson teaching kit, ready to teach and ready to run.
Open index.html for the lesson list.
There is also an advanced course/ with 10 ready-to-run programs (scripts
only, no decks) that combine the camera and sensor courses — see its README.
This offline package is one component of the hybrid CS3 AI Teaching Kit developed for the 2026 CS3 Research Experience for Teachers program at Columbia University. The companion online AI Training… See the full description on the dataset page: https://huggingface.co/datasets/mehmetkeremturkcan/cs3-ai-teaching-kit.teaching_tools_2025HS_direct_teaching_260806_freedrive_originalHS_direct_teaching_260806_cam2_remapkinesthetic_teaching_box_flip_20260707_175626This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
14
],
"names": [
"left_joint_1.pos",
"left_joint_2.pos",
"left_joint_3.pos",
"left_joint_4.pos",
"left_joint_5.pos"… See the full description on the dataset page: https://huggingface.co/datasets/collected-ai/kinesthetic_teaching_box_flip_20260707_175626.Teaching-Engine-Datasettplegacy-teachings
True Parents Legacy Teaching Archive
Digital archive of 3549 passages — sermons, speeches, prayers, and book excerpts by Sun Myung Moon (1920–2012) and Hak Ja Han Moon (1943–present), spanning 1946–2012.
Structure
Each record contains:
Field
Type
Description
id
string
URL slug (unique identifier)
title
string
Title of the sermon, speech, or passage
author
string
Speaker name
date
string
Publication date (YYYY-MM-DD)
year
string
Year only
tags… See the full description on the dataset page: https://huggingface.co/datasets/JonAuror/tplegacy-teachings.lerobot_teaching_HS_PicknPlace_1004This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "ur5e_freedrive",
"total_episodes": 100,
"total_frames": 76157,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/woodr/lerobot_teaching_HS_PicknPlace_1004.betbetter-teaching-dataset-2026
Bet Better Teaching Dataset 2026
One row per completed game across six leagues (NFL, MLB, NBA, NHL, AFL, NRL), built from Bet Better's own score, model and closing-price tables. Each row carries the final score, the model's projected margin and total before the game, the closing bookmaker prices, and the graded result, so a class can test questions like "are bookmaker prices well calibrated?" or "does a projected margin beat the closing line?" without any scraping.
DOI… See the full description on the dataset page: https://huggingface.co/datasets/edushinka/betbetter-teaching-dataset-2026.teaching-sfm-a01260048-supplementary
Supplementary Reproducibility Package — Manuscript A01260048
This package accompanies the article “Teaching Feature-Based 2D-to-3D Reconstruction for Metaverse-Oriented Content Generation: An OpenCV Structure-from-Motion Methodology for Computer Science Students.” It provides the executable OpenCV benchmark, pinned dependencies, machine-readable results, representative scene images, and figure sources used in the study.
Contents
benchmark_sfm.py: deterministic… See the full description on the dataset page: https://huggingface.co/datasets/malakhovks/teaching-sfm-a01260048-supplementary.teaching-cartpole-10k-v1
TeachingCartpole 10k (v1)
Offscreen-rendered single-pole CartPole episodes (RGB frames, physical state, and the
applied horizontal force) in the stable-worldmodel
HDF5 layout used to train LeWorldModel. Collected for
COMP 765 (McGill) Assignment 1, Q2(e). Models trained on it:
Autobrik/lewm-teaching-cartpole-fs5,
Autobrik/lewm-teaching-cartpole-fs1.
Contents
10,000 episodes, 1,507,974 transitions, split by episode (seed 43):
File
Episodes
Noisy LQR
Random… See the full description on the dataset page: https://huggingface.co/datasets/Autobrik/teaching-cartpole-10k-v1.Teaching-Dataset
Teaching Dataset
This dataset contains conversational data for training teaching and instruction-following models.
Dataset Structure
The dataset is provided in JSONL format where each line contains a conversation with multiple turns.
Each conversation consists of:
role: Either "user" or "assistant"
content: List containing the message content with type and text
Example
[
{
"role": "user",
"content": [{"type": "text", "text": "What is 2+2?"}]
}… See the full description on the dataset page: https://huggingface.co/datasets/drwlf/Teaching-Dataset.so100_teachingThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 2,
"total_frames": 1770,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/BENVBAI/so100_teaching.mdr-tb-india-2009-2011-teaching-dataset
MDR-TB India 2009-2011 Teaching Dataset
This repository contains a reproducibility-focused teaching package based on:
Nair D et al. Predictors of unfavourable treatment outcome in patients with multidrug-resistant tuberculosis in India. Public Health Action. 2017;7(1):32-38. doi:10.5588/pha.16.0055.
Files
article_table2_selected_baseline.csv (real extracted aggregate values)
article_table4_outcomes.csv (real extracted aggregate values)
article_table6_adrs.csv (real… See the full description on the dataset page: https://huggingface.co/datasets/hssling/mdr-tb-india-2009-2011-teaching-dataset.HS_direct_teaching_260806_cam1_remapcefr-english-teaching-corpus
The English Classroom Corpus
A 600-hour, CEFR-graded English language teaching corpus (A1–C2) with aligned native-speaker audio. Single author. Clean rights. Available under commercial licence.
⚠️ The data is not hosted in this repository. This is a dataset card for discovery purposes. Licensing enquiries: the-english-classroom.com/licensing or jennifer@the-englishclassroom.com
Dataset summary
The English Classroom Corpus is a complete English language… See the full description on the dataset page: https://huggingface.co/datasets/jennifertec/cefr-english-teaching-corpus.train_teaching_arithmetic_1-3digits_10000examplesHS_data_analysis_teaching_0829-freedrive-cam4teaching_motivational_quotesXLogoMiniProg
XLogoMiniProg: Program Synthesis Benchmark for Visual Programming in XLogoOnline Environment
Dataset |
Code | Paper
This repo contains the datasets for the ACL 2025 paper "Program Synthesis Benchmark for Visual Programming in XLogoOnline Environment"
Dataset
The datasets include the following files:
xlogomini-dataset-test.json: test set with 1085 samples
xlogomini-dataset-train.json: train set with 87,053 samples
xlogomini-dataset-validation.json: validation set… See the full description on the dataset page: https://huggingface.co/datasets/machine-teaching-group/XLogoMiniProg.teaching-claude-why
geodesic-research/teaching-claude-why
Auto-generated by dataset-builder.
Each config below is a separate dataset produced from a versioned YAML build
config. Load with:
from datasets import load_dataset
ds = load_dataset("geodesic-research/teaching-claude-why", "<config_name>", revision="<commit-sha>")
Pin revision= to the specific commit SHA you want; without it, you get the
current HEAD of the dataset repo, which may change when the builder re-pushes.
Configs… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/teaching-claude-why.lerobot_teaching_HS_PicknPlace_1004_replay_mergedadaption-bolivia-rural-teaching-guides
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-bolivia_rural_teaching_guides
This dataset contains pedagogical guides and dialogue scripts for multigrade teachers in rural Bolivia, covering subjects like math, science, and community organization. Each entry provides concrete, culturally relevant activities using local materials such as clay, seeds, and recycled items to engage students of varying ages. The content emphasizes the… See the full description on the dataset page: https://huggingface.co/datasets/Oliver369X/adaption-bolivia-rural-teaching-guides.africa-worldbank-teaching-staff-compensation-as-a-percentage-of-total-expenditure-in-tertiary-pu
Teaching staff compensation as a percentage of total expenditure in tertiary public institutions (%) | Africa (World Bank — Education Statistics) | Africa (World Bank)
Size category: n<1K - Formats: parquet - Sector: education - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-worldbank-teaching-staff-compensation-as-a-percentage-of-total-expenditure-in-tertiary-pu.adaption-catholic-moral-teachings
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-catholic_moral_teachings
This dataset consists of instruction and response pairs focused on core Catholic moral theology and ethical principles. It covers foundational topics such as the cardinal and theological virtues, the Ten Commandments, and Catholic Social Teaching. The entries provide doctrinally grounded explanations and practical applications of Catholic morality to… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/adaption-catholic-moral-teachings.teaching_equationsafrica-worldbank-teaching-staff-compensation-as-a-percentage-of-total-expenditure-in-pre-primary
Teaching staff compensation as a percentage of total expenditure in pre-primary public institutions (%) | Africa (World Bank — Education Statistics) | Africa (World Bank)
Size category: n<1K - Formats: parquet - Sector: education - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-worldbank-teaching-staff-compensation-as-a-percentage-of-total-expenditure-in-pre-primary.
