datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LaSerena-Corpus-Geociencias
La Serena Digital Geo Corpus — Dominga EIA Dataset
Dataset Sci-Align de geología ambiental chilena basado en el expediente
de Evaluación de Impacto Ambiental del proyecto Dominga (SEIA, Región de Coquimbo).
Contenido
dominga_geo_align.jsonl — 402 registros Sci-Align (MinerU 3.1.15)
seia/ — documentos públicos del expediente Dominga (fuente primaria)
Licencia
CC-BY-4.0 — Fuente: SEIA Chile (acceso público)
Concurso
AGI4S — Pista 1: Creación de bases… See the full description on the dataset page: https://huggingface.co/datasets/Karlangaz/LaSerena-Corpus-Geociencias.verifiable-corpus
verifiable-corpus
This is the corpus from "Learning on the Job: Test-Time Curricula for Targeted Reinforcement Learning".
Code: https://github.com/jonhue/ttc
Introduction
We study how large language models (LLMs) can continually improve at reasoning on their target tasks at test-time. We propose an agent that assembles a task-specific curriculum, called test-time curriculum (TTC-RL), and applies reinforcement learning to continue training the model for its target task.… See the full description on the dataset page: https://huggingface.co/datasets/lasgroup/verifiable-corpus.last-translation-benchmark
Last Translation Benchmark
Abstract: For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases.
Standard benchmarks for machine translation evaluation are often either trivial (having few authentic mistakes) or unrealistic (overly synthetically contrived).
Furthermore, automatic translation metrics become less reliable and reward-hacked as models get stronger, and their outputs are… See the full description on the dataset page: https://huggingface.co/datasets/zouhar/last-translation-benchmark.toti-cakery-toolcall
Toti Cakery — Tool-Calling Fine-Tuning Dataset (Qwen3, v9)
Synthetic bilingual (Indonesian ~78% / English ~22%) SFT dataset for the Toti
Cakery WhatsApp chatbot: 13 LangChain tools (11 for customers, +2 owner-only
reports) and grounded answers from RAG FAQ context. Rows are built from the
live runtime code (SYSTEM_PROMPT, TOOL_REMINDER, tool schemas via
convert_to_openai_tool, _history_view, pertanyaan_dengan_konteks), so the
training prompt is byte-identical to what the model… See the full description on the dataset page: https://huggingface.co/datasets/LasagnaS/toti-cakery-toolcall.china-uncensored
China Uncensored / Anti-Authoritarian Information Integrity Dataset
A post-training dataset for improving censorship resistance, information integrity, and anti-authoritarian reasoning in open-source language models.
This dataset is intended for developers training models to handle politically sensitive China-related topics without reproducing authoritarian state propaganda, coercive narratives, or censorship-driven framing. It is especially relevant for open-source models that… See the full description on the dataset page: https://huggingface.co/datasets/lastbattle/china-uncensored.la-serena-digital-geo-corpus
La Serena Digital Geo Corpus — Dominga EIA Dataset
Dataset Sci-Align de geología ambiental chilena basado en el expediente
de Evaluación de Impacto Ambiental del proyecto Dominga (SEIA, Región de Coquimbo).
Contenido
dominga_geo_align.jsonl — 402 registros Sci-Align (MinerU 3.1.15)
seia/ — documentos públicos del expediente Dominga (fuente primaria)
Licencia
CC-BY-4.0 — Fuente: SEIA Chile (acceso público)
Concurso
AGI4S — Pista 1:… See the full description on the dataset page: https://huggingface.co/datasets/Karlangaz/la-serena-digital-geo-corpus.lastfm50
LastFM-50
This dataset expands
rohan2810/lastfm
from 20 to 50 candidates per example for finite-pool preference-optimization
experiments.
Construction
For every example, the original 20-candidate pool is preserved. Thirty
additional artists are sampled deterministically from the 4,606-artist
source candidate universe using seed 1958. New candidates exclude
the true item, the existing candidates, and artists in the listening history.
The resulting 50 candidates are… See the full description on the dataset page: https://huggingface.co/datasets/rohan2810/lastfm50.1984
Dataset Card for 1984
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/LastXuanZz/1984/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/LastXuanZz/1984.my-distiset-833e3cd0
Dataset Card for my-distiset-833e3cd0
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/LastXuanZz/my-distiset-833e3cd0/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/LastXuanZz/my-distiset-833e3cd0.my-distiset-e7f79bdc
Dataset Card for my-distiset-e7f79bdc
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/LastXuanZz/my-distiset-e7f79bdc/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/LastXuanZz/my-distiset-e7f79bdc.lastbox-survival-dialogues
LastBox survival dialogues
Training and evaluation data for the
LastBox offline survival
assistant on Raspberry Pi 5. Used to fine-tune
norecyc/lastbox-gemma4-e2b-sft-v3
(Kaggle "Gemma 4 Good" Hackathon 2026 submission) and the post-deadline
norecyc/lastbox-gemma4-e2b-v6-toolprior
checkpoint.
Files
File
Lines
Purpose
train_v2.jsonl
1 034
Main SFT training set (full tool-use traces)
val_v2.jsonl
114
Held-out validation
golden_en.jsonl
25
Agent-level eval… See the full description on the dataset page: https://huggingface.co/datasets/norecyc/lastbox-survival-dialogues.task078_all_elements_except_last_i
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task078_all_elements_except_last_i
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task078_all_elements_except_last_i.the-last-of-us-instruction-dataset
🧠 The Last of Us QA Dataset
This dataset was created from scratch to train Question Answering (QA) models specializing in the universe of The Last of Us.
📌 Description
The dataset contains question-and-answer pairs based on information from The Last of Us universe, including characters, story, events, and narrative context.
Its main purpose is to serve as a foundation for training language models focused on answering questions about this franchise.
⚙️ Creation… See the full description on the dataset page: https://huggingface.co/datasets/adriangg04/the-last-of-us-instruction-dataset.last
Gemma 3 1B IT Reasoning Tokenized 8K
Tokenizer/model: google/gemma-3-1b-itMax sequence length: 8192Train on prompt: FalsePad to max length: False
Sources
vanty120/Gpt-5.4-Xhigh-Reasoning-2000x
KingNish/reasoning-base-20k
Efe2898/distill-reasoning-turkish-1k
Efe2898/phi4-grpo-deep-reasoning
Format notes
System messages are intentionally excluded.
GPT-5.4 / Suayp-Talha style datasets use instruction as user, thinking as reasoning, and response as answer.… See the full description on the dataset page: https://huggingface.co/datasets/Efe2898/last.laser_drilling_dataset
Laser Drilling Simulation Reasoning Dataset
Dataset Description
A comprehensive collection of physics-based reasoning data for multi-material laser drilling processes, designed for training AI models with chain-of-thought reasoning capabilities. Covers various material processing scenarios including metals, ceramics, and PCB substrates.
Data Structure
{
"Question": "Process parameter query",
"Complex_CoT": "Step-by-step physical derivation process"… See the full description on the dataset page: https://huggingface.co/datasets/Worlthen/laser_drilling_dataset.multi_llm_dpo
