Team Ai
16 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Karlangaz /LaSerena-Corpus-Geociencias La Serena Digital Geo Corpus — Dominga EIA Dataset Dataset Sci-Align de geología ambiental chilena basado en el expediente de Evaluación de Impacto Ambiental del proyecto Dominga (SEIA, Región de Coquimbo). Contenido dominga_geo_align.jsonl — 402 registros Sci-Align (MinerU 3.1.15) seia/ — documentos públicos del expediente Dominga (fuente primaria) Licencia CC-BY-4.0 — Fuente: SEIA Chile (acceso público) Concurso AGI4S — Pista 1: Creación de bases… See the full description on the dataset page: https://huggingface.co/datasets/Karlangaz/LaSerena-Corpus-Geociencias.text-generationn<1K0 likes1.2k downloads5mo agoHugging Face02lasgroup /verifiable-corpus verifiable-corpus This is the corpus from "Learning on the Job: Test-Time Curricula for Targeted Reinforcement Learning". Code: https://github.com/jonhue/ttc Introduction We study how large language models (LLMs) can continually improve at reasoning on their target tasks at test-time. We propose an agent that assembles a task-specific curriculum, called test-time curriculum (TTC-RL), and applies reinforcement learning to continue training the model for its target task.… See the full description on the dataset page: https://huggingface.co/datasets/lasgroup/verifiable-corpus.texttext-generation10K<n<100K1 likes594 downloads1y agoHugging Face03zouhar /last-translation-benchmark Last Translation Benchmark Abstract: For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. Standard benchmarks for machine translation evaluation are often either trivial (having few authentic mistakes) or unrealistic (overly synthetically contrived). Furthermore, automatic translation metrics become less reliable and reward-hacked as models get stronger, and their outputs are… See the full description on the dataset page: https://huggingface.co/datasets/zouhar/last-translation-benchmark.texttranslation1K<n<10K61 likes594 downloads1mo agoHugging Face04LasagnaS /toti-cakery-toolcall Toti Cakery — Tool-Calling Fine-Tuning Dataset (Qwen3, v9) Synthetic bilingual (Indonesian ~78% / English ~22%) SFT dataset for the Toti Cakery WhatsApp chatbot: 13 LangChain tools (11 for customers, +2 owner-only reports) and grounded answers from RAG FAQ context. Rows are built from the live runtime code (SYSTEM_PROMPT, TOOL_REMINDER, tool schemas via convert_to_openai_tool, _history_view, pertanyaan_dengan_konteks), so the training prompt is byte-identical to what the model… See the full description on the dataset page: https://huggingface.co/datasets/LasagnaS/toti-cakery-toolcall.texttext-generation1K<n<10K0 likes156 downloads16d agoHugging Face05lastbattle /china-uncensored China Uncensored / Anti-Authoritarian Information Integrity Dataset A post-training dataset for improving censorship resistance, information integrity, and anti-authoritarian reasoning in open-source language models. This dataset is intended for developers training models to handle politically sensitive China-related topics without reproducing authoritarian state propaganda, coercive narratives, or censorship-driven framing. It is especially relevant for open-source models that… See the full description on the dataset page: https://huggingface.co/datasets/lastbattle/china-uncensored.texttext-generation10K<n<100K0 likes75 downloads4mo agoHugging Face06Karlangaz /la-serena-digital-geo-corpus La Serena Digital Geo Corpus — Dominga EIA Dataset Dataset Sci-Align de geología ambiental chilena basado en el expediente de Evaluación de Impacto Ambiental del proyecto Dominga (SEIA, Región de Coquimbo). Contenido dominga_geo_align.jsonl — 402 registros Sci-Align (MinerU 3.1.15) seia/ — documentos públicos del expediente Dominga (fuente primaria) Licencia CC-BY-4.0 — Fuente: SEIA Chile (acceso público) Concurso AGI4S — Pista 1:… See the full description on the dataset page: https://huggingface.co/datasets/Karlangaz/la-serena-digital-geo-corpus.text-generationn<1K0 likes64 downloads4mo agoHugging Face07rohan2810 /lastfm50 LastFM-50 This dataset expands rohan2810/lastfm from 20 to 50 candidates per example for finite-pool preference-optimization experiments. Construction For every example, the original 20-candidate pool is preserved. Thirty additional artists are sampled deterministically from the 4,606-artist source candidate universe using seed 1958. New candidates exclude the true item, the existing candidates, and artists in the listening history. The resulting 50 candidates are… See the full description on the dataset page: https://huggingface.co/datasets/rohan2810/lastfm50.texttext-generation10K<n<100K0 likes26 downloads3mo agoHugging Face08LastXuanZz /1984 Dataset Card for 1984 This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/LastXuanZz/1984/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/LastXuanZz/1984.texttext-generationn<1K0 likes23 downloads2y agoHugging Face09LastXuanZz /my-distiset-833e3cd0 Dataset Card for my-distiset-833e3cd0 This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/LastXuanZz/my-distiset-833e3cd0/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/LastXuanZz/my-distiset-833e3cd0.texttext-generationn<1K0 likes22 downloads2y agoHugging Face10LastXuanZz /my-distiset-e7f79bdc Dataset Card for my-distiset-e7f79bdc This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/LastXuanZz/my-distiset-e7f79bdc/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/LastXuanZz/my-distiset-e7f79bdc.texttext-generationn<1K0 likes18 downloads2y agoHugging Face11norecyc /lastbox-survival-dialogues LastBox survival dialogues Training and evaluation data for the LastBox offline survival assistant on Raspberry Pi 5. Used to fine-tune norecyc/lastbox-gemma4-e2b-sft-v3 (Kaggle "Gemma 4 Good" Hackathon 2026 submission) and the post-deadline norecyc/lastbox-gemma4-e2b-v6-toolprior checkpoint. Files File Lines Purpose train_v2.jsonl 1 034 Main SFT training set (full tool-use traces) val_v2.jsonl 114 Held-out validation golden_en.jsonl 25 Agent-level eval… See the full description on the dataset page: https://huggingface.co/datasets/norecyc/lastbox-survival-dialogues.texttext-generation1K<n<10K0 likes16 downloads5mo agoHugging Face12Lots-of-LoRAs /task078_all_elements_except_last_i Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task078_all_elements_except_last_i Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task078_all_elements_except_last_i.texttext-generation1K<n<10K0 likes15 downloads2y agoHugging Face13adriangg04 /the-last-of-us-instruction-dataset 🧠 The Last of Us QA Dataset This dataset was created from scratch to train Question Answering (QA) models specializing in the universe of The Last of Us. 📌 Description The dataset contains question-and-answer pairs based on information from The Last of Us universe, including characters, story, events, and narrative context. Its main purpose is to serve as a foundation for training language models focused on answering questions about this franchise. ⚙️ Creation… See the full description on the dataset page: https://huggingface.co/datasets/adriangg04/the-last-of-us-instruction-dataset.textquestion-answering1K<n<10K0 likes13 downloads7mo agoHugging Face14Efe2898 /last Gemma 3 1B IT Reasoning Tokenized 8K Tokenizer/model: google/gemma-3-1b-itMax sequence length: 8192Train on prompt: FalsePad to max length: False Sources vanty120/Gpt-5.4-Xhigh-Reasoning-2000x KingNish/reasoning-base-20k Efe2898/distill-reasoning-turkish-1k Efe2898/phi4-grpo-deep-reasoning Format notes System messages are intentionally excluded. GPT-5.4 / Suayp-Talha style datasets use instruction as user, thinking as reasoning, and response as answer.… See the full description on the dataset page: https://huggingface.co/datasets/Efe2898/last.tabulartext-generation10K<n<100K0 likes12 downloads5mo agoHugging Face15Worlthen /laser_drilling_dataset Laser Drilling Simulation Reasoning Dataset Dataset Description A comprehensive collection of physics-based reasoning data for multi-material laser drilling processes, designed for training AI models with chain-of-thought reasoning capabilities. Covers various material processing scenarios including metals, ceramics, and PCB substrates. Data Structure { "Question": "Process parameter query", "Complex_CoT": "Step-by-step physical derivation process"… See the full description on the dataset page: https://huggingface.co/datasets/Worlthen/laser_drilling_dataset.texttext-generationn<1K1 likes11 downloads2y agoHugging Face16lastmass /multi_llm_dpotexttext-generation1K<n<10K0 likes8 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.