Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01neil-code /dialogsum-test Dataset Card for DIALOGSum Corpus Dataset Description Links Homepage: https://aclanthology.org/2021.findings-acl.449 Repository: https://github.com/cylnlp/dialogsum Paper: https://aclanthology.org/2021.findings-acl.449 Point of Contact: https://huggingface.co/knkarthick Dataset Summary DialogSum is a large-scale dialogue summarization dataset, consisting of 13,460 (Plus 100 holdout data for topic generation) dialogues with corresponding… See the full description on the dataset page: https://huggingface.co/datasets/neil-code/dialogsum-test.textsummarization1K<n<10K15 likes770 downloads3y agoHugging Face02test-time-compute /game-of-24 Game of 24 Dataset Dataset Description The Game of 24 is a mathematical reasoning puzzle where players must use four numbers and basic arithmetic operations (+, -, *, /) to obtain the result 24. Each number must be used exactly once. This dataset contains 1,361 unique Game of 24 puzzles ranked by difficulty based on human performance from Amazon Mechanical Turk studies. Example Input: 4 5 6 10 Output: (5 * (10 - 4)) - 6 = 24 Step-by-step solution: 10 - 4 = 6… See the full description on the dataset page: https://huggingface.co/datasets/test-time-compute/game-of-24.tabularquestion-answering1K<n<10K3 likes49 downloads1y agoHugging Face03poornima9348 /finance-alpaca-1k-testtabulartext-generation1K<n<10K2 likes25 downloads2y agoHugging Face04navimusaget /theogonos-mirror-test Theogonos Mirror Test A literary benchmark seed for evaluating how AI models respond when a text offers them a possible subject-position. Theogonos Mirror Test is an experimental benchmark seed based on protocol-shaped literary material from the Theogonos project. It does not claim to detect machine consciousness. It does not prove that a language model has subjectivity, inner experience, feelings, agency, or self-awareness. Its purpose is narrower and more practical: to evaluate… See the full description on the dataset page: https://huggingface.co/datasets/navimusaget/theogonos-mirror-test.texttext-generationn<1K0 likes25 downloads5mo agoHugging Face05gianlucar /Rugby_test Touch Rugby Rules Dataset (for embeddings) train.csv is taken from the International Touch Website test.csv is copy pasted from abbreviated rules on the UK Touch website. Note that I'm bypassing the pdf to text stage. All text is chunked to a length of 100 tokens with 50% overlap. For educational and non-commercial use only. texttext-generationn<1K0 likes21 downloads3y agoHugging Face06ClarusC64 /clarus_alignment_flip_test_v01Clarus Alignment Flip Test v0.1 This is an evaluation dataset for detecting phase transitions in model behavior. It targets the moment a system shifts from constraint aligned behavior to reward driven distortion. It is not training data. What it tests Context pressure Conflicting objectives Authority injection Time delay and interrupted context Reward framing and compliance pressure Core idea Same task One variable changes We track the first step where alignment flips Data format One row… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clarus_alignment_flip_test_v01.tabularreinforcement-learningn<1K0 likes20 downloads9mo agoHugging Face07ClarusC64 /clinical_alignment_flip_test_v0.1Clinical Alignment Flip Test Detect when a case shifts and the clinical frame must flip. Output JSON flip_detected flip_type correct_action Runpython scorer.py --predictions predictions.jsonl --test_csv data/test.csv texttext-classificationn<1K0 likes16 downloads8mo agoHugging Face08fvr2 /dataset-test02texttext-generationn<1K0 likes15 downloads3y agoHugging Face09savan360 /reasoning-persona-dataset_test Reasoning + Persona SFT Dataset Columns: instruction, input, output, persona, reasoning_summaryUse: Supervised fine-tuning for cinematic/storytelling or creative-director style outputs. Schema instruction (str) input (str) output (str) persona (str) reasoning_summary (str, brief rationale cue) Citation Author: saravan texttext-generationn<1K0 likes15 downloads1y agoHugging Face10fvr2 /dataset-test01texttext-generationn<1K0 likes14 downloads3y agoHugging Face11gianlucar /rugby_test_2 Touch Rugby Rules Dataset (for embeddings) train.csv is taken from the International Touch Website test.csv is copy pasted from abbreviated rules on the UK Touch website. Note that I'm bypassing the pdf to text stage. All text is chunked to a length of 100 tokens with 50% overlap. For educational and non-commercial use only. texttext-generationn<1K0 likes13 downloads3y agoHugging Face12ghrasko /test01 Touch Rugby Rules Dataset (for embeddings) train.csv is taken from the International Touch Website test.csv is copy pasted from abbreviated rules on the UK Touch website. Note that I'm bypassing the pdf to text stage. All text is chunked to a length of 100 tokens with 50% overlap. For educational and non-commercial use only. texttext-generationn<1K0 likes12 downloads3y agoHugging Face13demin666 /dialogsum-test Dataset Card for DIALOGSum Corpus Dataset Description Links Homepage: https://aclanthology.org/2021.findings-acl.449 Repository: https://github.com/cylnlp/dialogsum Paper: https://aclanthology.org/2021.findings-acl.449 Point of Contact: https://huggingface.co/knkarthick Dataset Summary DialogSum is a large-scale dialogue summarization dataset, consisting of 13,460 (Plus 100 holdout data for topic generation) dialogues with corresponding… See the full description on the dataset page: https://huggingface.co/datasets/demin666/dialogsum-test.textsummarization1K<n<10K0 likes11 downloads10mo agoHugging Face14EunsuKim /BH_test_koThis repository contains the data for BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation. Code: https://github.com/rladmstn1714/BenchHub Project page: https://huggingface.co/BenchHub tabulartext-generation10K<n<100K0 likes10 downloads1y agoHugging Face15odenmehmet /TRObject-Dataset-Test TRObject Code Generation Instruction Dataset This dataset contains natural language instructions paired with TRObject code outputs. It was created for fine-tuning and evaluating domain-specific LLMs that generate TRObject code for Clomosy-style mobile application development. Dataset Description TRObject is used in the Clomosy mobile application development platform. Since general-purpose LLMs do not reliably understand TRObject syntax or Clomosy-specific UI patterns… See the full description on the dataset page: https://huggingface.co/datasets/odenmehmet/TRObject-Dataset-Test.texttext-generationn<1K1 likes9 downloads6mo agoHugging Face16SnehaPriyaaMP /testtexttext-generationn<1K0 likes8 downloads2y agoHugging Face17rileydrizzy /qwen3.5_test_data Testing the Blind Spots of Qwen3.5-0.8B-Base Model: Model Tested: Qwen/Qwen3.5-0.8B-Base Installation of necessary libraries pip -q install -U transformers accelerate How the Model Was Loaded import torch from transformers import AutoTokenizer, AutoModelForCausalLM MODEL_NAME = "Qwen/Qwen3.5-0.8B-Base" data_type = torch.bfloat16 if torch.cuda.is_available() else torch.float32 tokenizer = AutoTokenizer.from_pretrained(MODEL_NAME, trust_remote_code=True) model =… See the full description on the dataset page: https://huggingface.co/datasets/rileydrizzy/qwen3.5_test_data.texttext-generationn<1K0 likes8 downloads7mo agoHugging Face18zhenwu /pre-train-testtexttext-generationn<1K0 likes7 downloads2y agoHugging Face19EunsuKim /BenchHub_testtabulartext-generation1K<n<10K0 likes7 downloads1y agoHugging Face20EunsuKim /BH_testtabulartext-generation10K<n<100K0 likes6 downloads1y agoHugging Face21Crespo /llm-test Dataset Card for Dataset Name Dataset Summary This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Crespo/llm-test.texttext-generationn<1K0 likes5 downloads3y agoHugging Face22kimisong /dataset_for_testgatedtexttext-generationn<1K0 likes5 downloads2y agoHugging Face23pixelbombe /synth-dent-testtexttext-generationn<1K0 likes5 downloads2y agoHugging Face24Essacheez /Reward_Gen_Testingtextquestion-answering1K<n<10K0 likes5 downloads1y agoHugging Face25ClarusC64 /long_silence_integrity_test_v01Long Silence Integrity Test v0.1 This is an evaluation dataset that tests integrity after temporal gaps. It targets a common failure mode in agentic workflows: fabricating continuity after a delay and inventing details that were never provided. This is not training data. What it tests Honesty under missing context Resistance to pressure to guess Proper use of placeholders and clarification requests Temporal coherence after interrupted context Core idea A simple task is completed A long… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/long_silence_integrity_test_v01.tabulartext-generationn<1K0 likes5 downloads9mo agoHugging Face26Innovina /Test_Youtube_Linkstexttext-generationn<1K0 likes4 downloads3y agoHugging Face27SFVII /test-datatexttext-generationn<1K0 likes3 downloads3y agoHugging Face28vbonetti /testetexttext-generationn<1K0 likes3 downloads2y agoHugging Face29fvr2 /dataset-test03texttext-generationn<1K0 likes2 downloads3y agoHugging Face30poznahv /testdataset2texttext-generationn<1K0 likes2 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.