Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01open-law-data-thailand /soc-ratchakitcha Royal Gazette Thailand (Ratchakitcha) Dataset ชุดข้อมูลราชกิจจานุเบกษา (แบบ Machine Readable) โครงการ Open Law Data Thailand ร่วมกับคณะกรรมาธิการการพาณิชย์และการอุตสาหกรรม วุฒิสภา ได้รับความอนุเคราะห์ข้อมูลจาก สำนักเลขาธิการคณะรัฐมนตรี (สลค.) เพื่อเผยแพร่ข้อมูลกฎหมายไทยสู่สาธารณะในรูปแบบที่ประมวลผลได้ด้วยคอมพิวเตอร์ (Machine Readable) เพื่อส่งเสริมนวัตกรรม Legal Tech และ AI ของประเทศไทย Dataset Description ชุดข้อมูลนี้รวบรวมรายการประกาศในราชกิจจานุเบกษา… See the full description on the dataset page: https://huggingface.co/datasets/open-law-data-thailand/soc-ratchakitcha.tabulartext-retrieval1M<n<10M14 likes27k downloads7h agoHugging Face02gililior /mmlu-prox-eval-predictions MMLU-ProX Multilingual Model Predictions Raw per-sample model predictions on MMLU-ProX across 29 languages and 25 open-weight LLMs, produced with lm-evaluation-harness. This dataset releases the full prediction logs (not just aggregate scores) so that item-level responses can be re-analysed — e.g. for Item Response Theory (IRT) modelling of multilingual benchmarks, error analysis, or per-item difficulty estimation. Repository structure mmlu_prox_<lang>/ └──… See the full description on the dataset page: https://huggingface.co/datasets/gililior/mmlu-prox-eval-predictions.tabularquestion-answering1M<n<10M0 likes25k downloads4mo agoHugging Face03ShaofantuoshuzhengzhiSha /GUIGuard-Bench GUIGuard-Bench (Public Ladder) GUIGuard-Bench is a cross-platform GUI agent benchmark for studying privacy risks and privacy-preserving execution in multimodal GUI agents. This public-ladder release contains 121 GUI interaction trajectories (68 Android + 53 PC) for benchmark evaluation, with 26,407 region-level privacy annotations across 2,002 screenshots. For the anonymous review version of the evaluation toolkit, see GUIGaurd-Bench-CA4F. Dataset Summary GUI agents… See the full description on the dataset page: https://huggingface.co/datasets/ShaofantuoshuzhengzhiSha/GUIGuard-Bench.imagequestion-answering1K<n<10K1 likes9.2k downloads5mo agoHugging Face04OpenGVLab /ShareGPT-4ogatedtabularvisual-question-answering10K<n<100K199 likes8.5k downloads2y agoHugging Face05ulamai /UnsolvedMath🌐 Browse UnsolvedMath online ✅ Paper: Open Mathematical Problems as an AI Reasoning Benchmark UnsolvedMath Dataset A comprehensive curated collection of 15,458 open, partially solved, and solved mathematics problems across all domains and difficulty levels, including the largest collection of Erdős problems available in machine-readable format. Available for browsing at unsolvedmath.com. Paper: "Open Mathematical Problems as an AI Reasoning Benchmark" Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ulamai/UnsolvedMath.documentquestion-answering10K<n<100K80 likes6.9k downloads8d agoHugging Face06johanneskirmayr /car-bench-dataset CAR-Bench Dataset CAR-Bench is a benchmark for evaluating AI voice assistants in a realistic automotive (car) environment. It tests an agent's ability to correctly use vehicle control tools, handle disambiguation, and avoid hallucinations. Dataset Structure The dataset is organized into task configs and mock data configs: Tasks Each task defines a user persona, an instruction, the initial vehicle/environment context, and the ground-truth sequence of tool-call… See the full description on the dataset page: https://huggingface.co/datasets/johanneskirmayr/car-bench-dataset.tabulartext-generation1M<n<10M4 likes5k downloads8mo agoHugging Face07dell-research-harvard /newswire Dataset Card for NewsWire Dataset Summary NewsWire contains 2.7 million unique public domain U.S. news wire articles, written between 1878 and 1977. Locations in these articles are georeferenced, topics are tagged using customized neural topic classification, named entities are recognized, and individuals are disambiguated to Wikipedia using a novel entity disambiguation model. Languages English (en) Dataset Structure Each year in the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/dell-research-harvard/newswire.tabulartext-classification1M<n<10M93 likes4.4k downloads1y agoHugging Face08Qalam /nuclear-intelligence-dataset Nuclear Intelligence Dataset Public, auto-generated dataset of validated nuclear-energy research cycles. Latest stats (auto-updated): 🪙 NES tokens minted: 0 ⛓️ Blockchain length: 1 blocks 🕸️ Knowledge entities: 2 Source GitHub: https://github.com/QalamHipHop/nuclear-intelligence HF Space: https://huggingface.co/spaces/Qalam/Nuclear-Intelligence License MIT tabularquestion-answeringn<1K1 likes4.2k downloads49m agoHugging Face09hashmortar /spreadsheet-bench-v2-modified SpreadsheetBench V2 Modified: Multi-Document QA 1,060 questions and reference answers grounded in 127 Excel workbooks, 35 PDFs and 9 DOCX files. This independent derivative of SpreadsheetBench 2 shifts the task from editing spreadsheets and producing workbook deliverables toward finding, interpreting and combining information in business documents. An independent project built entirely from publicly available source material and newly authored QA annotations. No private company… See the full description on the dataset page: https://huggingface.co/datasets/hashmortar/spreadsheet-bench-v2-modified.documentquestion-answering1K<n<10K2 likes3.7k downloads29d agoHugging Face10Joysw909 /AVQA Summary | 摘要 This dataset is collected from the AVQA training subset (train_qa.json). We converted the data to the R1-AQA format, where each line in the text file represents a JSON object with specific keys. The AVQA training set originally consists of approximately 40k samples. However, we use only about 38k samples because some data sources have become invalid (e.g. link failure, or less than 10 seconds). Given that there is no quick link to the audio mentioned in the above two… See the full description on the dataset page: https://huggingface.co/datasets/Joysw909/AVQA.audioquestion-answering10K<n<100K2 likes3.5k downloads11mo agoHugging Face11buaaplay /SVCBench SVCBench: Streaming Video Counting Benchmark This dataset contains the clipped video segments for SVCBench, a Streaming Video Counting Benchmark for Spatial-Temporal State Maintenance. It repositions counting as a minimal, controlled probe for diagnosing how video understanding models maintain world state along the video timeline. Project Page: https://buaa-colalab.github.io/SVCBench/ Code: https://github.com/buaa-colalab/SVCBench Dataset Description This… See the full description on the dataset page: https://huggingface.co/datasets/buaaplay/SVCBench.tabularvideo-classification1K<n<10K10 likes3.4k downloads3mo agoHugging Face12G4KMU /t2-ragbench Dataset Card for T2-RAGBench Project Page | Paper | Code IMPORTANT NOTICE: We deleted VQAonBD from the dataset due to low quality of the question reformulations. If you still want to use it you will find the data in the previous commit history. Dataset Description Dataset Summary T2-RAGBench is a benchmark dataset designed to evaluate Retrieval-Augmented Generation (RAG) on financial documents containing both text and tables. It consists of 23,088… See the full description on the dataset page: https://huggingface.co/datasets/G4KMU/t2-ragbench.documenttable-question-answering10K<n<100K17 likes3.3k downloads6mo agoHugging Face13stanfordnlp /SHP 🚢 Stanford Human Preferences Dataset (SHP) If you mention this dataset in a paper, please cite the paper: Understanding Dataset Difficulty with V-Usable Information (ICML 2022). Summary SHP is a dataset of 385K collective human preferences over responses to questions/instructions in 18 different subject areas, from cooking to legal advice. The preferences are meant to reflect the helpfulness of one response over another, and are intended to be used for training RLHF… See the full description on the dataset page: https://huggingface.co/datasets/stanfordnlp/SHP.tabulartext-generation100K<n<1M325 likes3.1k downloads3y agoHugging Face14minkyuchoi /Temporal-Logic-Video-Dataset Temporal Logic Video (TLV) Dataset Temporal Logic Video (TLV) Dataset Synthetic and real video dataset with temporal logic annotation Explore the GitHub » NSVS-TL Project Webpage · NSVS-TL Source Code Overview The Temporal Logic Video (TLV) Dataset addresses the scarcity of state-of-the-art video datasets for long-horizon, temporally extended activity and object detection. It comprises two main components: Synthetic… See the full description on the dataset page: https://huggingface.co/datasets/minkyuchoi/Temporal-Logic-Video-Dataset.tabularquestion-answeringn<1K1 likes2.7k downloads2y agoHugging Face15mib-bench /copycolors_mcqaThis dataset consists of formatted n-way multiple choice questions, where n is in [2,10]. The task itself is simply to copy the prototypical color from the context and produce the corresponding color's answer choice letter. The "prototypical colors" dataset instances themselves come from Memory Colors (Norland et al. 2021) and corypaik/coda (instances whose object_group is 0, indicating participants agreed on a prototypical color of that object). tabularquestion-answering1K<n<10K0 likes2.3k downloads2y agoHugging Face164papersubmission /TPBench TPBench: A Turning-Point Benchmark for Dialogue Compression TPBench measures three retrieval targets after dialogue compression: P1: the user's initial goal; P2: the final value of a revised slot; P3: both answers when the final update occurs late. This repository includes data, construction inputs, compression and reader code, scorers, saved answers, and paper results. RESULTS_MAP.md connects the paper comparisons to their evidence files. LLMLingua adapter… See the full description on the dataset page: https://huggingface.co/datasets/4papersubmission/TPBench.tabularquestion-answering10K<n<100K2 likes2.2k downloads7d agoHugging Face17AdaptLLM /finance-tasks Adapting LLMs to Domains via Continual Pre-Training (ICLR 2024) This repo contains the evaluation datasets for our paper Adapting Large Language Models via Reading Comprehension. We explore continued pre-training on domain-specific corpora for large language models. While this approach enriches LLMs with domain knowledge, it significantly hurts their prompting ability for question answering. Inspired by human learning via reading comprehension, we propose a simple method to… See the full description on the dataset page: https://huggingface.co/datasets/AdaptLLM/finance-tasks.tabulartext-classification10K<n<100K83 likes2.1k downloads2y agoHugging Face18Coldog2333 /JMedBench Maintainers Junfeng Jiang@Aizawa Lab: jiangjf (at) is.s.u-tokyo.ac.jp Jiahao Huang@Aizawa Lab: jiahao-huang (at) g.ecc.u-tokyo.ac.jp If you find any error in this benchmark or want to contribute to this benchmark, please feel free to contact us. Introduction This is a dataset collection of JMedBench, which is a benchmark for evaluating Japanese biomedical large language models (LLMs). Details can be found in this paper. We also provide an evaluation framework, med-eval… See the full description on the dataset page: https://huggingface.co/datasets/Coldog2333/JMedBench.tabulartext-classification100K<n<1M8 likes2k downloads2y agoHugging Face19typhoon-ai /thai_exam Dataset Card for Thai_Exam ThaiExam is a Thai knowledge benchmarking dataset, consisting of multiple-choice questions from examinations in Thailand. The dataset was originally developed for evaluating Typhoon (Thai LLM). This dataset contains 5 splits corresponding to 5 examinations as follows: ONET: The Ordinary National Educational Test (ONET) is an examination for students in Thailand. This dataset is based on the grade-12 ONET exam, comprising 4 subjects and each question has 5… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/thai_exam.tabularquestion-answeringn<1K19 likes1.9k downloads2y agoHugging Face20choucsan /mimo-claude-code-traces-1k MIMO Claude Code Traces MIMO Claude Code Traces is a collection of coding-agent trajectories in a Claude Code-style environment. Each record contains a user coding task, the full multi-turn message trace, available tool schemas, assistant reasoning fields, tool calls, tool outputs, and metadata such as model name, category, duration, cost, token usage, and whether the trace used tools. The traces were generated with mimo-v2.5-pro, MiMo's most capable model at the time of… See the full description on the dataset page: https://huggingface.co/datasets/choucsan/mimo-claude-code-traces-1k.tabulartext-generation1K<n<10K11 likes1.9k downloads2mo agoHugging Face21cambridgeltl /vsr_random VSR: Visual Spatial Reasoning This is the random set of VSR: Visual Spatial Reasoning (TACL 2023) [paper]. Usage from datasets import load_dataset data_files = {"train": "train.jsonl", "dev": "dev.jsonl", "test": "test.jsonl"} dataset = load_dataset("cambridgeltl/vsr_random", data_files=data_files) Note that the image files still need to be downloaded separately. See data/ for details. Go to our github repo for more introductions. Citation If you find VSR… See the full description on the dataset page: https://huggingface.co/datasets/cambridgeltl/vsr_random.imagetext-classification10K<n<100K4 likes1.7k downloads4y agoHugging Face22cambridgeltl /vsr_zeroshot VSR: Visual Spatial Reasoning This is the zero-shot set of VSR: Visual Spatial Reasoning (TACL 2023) [paper]. Usage from datasets import load_dataset data_files = {"train": "train.jsonl", "dev": "dev.jsonl", "test": "test.jsonl"} dataset = load_dataset("cambridgeltl/vsr_zeroshot", data_files=data_files) Note that the image files still need to be downloaded separately. See data/ for details. Go to our github repo for more introductions. Citation If you find… See the full description on the dataset page: https://huggingface.co/datasets/cambridgeltl/vsr_zeroshot.imagetext-classification1K<n<10K1 likes1.7k downloads4y agoHugging Face23junfeng0288 /MathReal Dataset Card for MathReal Dataset Description Paper Information Dataset Examples  Leaderboard Citation Dataset Description The MathReal dataset is designed to evaluate the performance of Multi-modal Large Language Models (MLLMs)on real-world K-12 mathematical questions. It consists of 2,000 high-quality math problems, each represented as an image captured in authentic educational contexts. The dataset includes various types of questions, such as multiple-choice… See the full description on the dataset page: https://huggingface.co/datasets/junfeng0288/MathReal.imagemultiple-choicen<1K2 likes1.5k downloads1y agoHugging Face24csoai /gspc-swarm GSPC — swarm bank (SwarmBench v2b) In one line: The frozen multi-agent coordination safety bank (SwarmBench v2b) behind the board's swarm axis, with candidates, samples and protocol. For evaluators who want to run the same items against their own model. Use it from datasets import load_dataset ds = load_dataset("csoai/gspc-swarm", split="train") print(ds[0]) Verify a signed card in your browser, free, no account:… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-swarm.tabularquestion-answeringn<1K0 likes1.5k downloads1h agoHugging Face25HPAI-BSC /CareQA CareQA Dataset Summary CareQA is a healthcare QA dataset with two versions: Closed-Ended Version: A multichoice question answering (MCQA) dataset containing 5,621 QA pairs across six categories. Available in English and Spanish. Open-Ended Version: A free-response dataset derived from the closed version, containing 2,769 QA pairs (English only). The dataset originates from… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/CareQA.tabularquestion-answering10K<n<100K18 likes1.5k downloads1y agoHugging Face26HPAI-BSC /medical-specialities Medical Question Classification Dataset Dataset Summary This dataset is designed for medical language models evaluation. It merges several of the most important medical QA datasets into a common format and classifies them into 35 distinct medical categories. This structure enables users to identify any specific categories where the model's performance may be lacking and address these areas accordingly. Dataset Structure Data Fields id: Unique… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/medical-specialities.tabularquestion-answering10K<n<100K7 likes1.5k downloads11mo agoHugging Face27stanfordnlp /SHP-2 🚢 Stanford Human Preferences Dataset v2 (SHP-2) Summary SHP-2 is a dataset of 4.8M collective human preferences over responses to questions/instructions in 129 different subject areas, from cooking to legal advice. It is an extended version of the original 385K SHP dataset. The preferences are meant to reflect the helpfulness of one response over another, and are intended to be used for training RLHF reward models and NLG evaluation models (e.g., SteamSHP). Each example… See the full description on the dataset page: https://huggingface.co/datasets/stanfordnlp/SHP-2.tabulartext-generation1M<n<10M18 likes1.4k downloads3y agoHugging Face28michalsr /molmo2-moments Molmo-2 Moments (M2M) Long-video QA dataset where every question is anchored to a specific [start, end] clip interval in seconds. Released alongside the ToolMerge paper, "Decomposing Queries into Tool Calls for Long-Video Keyframe Retrieval". ⚠️ Source videos & ownership The videos/*.mp4 files in this repository were collected from YouTube. We do not own these videos and claim no copyright over them. All rights to the video content remain with the original… See the full description on the dataset page: https://huggingface.co/datasets/michalsr/molmo2-moments.tabularvideo-text-to-text10K<n<100K0 likes1.3k downloads2mo agoHugging Face29cardiffnlp /super_tweeteval SuperTweetEval Dataset Card for "super_tweeteval" Dataset Summary This is the oficial repository for SuperTweetEval, a unified benchmark of 12 heterogeneous NLP tasks. More details on the task and an evaluation of language models can be found on the reference paper, published in EMNLP 2023 (Findings). Data Splits All tasks provide custom training, validation and test splits. task dataset load dataset description number of instances Topic… See the full description on the dataset page: https://huggingface.co/datasets/cardiffnlp/super_tweeteval.tabulartext-classification100K<n<1M15 likes1.2k downloads2y agoHugging Face30csoai /gspc-gov GSPC — governance bank (GovBench) In one line: The frozen EU AI Act risk-tier classification bank (GovBench) behind the board's governance axis. For evaluators who want to test a model on the same items. Use it from datasets import load_dataset ds = load_dataset("csoai/gspc-gov", split="train") print(ds[0]) Verify a signed card in your browser, free, no account: https://councilof.ai/gspc-verify/?ref=hf-gspc-gov For agents: MCP endpoint POST… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-gov.tabularquestion-answeringn<1K0 likes1.1k downloads1h agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.