datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
openvino-arc140v-lunarlake
OpenVINO local-inference on an Intel Arc 140V (Lunar Lake) iGPU
Reference performance data for running local models on a single Intel Core Ultra 7 258V
(Lunar Lake) laptop with the integrated Intel Arc 140V (Xe2) GPU, via OpenVINO. All
inference runs on the iGPU; the NPU stays idle throughout, confirmed by the telemetry here.
This is reference characterization shared by a non-expert contributor — careful measurements
on one machine, offered so others can compare and correct, not… See the full description on the dataset page: https://huggingface.co/datasets/blairducrayoppat/openvino-arc140v-lunarlake.arc-agi-3-schema-traces
ARC-AGI-3 Schema Gameplay Trajectories
This release contains 50 ARC-AGI-3 gameplay trajectories and a dependency-free
scoring utility. The trajectories are split evenly across two collections:
gpt_5_6_sol/: 25 GPT-5.6 Sol trajectories.
claude_fable_opus/: 25 trajectories from Claude Opus 4.8 and Claude Fable 5.
Each trajectory directory includes run.json, a streamed events.jsonl event
log, sanitized session data, snapshots, and the shareable text/image files
produced during… See the full description on the dataset page: https://huggingface.co/datasets/schema-harness/arc-agi-3-schema-traces.us-historical-layoffs-archive-warn-act-notices-removed-from-state-websites
Historical US layoffs archive: 6,799 WARN Act notices that state websites no longer list (2000-2025), recovered
Rebuilt 2026-09-25. Five state labor agencies — Connecticut, Michigan, New York,
North Carolina and Pennsylvania — retired the web pages their older WARN Act
layoff notices lived on. Their current pages start years later. This dataset is
every notice in our file that came from one of those retired pages and is not
on the agency's live page today: 6,799 notices, 6,799… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-historical-layoffs-archive-warn-act-notices-removed-from-state-websites.arc-agi-3-schema-traces
ARC-AGI-3 Schema Gameplay Trajectories
This release contains 50 ARC-AGI-3 gameplay trajectories and a dependency-free
scoring utility. The trajectories are split evenly across two collections:
gpt_5_6_sol/: 25 GPT-5.6 Sol trajectories.
claude_fable_opus/: 25 trajectories from Claude Opus 4.8 and Claude Fable 5.
Each trajectory directory includes run.json, a streamed events.jsonl event
log, sanitized session data, snapshots, and the shareable text/image files
produced during… See the full description on the dataset page: https://huggingface.co/datasets/JBrightmanAI/arc-agi-3-schema-traces.2025_Virtual_Cell_Challenge_Test_Datacbi-archive-raw
Central Bank of Ireland Archive: original source files
6,309 original files, 6.56 GB. Every PDF, spreadsheet, Word document and
archive gathered from the Central Bank of Ireland's public website, stored by
content hash so that a search result can be turned back into the document a
human would actually read.
This is the raw tier. If you want the text, you almost certainly want
aditya487/cbi-archive-corpus
instead: 5,568 documents and 89,242 page or pseudo-page rows as Parquet… See the full description on the dataset page: https://huggingface.co/datasets/aditya487/cbi-archive-raw.industrial-technical-archive
🚀 Latest Updates (Sep, 2026)
Version: v09.2026 (Verified)
Status: Integrated with 1,000,000+ records.
New Files: product-E-20-09-2026.csv & product-V-20-09-2026.csv.
QTE Technologies: Industrial & Scientific Knowledge Base
Wikidata Entity: Q138411149
IPFS CID: bafybeibogxxuhmzfrsuhcfd4qr4tmc4okhmrcwhp3266hq47ccuyjnjxoq
Official Neural Hub: qtetech.github.io
This is the permanent technical archive for QTE Technologies, ensuring long-term accessibility of… See the full description on the dataset page: https://huggingface.co/datasets/QTE-Technologies/industrial-technical-archive.ArcBench
ArcBench: ML Conference Oral Paper-Presentation Benchmark
This benchmark is from the paper Narrative-Driven Paper-to-Slide Generation via ArcDeck.
A curated benchmark dataset of 100 oral presentation paper-slide deck link pairs from top-tier machine learning conferences (CVPR, ICCV, ICLR, ICML, NeurIPS), spanning 2022–2025. Each entry provides rich metadata together with links to the original paper PDF and presentation slides, plus a script that downloads them all in one step.… See the full description on the dataset page: https://huggingface.co/datasets/ArcDeck/ArcBench.arct
The Argument Reasoning Comprehension Task: Identification and Reconstruction of Implicit Warrants
https://github.com/UKPLab/argument-reasoning-comprehension-task
@InProceedings{Habernal.et.al.2018.NAACL.ARCT,
title = {The Argument Reasoning Comprehension Task: Identification
and Reconstruction of Implicit Warrants},
author = {Habernal, Ivan and Wachsmuth, Henning and
Gurevych, Iryna and Stein, Benno},
publisher = {Association for… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/arct.ARCHIVE-TEXT-URLS
Internet Archive English Text URLs Dataset
Dataset Description
This dataset contains 11,151,637 direct download URLs to OCR-processed text files from the Internet Archive's digital library. All entries are English-language texts spanning books, documents, historical records, and various other written materials.
Dataset Summary
Total Rows: 11,151,637
Language: English
Source: Internet Archive
Format: CSV with metadata and direct text file URLs
Text… See the full description on the dataset page: https://huggingface.co/datasets/Navanjana/ARCHIVE-TEXT-URLS.w2t-llm-arc-easy-lora
W2T Llm Arc Easy Lora
This repository contains artifacts for the W2T paper:
Paper: W2T: LoRA Weights Already Know What They Can Do
Repo: Weight2Token
Summary
ARC-Easy LoRA checkpoints and prepared metadata used for performance prediction.
Source Status
Storage location: local
Verification status: confirmed
Files
See manifest.json for the exact local or remote source paths used to prepare this release.
Citation… See the full description on the dataset page: https://huggingface.co/datasets/Xiaolong-Han/w2t-llm-arc-easy-lora.cars_from_drom.ru_archive_2007-2025More information on the parsing process can be found here: https://github.com/zavzyatiy/drom_archive_parser.
This dataset is also published on Kaggle: https://www.kaggle.com/datasets/assaabramovich/resaled-cars-from-drom-ruarchive-2018-2023/.
Main dataset with all data: drom_archive_2007-2025_full.csv
Dataset with (almost) all configurations from Drom for cars in data: additional_data/drom-24-07-2025-all_main_cars_configurations.csv
Dataset with identification of regions for all cities in… See the full description on the dataset page: https://huggingface.co/datasets/zavzyatiy/cars_from_drom.ru_archive_2007-2025.prog-archivessinhala-tts-dataset-archive-20260429-082457
Sinhala TTS Dataset
Clean, segmented Sinhala speech from the "Unlimited History" YouTube series by @sunchare.
Stats
Metric
Value
Utterances
218
Train
208
Val
10
Hours
0.51
Mean duration
8.5s
Sample rate
22050 Hz
Pipeline
Raw YouTube audio -> HTDemucs -> VoiceFixer + DeepFilterNet3 ->
Diarization -> Silero-VAD -> ASR (faster-whisper: C:\Users\kosal\sinhala-tts\whisper-small-si-ct2) -> Quality filtering (SNR>=20.0dB)
Format… See the full description on the dataset page: https://huggingface.co/datasets/outlawmold/sinhala-tts-dataset-archive-20260429-082457.arc_agi_2_human_testing
ARC-AGI-2 Human testing data
This file contains data from human testing sessions on ARC-AGI tasks.
Each row represents a single test attempt by a human participant on a specific task-test pair in the "Public Train" or "Public Eval" ARC-AGI-2 datasets. Not all tasks in the released "Public Train"
sets were tested, so these results are not comprehensive. This data does not include tasks from "Semi Private Evaluation" or "Private Evaluation"
Column Descriptions… See the full description on the dataset page: https://huggingface.co/datasets/arcprize/arc_agi_2_human_testing.arctic-open-water-window
Copyright and license
Copyright (c) Heimdall Research. This package is released under Creative Commons Attribution 4.0 International (CC BY 4.0). The license covers Heimdall Research's contribution. Upstream inputs keep their own terms (see NOTICE.md in the package when present).
Attribution. Contains modified Copernicus Sentinel data 2015-2026. NSIDC sea-ice concentration CDRs: DOI 10.7265/b18j-z797 and DOI 10.7265/tm2n-1m33.
Arctic Open-Water Window
Not for… See the full description on the dataset page: https://huggingface.co/datasets/HeimdallResearch/arctic-open-water-window.jazz-music-archivesarc-agi-3-schema-traces-gpt56
ARC-AGI-3 Schema Gameplay Trajectories — GPT-5.6 Sol
This release contains every gpt-5.6-sol gameplay trajectory produced on our
cluster with the world_model_v5 agent harness — 100 runs across the 25 public
ARC-AGI-3 games — plus a dependency-free scoring utility.
It is the GPT-5.6 Sol member of a family built by the same harness and the same
sanitizer, so trajectories can be compared game by game:
arc-agi-3-schema-traces-fable5 — Claude Fable 5, best per game (25)… See the full description on the dataset page: https://huggingface.co/datasets/guanning/arc-agi-3-schema-traces-gpt56.arched-halls-decision-matrix
Arched Halls Decision Matrix / Macierz decyzyjna hal łukowych
Dataset summary
This Polish-language dataset describes 30 practical application scenarios for an arched hall (hala łukowa) in agriculture, storage, logistics, transport, industry, waste management, infrastructure, sports, public facilities, seasonal buildings, construction and energy.
Each record connects the intended use of an arched hall with qualitative decision factors such as indoor climate… See the full description on the dataset page: https://huggingface.co/datasets/halalukowa24/arched-halls-decision-matrix.ARC-Challenge-Explained-by-ChatGPTThis is a dataset with explanations from ChatGPT for the correct and incorrect answers in ARC Challenge. The explanations are generated by prompting ChatGPT with answer keys and in-context examples. We expect this dataset to be an useful source for understanding the commonsense reasoning ability of LLMs or training other LMs.
ArchEGraph-demo
ArchEGraph-demo
ArchEGraph-demo is a compact demo package of the ArchEGraph building-energy dataset for graph-based and weather-conditioned learning.
Dataset Summary
Total cases in manifest.csv: 300
Unique buildings: 75
Unique weather IDs: 48
n_steps: always 8,760
n_spaces range: 2 to 132
This package currently stores:
manifest.csv (index of all demo cases)
building/ (75 files)
geometry/ (75 files)
weather/ (48 files)
energy/ (300 files)
split/ (demo split CSV files)… See the full description on the dataset page: https://huggingface.co/datasets/ArchEGraph/ArchEGraph-demo.arc-agi-3-schema-traces-opus48
ARC-AGI-3 Schema Gameplay Trajectories — Claude Opus 4.8
This release contains the best claude-opus-4-8 / max trajectory for each of
the 25 public ARC-AGI-3 games, plus a dependency-free scoring utility. It is the
Opus 4.8 counterpart of
arc-agi-3-schema-traces-fable5,
produced by the same agent harness (world_model_v5) and the same sanitizer, so
the two collections can be compared game by game.
Each trajectory directory includes run.json, a streamed events.jsonl event
log… See the full description on the dataset page: https://huggingface.co/datasets/guanning/arc-agi-3-schema-traces-opus48.polaris-arctic-v2
Arctic v2
Arctic is one of six datasets in Polaris: Learning to Generate Table Descriptions from Retrieval
Feedback, alongside aw, lter, ecir, wikitables, and
wtr.
It holds 251 tables sampled at random from the Environmental Data Initiative (EDI), a repository of
long-term ecological research data — lake water temperature, soil chemistry, rainfall, coral
taxonomy — and 20 keyword queries over them. For each query–table pair, a person decided whether
that
table answers that… See the full description on the dataset page: https://huggingface.co/datasets/anhaidgroup/polaris-arctic-v2.polaris-arctic-v1
Arctic v1
Arctic is one of six datasets in Polaris: Learning to Generate Table Descriptions from Retrieval
Feedback, alongside aw, lter, ecir, wikitables, and
wtr.
It holds 251 tables sampled at random from the Environmental Data Initiative (EDI), a repository of
long-term ecological research data — lake water temperature, soil chemistry, rainfall, coral
taxonomy — and 20 keyword queries over them. For each query–table pair, a person decided whether that
table answers that… See the full description on the dataset page: https://huggingface.co/datasets/anhaidgroup/polaris-arctic-v1.ArcMMLU
Introduction
ArcMMLU is a Chinese benchmark specifically designed for evaluating LLMs on Library & Information Science (LIS). It aims to evaluate the knowledge and reasoning capabilities of LLMs in the LIS academic field, which covers four key sub-areas: Archival Science, Data Science, Library Science, and Information Science. Please refer to our paper for more information ArcMMLU: A Library and Information Science Benchmark for Large Language Models
It is important to note that the… See the full description on the dataset page: https://huggingface.co/datasets/patrickshitou/ArcMMLU.llm-cold-start-benchmark
LLM Container Cold-Start Benchmark
Measurements of how long it takes to bring a language model from cold storage to
a state where it can serve its first token, across 25 open-weight
models spanning 17 architecture families and
100.9 GiB of checkpoints, on a single NVIDIA T4.
Cold start is the latency a serverless or scale-to-zero inference platform pays
when it has no warm replica. It decomposes into weight transfer from storage,
deserialization into host memory, transfer to the… See the full description on the dataset page: https://huggingface.co/datasets/ArchCoder/llm-cold-start-benchmark.Sentinel2_arc_of_deforestion_csv_DataMunyarwanda-AI-AutoTrain
Munyarwanda AI - AutoTrain dataset
AutoTrain-ready version of arcange9/Munyarwanda-AI-Dataset v0.2.
Every example is pre-formatted in Qwen chat template as a single text column (5,542 train / 22 validation rows).
Intended recipe (Hugging Face AutoTrain, base model Qwen/Qwen3-0.6B):
LLM task, causal LM
text column: text
LoRA/PEFT + int4 quantization to fit free-tier GPUs
arc-agi-3-schema-traces-gpt56-xhigh
ARC-AGI-3 Schema Gameplay Trajectories — GPT-5.6 Sol (xhigh)
The best gpt-5.6-sol trajectory at xhigh reasoning effort for each of the
25 public ARC-AGI-3 games, produced with the world_model_v5 agent harness.
This release exists to make the cross-model comparison single-effort on all
sides. Its siblings are each one model at one effort, but the
gpt-5.6-sol collection in
arc-agi-3-schema-gameplay
is a mix of xhigh and max (16 games + 9 games), so it is not directly
comparable to… See the full description on the dataset page: https://huggingface.co/datasets/guanning/arc-agi-3-schema-traces-gpt56-xhigh.Sheboygan-People-Archive
Sheboygan People Archive
Public, machine-readable people corpus for SheVegas — original photography of Michael Brunette and people around him / Sheboygan life (portraits, friends, gatherings), separate from the places-focused Sheboygan Visual Archive.
Drive remains the master vault. This Hub dataset is the public intelligence layer: approved media + structured metadata only.
License
CC BY 4.0 — attribution required.
What goes here
Approved PUBLIC… See the full description on the dataset page: https://huggingface.co/datasets/SheVegas/Sheboygan-People-Archive.
