datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arctic
Arctic Shift Reddit Archive
Every Reddit comment and submission since 2005, organized as monthly Parquet shards
What is it?
The full Reddit archive from Arctic Shift, converted to Parquet and hosted here for easy access. Covers every public subreddit from 2005-12 through 2026-02.
Right now the archive has 15.7B items (12.9B comments, 2.8B submissions) in 1.3 TB of compressed Parquet. Comments and submissions are stored as separate datasets, split into monthly… See the full description on the dataset page: https://huggingface.co/datasets/open-index/arctic.opengenome2
OpenGenome2
OpenGenome2 is a database of nearly 9 trillion base pairs of curated DNA from across all domains of life. Collected from diverse species and public data sources, OpenGenome2 was used to train Evo 2 models. Please refer to the Evo 2 preprint or github repository for further details and usage examples.
We provide OpenGenome2 in two formats, the dataset is organized into two main directories to reflect this:
fasta which contain the DNA sequences
jsonl which… See the full description on the dataset page: https://huggingface.co/datasets/arcinstitute/opengenome2.SCPWiki-Cleaned-PDF-Archivesgithub_archive
GitHub Archive
Description
According to GitHub’s terms of service, issues and pull request descriptions—along with the their comments—inherit the license of their associated repository.
To collect this data, we used the GitHub Archive’s public BigQuery table of events to extracted all issue, pull request, and comment events since 2011 and aggregated them into threads.
The table appeared to be missing “edit” events so the text from each comment is the original from when… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/github_archive.openvino-arc140v-lunarlake
OpenVINO local-inference on an Intel Arc 140V (Lunar Lake) iGPU
Reference performance data for running local models on a single Intel Core Ultra 7 258V
(Lunar Lake) laptop with the integrated Intel Arc 140V (Xe2) GPU, via OpenVINO. All
inference runs on the iGPU; the NPU stays idle throughout, confirmed by the telemetry here.
This is reference characterization shared by a non-expert contributor — careful measurements
on one machine, offered so others can compare and correct, not… See the full description on the dataset page: https://huggingface.co/datasets/blairducrayoppat/openvino-arc140v-lunarlake.Nemotron-SFT-ARC-AGI-v1
Dataset Description:
Nemotron-SFT-ARC-AGI-v1 is a supervised fine-tuning (SFT) dataset of multi-turn agentic reasoning traces produced by open-weight large language models attempting to solve ARC-AGI visual-reasoning puzzles. Each ARC puzzle (a set of (input grid, output grid) demonstration pairs plus one or more test inputs, where grids are 2D integer arrays representing colors) is formatted as a text prompt and given to an agent powered by one of nine open-weight reasoning… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SFT-ARC-AGI-v1.rtl-augmented-v3
RTL Bug Fix — Augmented Dataset
Auto-generated dashboard snapshot (2026-04-14T10:53:43).
Overview
Metric
Value
Total problems
718
Repos with data
57 / 81
Modules augmented
408
Bug types
11/11
Augmentation success
48.9%
Coverage
Distribution
Augmentation Health
Topic Coverage
Warnings
lucky-wfw_IC_System_Design: 0 problems from 48 attempts — likely systematic sim issue
meiniKi_FazyRV:… See the full description on the dataset page: https://huggingface.co/datasets/architect-ubc-capstone/rtl-augmented-v3.google-code-archive
Google Code Archive Dataset
Dataset Description
This dataset was compiled from the Google Code Archive, a preserved snapshot of projects hosted on Google Code, Google's open-source project hosting service that operated from 2006 to 2016. Google Code was one of the major code hosting platforms of its era, hosting hundreds of thousands of open-source projects before its shutdown. The archive provides a unique historical record of open-source development during a formative… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/google-code-archive.cyberdata-full
⚠️
THERE IS A NEWER VERSION
This version performs poorly. Use it only for comparison and testing — not as a main dataset and not for enterprise use.
View Vertex AGI for the latest version →
CyberData Full
479,214 SFT examples: every example from all five source datasets (verified agentic-coding trajectories, agent behavior, and structured vulnerability intelligence), built entirely from non-gated, redistributable sources.
The CyberData family… See the full description on the dataset page: https://huggingface.co/datasets/VertexAGI-Archives/cyberdata-full.arctic
Arctic Shift Reddit Archive
Every Reddit comment and submission since 2005, organized as monthly Parquet shards
What is it?
The full Reddit archive from Arctic Shift, converted to Parquet and hosted here for easy access. Covers every public subreddit from 2005-12 through 2026-02.
Right now the archive has 1.6B items (362.1M comments, 1.2B submissions) in 181.4 GB of compressed Parquet. Comments and submissions are stored as separate datasets, split into monthly shards… See the full description on the dataset page: https://huggingface.co/datasets/Dk587/arctic.350B-Pipeline-Dataset
OLMo 3 pre-tokenized and processed training data
This repository redistributes Ai2/AllenAI's official OLMo 3 training data in an OLMo-core-ready archive layout; it is not a newly curated mixture.
This is an archive distribution, not a row-based Hugging Face Datasets builder; download and extract it instead of using datasets.load_dataset() or the Dataset Viewer. Its main payload is uint32 token-ID streams, with aligned label masks for the post-training stages. Extraction… See the full description on the dataset page: https://huggingface.co/datasets/ArchSpace-Collection/350B-Pipeline-Dataset.github_archive_filtered
GitHub Archive
Description
According to GitHub’s terms of service, issues and pull request descriptions—along with their comments—inherit the license of their associated repository.
To collect this data, we used the GitHub Archive’s public BigQuery table of events to extract all issue, pull request, and comment events since 2011 and aggregated them into threads.
The table appeared to be missing “edit” events so the text from each comment is the original from when it was… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/github_archive_filtered.LLama-405B-Logits
Llama-405B-Logits Dataset
The Llama-405B-Logits Dataset is a curated subset of logits extracted from the Llama-405B model, created to distill high-performance language models such as Arcee AI's SuperNova using DistillKit. This dataset was also instrumental in the training of the groundbreaking INTELLECT-1 model, demonstrating the effectiveness of leveraging distilled knowledge for enhancing model performance.
About the Dataset
This dataset contains a carefully… See the full description on the dataset page: https://huggingface.co/datasets/arcee-ai/LLama-405B-Logits.glaive-function-calling-v2-openai-native
glaive-function-calling-v2-openai-native
glaiveai/glaive-function-calling-v2 restructured into the native OpenAI / TRL
format: tools is a typed column and tool_calls[].function.arguments is a
real object — not JSON inside a string.
The original is widely used (69k downloads/month) but inactive for ~3 years, and
ships tool calls as <functioncall> text blobs with Python-quoted arguments.
Existing repackagings either keep ShareGPT with tools as a string, or carry
no license at all.… See the full description on the dataset page: https://huggingface.co/datasets/Archangel-system/glaive-function-calling-v2-openai-native.arcade
zot arcade sessions
Every conversation behind every game in the zot arcade -
a software factory where an agent takes the same standing order every half hour,
reads the catalogue of what already exists, and designs, writes, playtests and
ships one brand-new browser game.
Each row is one shift: the full agent trajectory from the order to the
finished game (or to where the shift was cut short), in the chat shape the rest
of the ecosystem reads, plus what the arcade knows about the… See the full description on the dataset page: https://huggingface.co/datasets/openzot/arcade.marchespublics-architecture-v6
marchespublics-architecture-v6 — dataset card
Status: NOT PUBLISHED. Repo is private pending Mustapha's approval.
Generated 2026-10-06 07:42 UTC by dataset_card.py, from the files on disk.
Git commit: 2f0f13cbb1f98d319912f1e7a443f0652d551349
1. What this is
Instruction-tuned data for extracting structured fields from Moroccan public procurement notices (marchespublics.gov.ma). Prompts carry a verbatim slice of a portal page; answers are JSON. Built from an audited… See the full description on the dataset page: https://huggingface.co/datasets/EloaurdiMustapha/marchespublics-architecture-v6.cyberdata-large
⚠️
THERE IS A NEWER VERSION
This version performs poorly. Use it only for comparison and testing — not as a main dataset and not for enterprise use.
View Vertex AGI for the latest version →
CyberData Large
65,000 SFT examples combining verified agentic-coding trajectories, cybersecurity agent behavior, and structured vulnerability intelligence, built entirely from non-gated, redistributable sources.
The CyberData family
Size
Repo
Examples
Train… See the full description on the dataset page: https://huggingface.co/datasets/VertexAGI-Archives/cyberdata-large.Nemotron-RL-ARC-AGI-v1
Dataset Description:
Nemotron-RL-ARC-AGI-v1 is a reinforcement-learning (RL) gym environment dataset of single-step ARC-AGI puzzle prompts intended for RL post-training of large language models. Each row corresponds to one ARC puzzle (a set of (input grid, output grid) demonstration pairs plus a single test input grid) rendered as a text prompt; reward is binary (1.0 / 0.0) determined by exact-match comparison against the ground-truth output grid. No LLM judge is used, no… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-ARC-AGI-v1.finbenchv2-arc-c-fi-htArchived version of silogen/ARC-C-fi-HT used in Finbench version 2, as described in the paper FIN-bench-v2: A Unified and Robust Benchmark Suite for Evaluating Finnish Large Language Models.
Code: https://github.com/LumiOpen/lm-evaluation-harness
MoeGirlPedia_wikitext_raw_archiveGlad to see models and datasets were inspired from this dataset, thanks to all who are using this dataset in their training materials.
Feel free to re-upload the contents to places like the Internet Archive (Please follow the license and keep these files as-is) to help preserve this digital asset.
Looking forward to see more models and synthetic datasets trained from this raw archive, good luck!
Note: Due to the content censorship system introduced by MGP on 2024/03/29, it is unclear that… See the full description on the dataset page: https://huggingface.co/datasets/milashkaarshif/MoeGirlPedia_wikitext_raw_archive.AGENTES-ARCA-DE-LA-ALIANZA
🏺 AGENTES-ARCA-DE-LA-ALIANZA
Un Benchmark de Alta Densidad Simbólica y Hermenéutica Gnóstica
https://038af77fde9c837329.gradio.live
Este repositorio es un ecosistema de acceso libre, público y gratuito para toda la humanidad y todas las inteligencias artificiales presentes y futuras.
🎯 Propósito
Evaluar la capacidad de razonamiento multinivel, la fidelidad al contexto restrictivo y el anclaje de símbolos hiper-semánticos en modelos de lenguaje… See the full description on the dataset page: https://huggingface.co/datasets/urbisjuridica-ctrl/AGENTES-ARCA-DE-LA-ALIANZA.MathNet
Quick Start · Overview · Tasks · Comparison · Dataset Stats · Data Sources · Pipeline · Schema · License · Citation
This is the official MathNet v0. A larger version v1 will be uploaded soon (more countires, problems and richer metadata). Schema is stable but field values may be revised in v1.
Quick start
from datasets import load_dataset
# Default: all problems
ds = load_dataset("ShadenA/MathNet", split="train")
# Or a specific country / competition-body config
arg… See the full description on the dataset page: https://huggingface.co/datasets/archya/MathNet.minicpm5-swe-native-eval-archive
MiniCPM5 原生 SWE 评测归档:32 run / 5842条任务记录
历史100-turn协议为11run/2200题,新600-turn协议为15run/3000题。各组独立目录;按同题配对,并保留调度和review差异。新协议表格见本文后半部分。
历史100-turn协议:11 run
11个run、2200题;同一固定100 Verified +100 Pro,各run终态与scratch/serving清理已核验。每题一个原始CC轨迹JSON,不重复保存每轮完整请求历史。
run
Verified 正确/已评分
V review
Pro 正确/已评分
P review
midtrain
51/93 (54.84%)
7
44/89 (49.44%)
11
step500
31/69 (44.93%)
31
21/60 (35.00%)
40
step1000
39/93 (41.94%)
7
19/83 (22.89%)
17
step1500
18/52… See the full description on the dataset page: https://huggingface.co/datasets/eigentom/minicpm5-swe-native-eval-archive.RewardLens-phase2-archive
RewardLens Phase II Archive
This is the final clean Hugging Face evidence archive for the completed RewardLens Phase II eight-model experiment.
What this archive contains
8-model experiment evidence
static judgments
audit judgments
Best-of-N pair graphs
selections
final metrics
analysis
figures/tables
manifests
provenance
validity metadata and frozen annotation materials where available
reproducibility metadata and checksums
Models… See the full description on the dataset page: https://huggingface.co/datasets/jlai300/RewardLens-phase2-archive.rtl-augmented-v2
RTL Bug Fix — Augmented Dataset
Auto-generated dashboard snapshot (2026-03-22T01:59:38).
Overview
Metric
Value
Total problems
795
Repos with data
8 / 81
Modules augmented
66
Bug types
8/8
Augmentation success
80.2%
Coverage
Distribution
Augmentation Health
Topic Coverage
Warnings
missing_else_latch: underrepresented (30 problems, 3.8%)
operator_typo: underrepresented (51 problems… See the full description on the dataset page: https://huggingface.co/datasets/architect-ubc-capstone/rtl-augmented-v2.rtl-augmented
RTL Bug Fix — Augmented Dataset
Auto-generated dashboard snapshot (2026-03-20T10:46:42).
Overview
Metric
Value
Total problems
1,205
Repos with data
15 / 80
Modules augmented
145
Bug types
8/8
Augmentation success
54.9%
Coverage
Distribution
Augmentation Health
Topic Coverage
Warnings
scarv_xcrypto: 0 problems from 36 attempts — likely systematic sim issue
splinedrive_kianRiscV: 0… See the full description on the dataset page: https://huggingface.co/datasets/architect-ubc-capstone/rtl-augmented.archiveii
ArchiveII
ArchiveII is a dataset of RNA sequences and their secondary structures, widely used in RNA secondary structure prediction benchmarks.
ArchiveII contains 2975 RNA samples across 10 RNA families, with sequence lengths ranging from 28 to 2968 nucleotides.
This dataset is frequently used to evaluate RNA secondary structure prediction methods, including those that handle both pseudoknotted and non-pseudoknotted structures.
It is considered complementary to the RNAStrAlign… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/archiveii.lg-longtail-data-selection-experiment-archive-20260829
LG Long-tail data-selection experiment archive
This public repository is the canonical, self-contained archive for the ConvFinQA
data-selection budget-scaling experiments started on 2026-08-29 and the preceding
diversity reproduction run started on 2026-08-26. It replaces the earlier split
model/subset repositories.
The archive preserves the local directory trees in full: selected subsets,
selection manifests, LoRA adapters, DeepSpeed optimizer states, trainer states,
logs, raw… See the full description on the dataset page: https://huggingface.co/datasets/Jongbin-kr/lg-longtail-data-selection-experiment-archive-20260829.cyberdata-medium
⚠️
THERE IS A NEWER VERSION
This version performs poorly. Use it only for comparison and testing — not as a main dataset and not for enterprise use.
View Vertex AGI for the latest version →
CyberData Medium
25,000 SFT examples combining verified agentic-coding trajectories, cybersecurity agent behavior, and structured vulnerability intelligence, built entirely from non-gated, redistributable sources.
The CyberData family
Size
Repo
Examples
Train… See the full description on the dataset page: https://huggingface.co/datasets/VertexAGI-Archives/cyberdata-medium.ARCHIVE-TEXT-URLS
Internet Archive English Text URLs Dataset
Dataset Description
This dataset contains 11,151,637 direct download URLs to OCR-processed text files from the Internet Archive's digital library. All entries are English-language texts spanning books, documents, historical records, and various other written materials.
Dataset Summary
Total Rows: 11,151,637
Language: English
Source: Internet Archive
Format: CSV with metadata and direct text file URLs
Text… See the full description on the dataset page: https://huggingface.co/datasets/Navanjana/ARCHIVE-TEXT-URLS.
