Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01open-index /arctic Arctic Shift Reddit Archive Every Reddit comment and submission since 2005, organized as monthly Parquet shards What is it? The full Reddit archive from Arctic Shift, converted to Parquet and hosted here for easy access. Covers every public subreddit from 2005-12 through 2026-02. Right now the archive has 15.7B items (12.9B comments, 2.8B submissions) in 1.3 TB of compressed Parquet. Comments and submissions are stored as separate datasets, split into monthly… See the full description on the dataset page: https://huggingface.co/datasets/open-index/arctic.text-generation1B<n<10B30 likes134k downloads2mo agoHugging Face02arcinstitute /opengenome2 OpenGenome2 OpenGenome2 is a database of nearly 9 trillion base pairs of curated DNA from across all domains of life. Collected from diverse species and public data sources, OpenGenome2 was used to train Evo 2 models. Please refer to the Evo 2 preprint or github repository for further details and usage examples. We provide OpenGenome2 in two formats, the dataset is organized into two main directories to reflect this: fasta which contain the DNA sequences jsonl which… See the full description on the dataset page: https://huggingface.co/datasets/arcinstitute/opengenome2.text-generationn>1T162 likes47k downloads1mo agoHugging Face03AiAF /SCPWiki-Cleaned-PDF-Archivesdocumenttext-generationn<1K1 likes5.4k downloads1y agoHugging Face04common-pile /github_archive GitHub Archive Description According to GitHub’s terms of service, issues and pull request descriptions—along with the their comments—inherit the license of their associated repository. To collect this data, we used the GitHub Archive’s public BigQuery table of events to extracted all issue, pull request, and comment events since 2011 and aggregated them into threads. The table appeared to be missing “edit” events so the text from each comment is the original from when… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/github_archive.texttext-generation10M<n<100M2 likes3.3k downloads1y agoHugging Face05blairducrayoppat /openvino-arc140v-lunarlake OpenVINO local-inference on an Intel Arc 140V (Lunar Lake) iGPU Reference performance data for running local models on a single Intel Core Ultra 7 258V (Lunar Lake) laptop with the integrated Intel Arc 140V (Xe2) GPU, via OpenVINO. All inference runs on the iGPU; the NPU stays idle throughout, confirmed by the telemetry here. This is reference characterization shared by a non-expert contributor — careful measurements on one machine, offered so others can compare and correct, not… See the full description on the dataset page: https://huggingface.co/datasets/blairducrayoppat/openvino-arc140v-lunarlake.tabulartext-generation1K<n<10K0 likes3k downloads2d agoHugging Face06nvidia /Nemotron-SFT-ARC-AGI-v1 Dataset Description: Nemotron-SFT-ARC-AGI-v1 is a supervised fine-tuning (SFT) dataset of multi-turn agentic reasoning traces produced by open-weight large language models attempting to solve ARC-AGI visual-reasoning puzzles. Each ARC puzzle (a set of (input grid, output grid) demonstration pairs plus one or more test inputs, where grids are 2D integer arrays representing colors) is formatted as a text prompt and given to an agent powered by one of nine open-weight reasoning… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SFT-ARC-AGI-v1.texttext-generation100K<n<1M24 likes2.1k downloads4mo agoHugging Face07architect-ubc-capstone /rtl-augmented-v3 RTL Bug Fix — Augmented Dataset Auto-generated dashboard snapshot (2026-04-14T10:53:43). Overview Metric Value Total problems 718 Repos with data 57 / 81 Modules augmented 408 Bug types 11/11 Augmentation success 48.9% Coverage Distribution Augmentation Health Topic Coverage Warnings lucky-wfw_IC_System_Design: 0 problems from 48 attempts — likely systematic sim issue meiniKi_FazyRV:… See the full description on the dataset page: https://huggingface.co/datasets/architect-ubc-capstone/rtl-augmented-v3.text-generation1K<n<10K0 likes1.8k downloads6mo agoHugging Face08nyuuzyou /google-code-archive Google Code Archive Dataset Dataset Description This dataset was compiled from the Google Code Archive, a preserved snapshot of projects hosted on Google Code, Google's open-source project hosting service that operated from 2006 to 2016. Google Code was one of the major code hosting platforms of its era, hosting hundreds of thousands of open-source projects before its shutdown. The archive provides a unique historical record of open-source development during a formative… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/google-code-archive.texttext-generation10M<n<100M73 likes1.7k downloads8mo agoHugging Face09VertexAGI-Archives /cyberdata-full ⚠️ THERE IS A NEWER VERSION This version performs poorly. Use it only for comparison and testing — not as a main dataset and not for enterprise use. View Vertex AGI for the latest version → CyberData Full 479,214 SFT examples: every example from all five source datasets (verified agentic-coding trajectories, agent behavior, and structured vulnerability intelligence), built entirely from non-gated, redistributable sources. The CyberData family… See the full description on the dataset page: https://huggingface.co/datasets/VertexAGI-Archives/cyberdata-full.text-generation100K<n<1M0 likes1.5k downloads5d agoHugging Face10Dk587 /arctic Arctic Shift Reddit Archive Every Reddit comment and submission since 2005, organized as monthly Parquet shards What is it? The full Reddit archive from Arctic Shift, converted to Parquet and hosted here for easy access. Covers every public subreddit from 2005-12 through 2026-02. Right now the archive has 1.6B items (362.1M comments, 1.2B submissions) in 181.4 GB of compressed Parquet. Comments and submissions are stored as separate datasets, split into monthly shards… See the full description on the dataset page: https://huggingface.co/datasets/Dk587/arctic.tabulartext-generation1B<n<10B1 likes1.4k downloads7mo agoHugging Face11ArchSpace-Collection /350B-Pipeline-Dataset OLMo 3 pre-tokenized and processed training data This repository redistributes Ai2/AllenAI's official OLMo 3 training data in an OLMo-core-ready archive layout; it is not a newly curated mixture. This is an archive distribution, not a row-based Hugging Face Datasets builder; download and extract it instead of using datasets.load_dataset() or the Dataset Viewer. Its main payload is uint32 token-ID streams, with aligned label masks for the post-training stages. Extraction… See the full description on the dataset page: https://huggingface.co/datasets/ArchSpace-Collection/350B-Pipeline-Dataset.texttext-generation0 likes1.3k downloads2mo agoHugging Face12common-pile /github_archive_filtered GitHub Archive Description According to GitHub’s terms of service, issues and pull request descriptions—along with their comments—inherit the license of their associated repository. To collect this data, we used the GitHub Archive’s public BigQuery table of events to extract all issue, pull request, and comment events since 2011 and aggregated them into threads. The table appeared to be missing “edit” events so the text from each comment is the original from when it was… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/github_archive_filtered.texttext-generation10M<n<100M2 likes1.3k downloads1y agoHugging Face13arcee-ai /LLama-405B-Logits Llama-405B-Logits Dataset The Llama-405B-Logits Dataset is a curated subset of logits extracted from the Llama-405B model, created to distill high-performance language models such as Arcee AI's SuperNova using DistillKit. This dataset was also instrumental in the training of the groundbreaking INTELLECT-1 model, demonstrating the effectiveness of leveraging distilled knowledge for enhancing model performance. About the Dataset This dataset contains a carefully… See the full description on the dataset page: https://huggingface.co/datasets/arcee-ai/LLama-405B-Logits.text-generation10K<n<100K14 likes1.1k downloads2y agoHugging Face14Archangel-system /glaive-function-calling-v2-openai-native glaive-function-calling-v2-openai-native glaiveai/glaive-function-calling-v2 restructured into the native OpenAI / TRL format: tools is a typed column and tool_calls[].function.arguments is a real object — not JSON inside a string. The original is widely used (69k downloads/month) but inactive for ~3 years, and ships tool calls as <functioncall> text blobs with Python-quoted arguments. Existing repackagings either keep ShareGPT with tools as a string, or carry no license at all.… See the full description on the dataset page: https://huggingface.co/datasets/Archangel-system/glaive-function-calling-v2-openai-native.texttext-generation10K<n<100K2 likes1.1k downloads27d agoHugging Face15openzot /arcade zot arcade sessions Every conversation behind every game in the zot arcade - a software factory where an agent takes the same standing order every half hour, reads the catalogue of what already exists, and designs, writes, playtests and ships one brand-new browser game. Each row is one shift: the full agent trajectory from the order to the finished game (or to where the shift was cut short), in the chat shape the rest of the ecosystem reads, plus what the arcade knows about the… See the full description on the dataset page: https://huggingface.co/datasets/openzot/arcade.text-generationn<1K2 likes909 downloads29m agoHugging Face16EloaurdiMustapha /marchespublics-architecture-v6 marchespublics-architecture-v6 — dataset card Status: NOT PUBLISHED. Repo is private pending Mustapha's approval. Generated 2026-10-06 07:42 UTC by dataset_card.py, from the files on disk. Git commit: 2f0f13cbb1f98d319912f1e7a443f0652d551349 1. What this is Instruction-tuned data for extracting structured fields from Moroccan public procurement notices (marchespublics.gov.ma). Prompts carry a verbatim slice of a portal page; answers are JSON. Built from an audited… See the full description on the dataset page: https://huggingface.co/datasets/EloaurdiMustapha/marchespublics-architecture-v6.text-generation10K<n<100K0 likes847 downloads5d agoHugging Face17VertexAGI-Archives /cyberdata-large ⚠️ THERE IS A NEWER VERSION This version performs poorly. Use it only for comparison and testing — not as a main dataset and not for enterprise use. View Vertex AGI for the latest version → CyberData Large 65,000 SFT examples combining verified agentic-coding trajectories, cybersecurity agent behavior, and structured vulnerability intelligence, built entirely from non-gated, redistributable sources. The CyberData family Size Repo Examples Train… See the full description on the dataset page: https://huggingface.co/datasets/VertexAGI-Archives/cyberdata-large.text-generation10K<n<100K0 likes743 downloads5d agoHugging Face18nvidia /Nemotron-RL-ARC-AGI-v1 Dataset Description: Nemotron-RL-ARC-AGI-v1 is a reinforcement-learning (RL) gym environment dataset of single-step ARC-AGI puzzle prompts intended for RL post-training of large language models. Each row corresponds to one ARC puzzle (a set of (input grid, output grid) demonstration pairs plus a single test input grid) rendered as a text prompt; reward is binary (1.0 / 0.0) determined by exact-match comparison against the ground-truth output grid. No LLM judge is used, no… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-ARC-AGI-v1.texttext-generation10K<n<100K11 likes694 downloads4mo agoHugging Face19TurkuNLP /finbenchv2-arc-c-fi-htArchived version of silogen/ARC-C-fi-HT used in Finbench version 2, as described in the paper FIN-bench-v2: A Unified and Robust Benchmark Suite for Evaluating Finnish Large Language Models. Code: https://github.com/LumiOpen/lm-evaluation-harness texttext-classification1K<n<10K0 likes634 downloads10mo agoHugging Face20milashkaarshif /MoeGirlPedia_wikitext_raw_archiveGlad to see models and datasets were inspired from this dataset, thanks to all who are using this dataset in their training materials. Feel free to re-upload the contents to places like the Internet Archive (Please follow the license and keep these files as-is) to help preserve this digital asset. Looking forward to see more models and synthetic datasets trained from this raw archive, good luck! Note: Due to the content censorship system introduced by MGP on 2024/03/29, it is unclear that… See the full description on the dataset page: https://huggingface.co/datasets/milashkaarshif/MoeGirlPedia_wikitext_raw_archive.texttext-generation100K<n<1M38 likes564 downloads8mo agoHugging Face21urbisjuridica-ctrl /AGENTES-ARCA-DE-LA-ALIANZA 🏺 AGENTES-ARCA-DE-LA-ALIANZA Un Benchmark de Alta Densidad Simbólica y Hermenéutica Gnóstica https://038af77fde9c837329.gradio.live Este repositorio es un ecosistema de acceso libre, público y gratuito para toda la humanidad y todas las inteligencias artificiales presentes y futuras. 🎯 Propósito Evaluar la capacidad de razonamiento multinivel, la fidelidad al contexto restrictivo y el anclaje de símbolos hiper-semánticos en modelos de lenguaje… See the full description on the dataset page: https://huggingface.co/datasets/urbisjuridica-ctrl/AGENTES-ARCA-DE-LA-ALIANZA.texttext-generationn<1K0 likes496 downloads10d agoHugging Face22archya /MathNet Quick Start · Overview · Tasks · Comparison · Dataset Stats · Data Sources · Pipeline · Schema · License · Citation This is the official MathNet v0. A larger version v1 will be uploaded soon (more countires, problems and richer metadata). Schema is stable but field values may be revised in v1. Quick start from datasets import load_dataset # Default: all problems ds = load_dataset("ShadenA/MathNet", split="train") # Or a specific country / competition-body config arg… See the full description on the dataset page: https://huggingface.co/datasets/archya/MathNet.imagequestion-answering10K<n<100K0 likes464 downloads5mo agoHugging Face23eigentom /minicpm5-swe-native-eval-archive MiniCPM5 原生 SWE 评测归档:32 run / 5842条任务记录 历史100-turn协议为11run/2200题,新600-turn协议为15run/3000题。各组独立目录;按同题配对,并保留调度和review差异。新协议表格见本文后半部分。 历史100-turn协议:11 run 11个run、2200题;同一固定100 Verified +100 Pro,各run终态与scratch/serving清理已核验。每题一个原始CC轨迹JSON,不重复保存每轮完整请求历史。 run Verified 正确/已评分 V review Pro 正确/已评分 P review midtrain 51/93 (54.84%) 7 44/89 (49.44%) 11 step500 31/69 (44.93%) 31 21/60 (35.00%) 40 step1000 39/93 (41.94%) 7 19/83 (22.89%) 17 step1500 18/52… See the full description on the dataset page: https://huggingface.co/datasets/eigentom/minicpm5-swe-native-eval-archive.texttext-generation1K<n<10K0 likes449 downloads3d agoHugging Face24jlai300 /RewardLens-phase2-archive RewardLens Phase II Archive This is the final clean Hugging Face evidence archive for the completed RewardLens Phase II eight-model experiment. What this archive contains 8-model experiment evidence static judgments audit judgments Best-of-N pair graphs selections final metrics analysis figures/tables manifests provenance validity metadata and frozen annotation materials where available reproducibility metadata and checksums Models… See the full description on the dataset page: https://huggingface.co/datasets/jlai300/RewardLens-phase2-archive.visual-question-answering0 likes418 downloads25d agoHugging Face25architect-ubc-capstone /rtl-augmented-v2 RTL Bug Fix — Augmented Dataset Auto-generated dashboard snapshot (2026-03-22T01:59:38). Overview Metric Value Total problems 795 Repos with data 8 / 81 Modules augmented 66 Bug types 8/8 Augmentation success 80.2% Coverage Distribution Augmentation Health Topic Coverage Warnings missing_else_latch: underrepresented (30 problems, 3.8%) operator_typo: underrepresented (51 problems… See the full description on the dataset page: https://huggingface.co/datasets/architect-ubc-capstone/rtl-augmented-v2.text-generation1K<n<10K0 likes378 downloads7mo agoHugging Face26architect-ubc-capstone /rtl-augmented RTL Bug Fix — Augmented Dataset Auto-generated dashboard snapshot (2026-03-20T10:46:42). Overview Metric Value Total problems 1,205 Repos with data 15 / 80 Modules augmented 145 Bug types 8/8 Augmentation success 54.9% Coverage Distribution Augmentation Health Topic Coverage Warnings scarv_xcrypto: 0 problems from 36 attempts — likely systematic sim issue splinedrive_kianRiscV: 0… See the full description on the dataset page: https://huggingface.co/datasets/architect-ubc-capstone/rtl-augmented.text-generation1K<n<10K0 likes354 downloads7mo agoHugging Face27multimolecule /archiveii ArchiveII ArchiveII is a dataset of RNA sequences and their secondary structures, widely used in RNA secondary structure prediction benchmarks. ArchiveII contains 2975 RNA samples across 10 RNA families, with sequence lengths ranging from 28 to 2968 nucleotides. This dataset is frequently used to evaluate RNA secondary structure prediction methods, including those that handle both pseudoknotted and non-pseudoknotted structures. It is considered complementary to the RNAStrAlign… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/archiveii.texttext-generation1K<n<10K0 likes312 downloads3mo agoHugging Face28Jongbin-kr /lg-longtail-data-selection-experiment-archive-20260829 LG Long-tail data-selection experiment archive This public repository is the canonical, self-contained archive for the ConvFinQA data-selection budget-scaling experiments started on 2026-08-29 and the preceding diversity reproduction run started on 2026-08-26. It replaces the earlier split model/subset repositories. The archive preserves the local directory trees in full: selected subsets, selection manifests, LoRA adapters, DeepSpeed optimizer states, trainer states, logs, raw… See the full description on the dataset page: https://huggingface.co/datasets/Jongbin-kr/lg-longtail-data-selection-experiment-archive-20260829.text-generation0 likes310 downloads18d agoHugging Face29VertexAGI-Archives /cyberdata-medium ⚠️ THERE IS A NEWER VERSION This version performs poorly. Use it only for comparison and testing — not as a main dataset and not for enterprise use. View Vertex AGI for the latest version → CyberData Medium 25,000 SFT examples combining verified agentic-coding trajectories, cybersecurity agent behavior, and structured vulnerability intelligence, built entirely from non-gated, redistributable sources. The CyberData family Size Repo Examples Train… See the full description on the dataset page: https://huggingface.co/datasets/VertexAGI-Archives/cyberdata-medium.text-generation10K<n<100K0 likes307 downloads5d agoHugging Face30Navanjana /ARCHIVE-TEXT-URLS Internet Archive English Text URLs Dataset Dataset Description This dataset contains 11,151,637 direct download URLs to OCR-processed text files from the Internet Archive's digital library. All entries are English-language texts spanning books, documents, historical records, and various other written materials. Dataset Summary Total Rows: 11,151,637 Language: English Source: Internet Archive Format: CSV with metadata and direct text file URLs Text… See the full description on the dataset page: https://huggingface.co/datasets/Navanjana/ARCHIVE-TEXT-URLS.texttext-generation1M<n<10M1 likes298 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.