datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lichess_elo_binnededgar-corpusThe dataset contains annual filings (10K) of all publicly traded firms from 1993-2020. The table data is stripped but all text is retained.
This dataset allows easy access to the EDGAR-CORPUS dataset based on the paper EDGAR-CORPUS: Billions of Tokens Make The World Go Round (See References in README.md for details).lichess-elo-binned-2021-and-20242025-challenge-demosThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "R1Pro",
"total_episodes": 10000,
"total_frames": 119094660,
"total_tasks": 50,
"total_videos": 90000,
"chunks_size": 10000,
"fps": 30,
"splits": {
"train": "0:10000"
},
"data_path": "data/task-{episode_chunk:04d}/episode_{episode_index:08d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/elonelonelon/2025-challenge-demos.elonmusk01-backupchatbot-arena-elo
LMSYS Chatbot Arena ELO Scores
This dataset is a datasets-friendly version of Chatbot Arena ELO scores,
updated daily from the leaderboard API at
https://huggingface.co/spaces/lmarena-ai/chatbot-arena-leaderboard.
Updated: 20250717
Loading Data
from datasets import load_dataset
dataset = load_dataset("mathewhe/chatbot-arena-elo", split="train")
The main branch of this dataset will always be updated to the latest ELO and
leaderboard version. If you need a fixed dataset… See the full description on the dataset page: https://huggingface.co/datasets/mathewhe/chatbot-arena-elo.imagenet-elo-arena-images
ImageNet Elo Arena (public)
Isolated resources for the elo-v1 human study. See the ImageNetArena elo-arena branch for the locked protocol.
marchespublics-architecture-v6
marchespublics-architecture-v6 — dataset card
Status: NOT PUBLISHED. Repo is private pending Mustapha's approval.
Generated 2026-10-06 07:42 UTC by dataset_card.py, from the files on disk.
Git commit: 2f0f13cbb1f98d319912f1e7a443f0652d551349
1. What this is
Instruction-tuned data for extracting structured fields from Moroccan public procurement notices (marchespublics.gov.ma). Prompts carry a verbatim slice of a portal page; answers are JSON. Built from an audited… See the full description on the dataset page: https://huggingface.co/datasets/EloaurdiMustapha/marchespublics-architecture-v6.elosysELOQ
ELOQ
Description
ELOQ is a framework to generate out-of-scope questions for a given corpus.
License
The annotations (labels) created for this dataset are licensed under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.
The news content (titles, snippets, article texts) was crawled from publicly available news websites. Copyright of the news content remains with the original publishers.
This dataset is distributed for research purposes… See the full description on the dataset page: https://huggingface.co/datasets/zhiyuanpeng/ELOQ.Elon_Muskcots_yolo_dataset
🪸 CSIRO Crown-of-Thorns Starfish (COTS) Detection Dataset — YOLO Format
This dataset is a modified version of the CSIRO COTS and COTS Scars Dataset, originally released under the Creative Commons Attribution 4.0 License (CC BY 4.0).
The original dataset contains images and annotations for Crown-of-Thorns Starfish (COTS) and COTS scars, collected to support coral reef monitoring and control efforts on the Great Barrier Reef (GBR).
These starfish are coral predators, and their… See the full description on the dataset page: https://huggingface.co/datasets/eloise54/cots_yolo_dataset.elonmusklichess_elo_binned_debugcodeforces-editorial-elo-512-2026-04-28
Codeforces Editorial ELO 512 - 2026-04-28
A 512-example subset sampled from open-r1/codeforces (verifiable, train) for Plan-CRL Codeforces feedback experiments that need a non-empty trusted reference field.
Important: open-r1/codeforces does not expose code reference solutions. This subset fills reference_solution from the source dataset's editorial field. Treat it as a natural-language editorial/rationale, not canonical reference code.
Selection seed: 20260428.
Filtering:
rating… See the full description on the dataset page: https://huggingface.co/datasets/1337xyz1337xyz/codeforces-editorial-elo-512-2026-04-28.InfLLM-V2-data-5B-v2
InfLLM-V2 Long-Context Training Dataset with 5B Tokens
Project Links: [Paper] [InfLLM-V2 Models] [CUDA Kernel Code]
🚀 About InfLLM-V2
InfLLM-V2 is a native sparse attention framework designed for the efficient processing of long-sequence texts. Its core advantage is the ability to maintain high performance comparable to dense attention in short-text scenarios—without any extra parameters—while seamlessly switching to a sparse mode for long-text scenarios, achieving… See the full description on the dataset page: https://huggingface.co/datasets/elonmuskceo/InfLLM-V2-data-5B-v2.video_to_figurearena-elo-by-axis
Arena Elo by axis (not a board score)
A pairwise-arena Elo reference computed per axis. elo_reference.json carries the
axis list, a content_id over the payload, and a leaderboard row per model: elo, games, winrate and a
confidence interval. This is an arena number over head-to-head rounds — it is not a GSPC board score and
does not write to any slot.
The live board is the authority
GET https://councilof.ai/api/gspc — quote totals.public_count. This Hub card is… See the full description on the dataset page: https://huggingface.co/datasets/csoai/arena-elo-by-axis.Robustness
ELOQUENT Robustness and Consistency Task
This dataset contains the sample and test datasets for the Robustness and Consistency task, which is part of the ELOQUENT lab. This dataset is for participants to generate texts for prompt variants, to investigate prompt style conditioned variation.
Robustness task
ELOQUENT lab
CLEF conference 9-12 September 2025
The task in brief (this is a simple task to execute!)
This dataset provides a number of questions in several… See the full description on the dataset page: https://huggingface.co/datasets/Eloquent/Robustness.DD-Elo-DataminiCodeProps
Getting Started
First install Lean 4. Then clone this repo:
git clone --recurse-submodules https://huggingface.co/datasets/elohn/miniCodeProps
The outer LeanSrc folder is a Lean Project. You can open that folder directly in VSCode and check that the proofs in LeanSrc/Sorts.lean type check after following the instructions for working on an existing lean project in the Lean 4 documentation.
The main miniCodeProps folder handles extracting the benchmark and calculating baselines. If… See the full description on the dataset page: https://huggingface.co/datasets/elohn/miniCodeProps.dataset
EloPhanto Agent Trajectories
Real-world conversation and tool-use trajectories collected automatically from EloPhanto – an open-source autonomous AI agent with an evolving self-model (identity, ego, affect, autonomous mind) that builds zero-human businesses, ships code, manages on-chain assets, and grows audiences without human supervision.
This dataset is automatically generated: every time an EloPhanto instance completes a task, the full conversation – user turn, assistant… See the full description on the dataset page: https://huggingface.co/datasets/EloPhanto/dataset.elon-style-dataset
Elon Style Dataset — elon-style-dataset
Fine-tuning dataset for training a conversational model that mimics
Elon Musk's private texting style: short punchy replies, sarcasm, visionary takes,
crypto/tech opinions, and authentic multi-turn cadence.
Dataset Stats
Split
Examples
Avg Response Length
train
20,900
~88 chars
validation
1,100
~91 chars
47–50% of responses contain emoji (authentic texting energy)
57% short replies (<60 chars) — punchy, human
33%… See the full description on the dataset page: https://huggingface.co/datasets/ceoelonmusk/elon-style-dataset.Omni-LIVOInfLLM-V2-data-5B
InfLLM-V2 Long-Context Training Dataset with 5B Tokens
Project Links: [Paper] [InfLLM-V2 Models] [CUDA Kernel Code]
🚀 About InfLLM-V2
InfLLM-V2 is a native sparse attention framework designed for the efficient processing of long-sequence texts. Its core advantage is the ability to maintain high performance comparable to dense attention in short-text scenarios—without any extra parameters—while seamlessly switching to a sparse mode for long-text scenarios, achieving… See the full description on the dataset page: https://huggingface.co/datasets/elonmuskceo/InfLLM-V2-data-5B.GroundCUA
GroundCUA: Grounding Computer Use Agents on Human Demonstrations
🌐 Website |
📑 Paper |
🤗 Dataset |
🤖 Models
GroundCUA Dataset
GroundCUA is a large and diverse dataset of real UI screenshots paired with structured annotations for building multimodal computer use agents. It covers 87 software platforms across productivity tools, browsers, creative tools, communication apps, development environments, and system utilities. GroundCUA is designed for research on GUI… See the full description on the dataset page: https://huggingface.co/datasets/ElonMusk-v2/GroundCUA.Hellenic-greek-parliamentary-speech
HParl: Hellenic Parliamentary Speech Corpus
Dataset Description
Note: This is a processed version of the original HParl dataset. This dataset is not created or maintained by the original authors.
Link to the original source: https://inventory.clarin.gr/corpus/1602
HParl is a 120-hour speech corpus for Modern Greek, originally collected from parliamentary proceedings of the Hellenic Parliament by the Institute for Language and Speech Processing. This version has been… See the full description on the dataset page: https://huggingface.co/datasets/Elormiden/Hellenic-greek-parliamentary-speech.eloquent-2026marchespublics-architecture-v4
marchespublics-architecture-v4
SFT corpus for fine-tuning a model to read the structure of Moroccan public-procurement
notices on marchespublics.gov.ma, rather than to
answer ad-hoc questions about one card.
v3 of this corpus taught a model to answer questions about a single avis. v4 teaches the
schema those cards are instances of: which feed a notice comes from, which fields that
feed can structurally publish, the entity graph (avis → lots → bordereau → résultat →
fournisseur →… See the full description on the dataset page: https://huggingface.co/datasets/EloaurdiMustapha/marchespublics-architecture-v4.elonmusk
AMI meeting subset (test audio)
Subset of the AMI Meeting Corpus (https://groups.inf.ed.ac.uk/ami/corpus/), CC BY 4.0.
Meetings: ES2008a-d, IS1009a-d.
ihm/: Headset mix (all individual headsets mixed to one file), as released.
sdm/: Single distant microphone, Array1-01 of the microphone array, converted to mono 16 kHz.
Attribution: Carletta et al., The AMI Meeting Corpus. Used for development/testing only.
