datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
marchespublics-architecture-v6
marchespublics-architecture-v6 — dataset card
Status: NOT PUBLISHED. Repo is private pending Mustapha's approval.
Generated 2026-10-06 07:42 UTC by dataset_card.py, from the files on disk.
Git commit: 2f0f13cbb1f98d319912f1e7a443f0652d551349
1. What this is
Instruction-tuned data for extracting structured fields from Moroccan public procurement notices (marchespublics.gov.ma). Prompts carry a verbatim slice of a portal page; answers are JSON. Built from an audited… See the full description on the dataset page: https://huggingface.co/datasets/EloaurdiMustapha/marchespublics-architecture-v6.codeforces-editorial-elo-512-2026-04-28
Codeforces Editorial ELO 512 - 2026-04-28
A 512-example subset sampled from open-r1/codeforces (verifiable, train) for Plan-CRL Codeforces feedback experiments that need a non-empty trusted reference field.
Important: open-r1/codeforces does not expose code reference solutions. This subset fills reference_solution from the source dataset's editorial field. Treat it as a natural-language editorial/rationale, not canonical reference code.
Selection seed: 20260428.
Filtering:
rating… See the full description on the dataset page: https://huggingface.co/datasets/1337xyz1337xyz/codeforces-editorial-elo-512-2026-04-28.dataset
EloPhanto Agent Trajectories
Real-world conversation and tool-use trajectories collected automatically from EloPhanto – an open-source autonomous AI agent with an evolving self-model (identity, ego, affect, autonomous mind) that builds zero-human businesses, ships code, manages on-chain assets, and grows audiences without human supervision.
This dataset is automatically generated: every time an EloPhanto instance completes a task, the full conversation – user turn, assistant… See the full description on the dataset page: https://huggingface.co/datasets/EloPhanto/dataset.elon-style-dataset
Elon Style Dataset — elon-style-dataset
Fine-tuning dataset for training a conversational model that mimics
Elon Musk's private texting style: short punchy replies, sarcasm, visionary takes,
crypto/tech opinions, and authentic multi-turn cadence.
Dataset Stats
Split
Examples
Avg Response Length
train
20,900
~88 chars
validation
1,100
~91 chars
47–50% of responses contain emoji (authentic texting energy)
57% short replies (<60 chars) — punchy, human
33%… See the full description on the dataset page: https://huggingface.co/datasets/ceoelonmusk/elon-style-dataset.rukh-elo-bins
chorcat/rukh-elo-bins
A balanced sample of rukh-games-1800: up to n_per_bin games per 100-Elo bin of the average rating of both players, for Elo conditioning experiments.
Part of Rukh, a chess language model built from scratch
as a course on generative and agentic AI. Every derived dataset ships with the exact filters and
counts of its manifest.json, so it can be regenerated with rukh data elo-bins.
Files
File
Bytes
SHA-256
games.parquet
65955015… See the full description on the dataset page: https://huggingface.co/datasets/chorcat/rukh-elo-bins.lichess_pretrain_elo_cutoff
Lichess pretraining games, binned by Elo
All 307,426,466 non-bullet Lichess games in our chess pretraining set, split into 11 bins of
200 Elo by the game's average rating, from 800 to 3000. Each bin is provided both as text
(one game per line, SAN movetext) and as pre-tokenized uint16 shards ready for the
LLM-Pretraining trainer. In total that is 69,516,616,889 tokens (69.52B).
Bins
A game goes in bin [lo, hi) when lo <= avg_elo < hi, where
avg_elo = (white_elo +… See the full description on the dataset page: https://huggingface.co/datasets/Pre-to-Post-2/lichess_pretrain_elo_cutoff.Elongated_CACAPO_for_E2EDataset information can be found in the JSON file named "elongated_training_cacapo_updated-02_22_2023_23_23_20.json", which was created with the interactive dataset creator provided by Huggingface.
preference_prediction
Dataset Card for Preference Prediction
Updates
10.03.2025: 🔥 release of the private test set, baseline, and evaluation script
31.01.2025: 🚀 release of the development set
Dataset Description
Our dataset is introduced in the Preference Prediction task, which is part of the 2025 ELOQUENT lab and aims to evaluate the capability of systems to predict human preferences for different outputs from generative large language models (LLMs) and explain… See the full description on the dataset page: https://huggingface.co/datasets/Eloquent/preference_prediction.codeforces-elo-512-2026-04-28
Codeforces ELO 512 - 2026-04-28
A 512-example subset sampled from open-r1/codeforces (verifiable, train) for Plan-CRL Codeforces feedback experiments.
Selection seed: 20260428.
Filtering:
rating present
complete official tests
stdio-only tasks
no generated checker
official tests present
at most 16 official test cases
Bucket balance:
elo_0800_1000: 128
elo_1100_1500: 128
elo_1600_2200: 128
elo_2300_3500: 128
Useful fields for manual inspection:
prompt
reference_solution
tests… See the full description on the dataset page: https://huggingface.co/datasets/1337xyz1337xyz/codeforces-elo-512-2026-04-28.elocus
E-Locus
Dataset Language:Greek (primary), with English alternative titles, abstracts and bibliographic terms throughout, and occasional French or German material in older volumes. The content column carries the full Greek text of each thesis; the alt_title and most of the subjects field are English by convention of the source repository.
Dataset Info:This dataset consists of text-extracted PDFs from E-Locus (elocus.lib.uoc.gr), the Institutional Repository of the Library and… See the full description on the dataset page: https://huggingface.co/datasets/glossAPI/elocus.ciclope-mitologias-verbales
Cíclope: Mitologías Verbales
Descripción
Cíclope es un sistema de lectura de segundo orden que genera TSRs (Thematic Semantic Reports) mediante una arquitectura de 7 capas progresivas. Cada TSR es un documento monolítico que consolida:
CAPA 0: Semilla conceptual (quote detonante)
CAPA 1: Bibliografía verificada
CAPA 2: Genealogía conceptual
CAPA 3: Problematización contemporánea
CAPA 4: Resonancias con Reflejos Híbridos
CAPA 5: Meta-análisis
CAPA 6: Guion de taller
CAPA… See the full description on the dataset page: https://huggingface.co/datasets/EloiseCry/ciclope-mitologias-verbales.
