datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
speculators-ci-datasets
speculator-tutorial
Raw vs. on-policy regenerated conversation data for training speculative-decoding
drafters (EAGLE-3 / DFlash / DSpark style), with the original source data kept alongside
so you can see exactly what regeneration changes and why it matters.
Prompts come from UltraChat-200k. The verifier / teacher model is Qwen/Qwen3-8B.
Why regenerate at all?
A speculative-decoding drafter is trained to predict what the verifier would say next.
If you train it… See the full description on the dataset page: https://huggingface.co/datasets/inference-optimization/speculators-ci-datasets.Alexandria_geometry_optimization_paths_PBE_2D
Cite this dataset Schmidt, J., Hoffmann, N., Wang, H., Borlido, P., Carriço, P. J. M. A., Cerqueira, T. F. T., Botti, S., and Marques, M. A. L. Alexandria geometry optimization paths PBE 2D. ColabFit, 2025. https://doi.org/10.60732/8781419f
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_6pieq95jrqpn_0
Visit the ColabFit… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/Alexandria_geometry_optimization_paths_PBE_2D.Alexandria_geometry_optimization_paths_PBE_3D
Cite this dataset Schmidt, J., Hoffmann, N., Wang, H., Borlido, P., Carriço, P. J. M. A., Cerqueira, T. F. T., Botti, S., and Marques, M. A. L. Alexandria geometry optimization paths PBE 3D. ColabFit, 2024. https://doi.org/10.60732/c88da7df
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_s6gf4z2hcjqy_0
Visit the ColabFit… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/Alexandria_geometry_optimization_paths_PBE_3D.Alexandria_geometry_optimization_paths_PBE_1D
Cite this dataset Schmidt, J., Hoffmann, N., Wang, H., Borlido, P., Carriço, P. J. M. A., Cerqueira, T. F. T., Botti, S., and Marques, M. A. L. Alexandria geometry optimization paths PBE 1D. ColabFit, 2025. https://doi.org/10.60732/12246d46
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_xnio123pebli_0
Visit the ColabFit… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/Alexandria_geometry_optimization_paths_PBE_1D.4c_optimizationThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "koch_follower",
"total_episodes": 25,
"total_frames": 8494,
"total_tasks": 1,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:25"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ethanCSL/4c_optimization.repro-optimal-regret-for-policy-optimization-in-contextual-bandits-traces
Agent traces
Agent sessions published from a Trackio Logbook.
CARMO-UltraFeedbacknl-optimization-instantiation-metrics
Natural-Language Optimization Instantiation Metrics (v1.2)
This dataset is a text-free metrics release for a retrieval-assisted natural-language optimization instantiation pipeline. It contains per-example and aggregate evaluation outcomes for a frozen pipeline evaluated on the NLP4LP benchmark, the OptMath external validation domain (added in v1.1), and, as of v1.2, a 13-method schema-retrieval/grounding diagnostic suite (per-query breakdowns, bottleneck taxonomy, threshold… See the full description on the dataset page: https://huggingface.co/datasets/SoroushVahidi/nl-optimization-instantiation-metrics.manufacturing-cost-optimization-2026Q3
Manufacturing Quarterly Cost Optimization Dataset (2026Q3)
Unified quarterly cost-analysis dataset for the manufacturing group, merged from
the China / Japan / India factory datasets hosted on Hugging Face.
Contents
11,100 records (>= 10,000) covering 9 plants across 3 regions.
Source datasets:
toolathon123/manufacturing-cn-energy-2026Q3 — China energy & raw material (4,200 rows)
toolathon123/manufacturing-jp-maintenance-2026Q3 — Japan maintenance & downtime (3… See the full description on the dataset page: https://huggingface.co/datasets/toolathon123/manufacturing-cost-optimization-2026Q3.synthesized-cloud-optimization-recommendations
Synthesized Cloud-Optimization Recommendations
18 scenarios that pair cloud telemetry with a hand-crafted optimization
recommendation. Use them to train models or to evaluate AI agents.
Summary
Each scenario has multi-tier telemetry, a Terraform file describing the
deployed infrastructure, and a gold-standard recommendation.
The dataset is built around a simple input-output mapping. The input is
telemetry plus the infrastructure. The output is an optimization… See the full description on the dataset page: https://huggingface.co/datasets/ameau01/synthesized-cloud-optimization-recommendations.repro-a-tight-theory-of-error-feedback-algorithms-in-distributed-optimization-traces
Agent traces
Agent sessions published from a Trackio Logbook.
repro-on-regret-bounds-of-thompson-sampling-for-bayesian-optimization-traces
Agent traces
Agent sessions published from a Trackio Logbook.
AI-Code-Optimization-for-Sustainability-Dataset
AI Code Optimization for Sustainability: Dataset
Refactoring Python Code for Energy-Efficiency using Qwen3: Dataset based on HumanEval, MBPP, and Mercury
📄 Read the Paper | Zenodo Mirror | DOI: 10.5281/zenodo.18377893 | About the author
This dataset is a part of a Master thesis research internship investigating the use of LLMs to optimize Python code for energy efficiency.
The research was conducted as part of the Greenify My Code (GMC) project at the Netherlands Organisation for… See the full description on the dataset page: https://huggingface.co/datasets/BambusControl/AI-Code-Optimization-for-Sustainability-Dataset.quantum-optimization
Neura Parse — Quantum Optimization, Annealing & Finance: QAOA, Adiabatic Methods & the Advantage Question
A research-plus-practitioner vertical on quantum approaches to combinatorial and continuous optimization and their most-piloted enterprise use cases. Covers QAOA theory and variants, adiabatic/annealing methods and D-Wave, QUBO/Ising encodings, amplitude-estimation Monte Carlo for finance, and the rigorous question of whether and where quantum beats classical (including… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-optimization.AMPO-OPT-Selectionrepro-finite-and-corruption-robust-regret-bounds-in-online-inverse-linear-optimization-traces
Agent traces
Agent sessions published from a Trackio Logbook.
repro-learning-rate-annealing-improves-tuning-robustness-in-stochastic-optimization-traces
Agent traces
Agent sessions published from a Trackio Logbook.
repro-flat-minima-and-generalization-insights-from-stochastic-convex-optimization-traces
Agent traces
Agent sessions published from a Trackio Logbook.
CARMO-UltraFeedback-BinarizedPortfolio-OptimizationF1-driver-car-setup-coupling-optimization-recommendations-v0.1What this dataset tests
Whether a system can propose setup adjustmentsthat increase driver-car coupling resonance.
Focus
Setup candidatespredicted resonance gainstability trade-offcondition sensitivitypersonalized setup profile
Required outputs
setup adjustment candidates
predicted resonance gain
stability trade-off index
track condition sensitivity
personalized setup profile
All indices0 to 1
Higher gainmeans larger coupling improvement.
Constraints
Setup recommendations only.Do… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/F1-driver-car-setup-coupling-optimization-recommendations-v0.1.chatgpt_optimization_exampleThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/chatgpt_optimization_example.data-mlops-infra-predictive-optimization-finalembedding-ie-optimizationclassification-ie-optimizationfactory-optimization
CatQualia factory operations corpus — optimisation, upgrades and decision records
1,094 rows · 1,323,000 bytes · JSON Lines.
What this is
Operational material from a self-improving research factory: what was optimised, what tools were upgraded, which decisions the operator made, and the mechanics of the firing cycle that schedules the work.
Provenance
This group merges 8 source corpora. Every row carries a _source_dataset field
naming the file it… See the full description on the dataset page: https://huggingface.co/datasets/CatQualia/factory-optimization.nigerian_energy_and_utilities_ai_grid_optimization
Nigerian Energy & Utilities – AI Grid Optimization | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: parquet - Sector: energy - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/nigerian_energy_and_utilities_ai_grid_optimization.vision-embedding-ie-optimizationAMPO-Coreset-selectionlaguna-xs-ultrachat-responses
