datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
speculators-ci-datasets
speculator-tutorial
Raw vs. on-policy regenerated conversation data for training speculative-decoding
drafters (EAGLE-3 / DFlash / DSpark style), with the original source data kept alongside
so you can see exactly what regeneration changes and why it matters.
Prompts come from UltraChat-200k. The verifier / teacher model is Qwen/Qwen3-8B.
Why regenerate at all?
A speculative-decoding drafter is trained to predict what the verifier would say next.
If you train it… See the full description on the dataset page: https://huggingface.co/datasets/inference-optimization/speculators-ci-datasets.Alexandria_geometry_optimization_paths_PBE_2D
Cite this dataset Schmidt, J., Hoffmann, N., Wang, H., Borlido, P., Carriço, P. J. M. A., Cerqueira, T. F. T., Botti, S., and Marques, M. A. L. Alexandria geometry optimization paths PBE 2D. ColabFit, 2025. https://doi.org/10.60732/8781419f
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_6pieq95jrqpn_0
Visit the ColabFit… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/Alexandria_geometry_optimization_paths_PBE_2D.Alexandria_geometry_optimization_paths_PBE_3D
Cite this dataset Schmidt, J., Hoffmann, N., Wang, H., Borlido, P., Carriço, P. J. M. A., Cerqueira, T. F. T., Botti, S., and Marques, M. A. L. Alexandria geometry optimization paths PBE 3D. ColabFit, 2024. https://doi.org/10.60732/c88da7df
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_s6gf4z2hcjqy_0
Visit the ColabFit… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/Alexandria_geometry_optimization_paths_PBE_3D.QP-Benchmark
QP-Benchmark
Benchmark instances for Warm-IP,
an open-source solver for large-scale convex quadratic programs (QPs),
from the paper Warm-IP: A Path-Following ADMM Warm Start for
Interior-Point Quadratic Programming (Aslani, Tefagh, Jhanwar,
Zarepisheh; preprint link to be added). Every instance is a convex QP
minimize 0.5 x'Qx + q'x + c
subject to constraints (one-sided or two-sided; see below)
stored as an HDF5 (.h5) file, one folder per family:
MPC_data/ 64… See the full description on the dataset page: https://huggingface.co/datasets/Radiotherapy-Optimization/QP-Benchmark.tasktrove-hparam-optimization
TaskTrove hyperparameter optimization artifacts
This dataset is the campaign-level artifact archive for the TaskTrove hyperparameter and backend experiment. It
contains experiment configurations, reward and timing curves, evaluation records, checkpoint-cleanup records, the
A0 same-source base evaluation workspace, and the generated analysis website.
The experiment report, policy, and tracker remain in the parent experiment directory. See
PUBLISHED_ARTIFACTS.md for the separately… See the full description on the dataset page: https://huggingface.co/datasets/penfever/tasktrove-hparam-optimization.Qwen3-8B-Regenerated-Collectionexperimental-optimizationmanifest-digital-identity-optimization
Manifest of Digital Identity Optimization (DIO) & Ontology of Digital Identity (ODI) — Hugging Face Distribution Layer
Version / Verze: 1.0.3 (Hugging Face Distribution Layer)
Author / Autor: Daniel Beránek
Date of public articulation / Datum veřejné artikulace: 2026-07-26
Primary public node / Primární veřejný uzel: https://danielberanek.cz/manifest-dio/
Canonical archival record / Kanonický archivní záznam: Zenodo, DOI: https://doi.org/10.5281/zenodo.21610934
License /… See the full description on the dataset page: https://huggingface.co/datasets/danielberanek/manifest-digital-identity-optimization.Alexandria_geometry_optimization_paths_PBE_1D
Cite this dataset Schmidt, J., Hoffmann, N., Wang, H., Borlido, P., Carriço, P. J. M. A., Cerqueira, T. F. T., Botti, S., and Marques, M. A. L. Alexandria geometry optimization paths PBE 1D. ColabFit, 2025. https://doi.org/10.60732/12246d46
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_xnio123pebli_0
Visit the ColabFit… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/Alexandria_geometry_optimization_paths_PBE_1D.4c_optimizationThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "koch_follower",
"total_episodes": 25,
"total_frames": 8494,
"total_tasks": 1,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:25"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ethanCSL/4c_optimization.speculators_benchmarks_tool_calldflash-code-multilingual-teacher-responses-qwen235b
Code + Multilingual Teacher Responses (Qwen3-235B-A22B-Instruct-2507)
This repo now contains 302,800 total samples across the main blended
data.jsonl / .parquet file plus a second Nemotron-only file
(nemotron_code_teacher_responses.jsonl / .parquet). All responses were
generated by Qwen3-235B-A22B-Instruct-2507 in non-thinking mode
(enable_thinking=false) to match downstream speculator training and eval.
Built in two batches: an initial 59,506-row batch (50K code + 9.5K… See the full description on the dataset page: https://huggingface.co/datasets/inference-optimization/dflash-code-multilingual-teacher-responses-qwen235b.SWE-bench_Multilingualcloudspotting_mars_optimizationsimready-isaac-optimization-test-assets
SimReady Isaac Optimization Test Assets
This dataset contains four unique real-source OpenUSD packages that each
intentionally fail one requirement in Isaac-Optimization@1.0.0. Together they
provide an end-to-end Retrieve → Derive → validate proof for all three
optimization features.
Package
Real source
Target
Observable derived change
License
obs_joystick_a01
Alexandre Abreu's published SimReady Foundations joystick sample
FET_102_ISAAC@0.1.0 / OPT.001
Active rigid… See the full description on the dataset page: https://huggingface.co/datasets/caslan-nv/simready-isaac-optimization-test-assets.Qwen3.5-0.8B-responsessome-molecular-optimization-experiment
OpenAI4S Molecular Optimization scenarios-v3 交付包
本目录是五个 admet_molecular_optimization 场景的重构交付版本。目录布局对齐 OpenAI4S-Eval scenarios-v3:公共 workflow 与评测规范位于顶层,Agent queries、ground-truth repositories 和两组 OpenAI4S worker codebases 分开保存。
所有内容均为真实文件,不使用符号链接。../final/ 是只读迁移来源,不属于本交付包。
场景映射
ID
Scenario
Query
GT
Doubao
GPT
scenario-1
Scaffold and Warhead Preserving Optimization
queries/scenario-1.md
GT-Codebase/scenario-1/
OpenAI4S-Doubao-Codebase/scenario-1/… See the full description on the dataset page: https://huggingface.co/datasets/Riiiiiiin/some-molecular-optimization-experiment.repro-optimal-regret-for-policy-optimization-in-contextual-bandits-traces
Agent traces
Agent sessions published from a Trackio Logbook.
Qwen3.5-4B-responsesrepro-off-policy-learning-in-large-action-spaces-optimization-matters-more-than-estimation
Reproduction: Off-Policy Learning in Large Action Spaces - Optimization Matters More Than Estimation
Paper Information
Title: Off-Policy Learning in Large Action Spaces: Optimization Matters More Than Estimation
OpenReview ID: srIStBTJiu
Conference: ICML 2026
Task: Compare optimization landscapes of IPS vs PWLL for off-policy policy learning
Reproduction Summary
This reproduction evaluates the paper's core thesis: optimization landscape (not… See the full description on the dataset page: https://huggingface.co/datasets/sabaridsnfuji/repro-off-policy-learning-in-large-action-spaces-optimization-matters-more-than-estimation.Meta_Plan_Optimization
MPO Datasets
This folder contains the datasets for the MPO experiments.
Paper: https://hf.co/papers/2503.02682
Code: https://github.com/WeiminXiong/MPO
File Structure
alfworld_metaplan_preference_pairs.json: includes comparison data for the DPO optimization phase of the ALFWorld meta planner.
sciworld_metaplan_preference_pairs.json: includes comparison data for the DPO optimization phase of the SciWorld meta planner.
alfworld_metaplan_sft.json: includes the metaplan data… See the full description on the dataset page: https://huggingface.co/datasets/xwm/Meta_Plan_Optimization.optimization-os-instance-dataset
Optimization OS — Instance Dataset
Synthetic optimization instances for scheduling, routing, assignment, inventory, facility location, and packing.
Version: 1.0.0Instances: 36Problem types: scheduling, routing, assignment, inventory, facility_location, packing
lm-optimization-data
LM optimization: reasoning documents
Generated reasoning documents from https://github.com/LamShiuChing/LM-optimization. The datasets are recipes
(src/tasks.py generates them from a seed); each <version>.jsonl here is a 100k-document sample of one
version, one JSON object per line with task, prompt, cot (the step trace) and answer. A document
trains as prompt|cot=answer. The format is described in docs/language.md and docs/dataset.md of the repo.
Longbench_Samples_SpecdecQwen3-8b-sharegpt-5kCARMO-UltraFeedbackemgena_r1_compiler_ast_optimization_reasoner_mcp_teaser
🧠 Code-Reasoning - AST Bytecode Optimization & Inlining Reasoner (Evaluation Teaser)
⚡ Official Free Evaluation Teaser (50 Verified Scenarios + Executable MCP Server)🏆 Get the Full Production Package & Commercial EULA on Gumroad:👉 Purchase Full Package on Gumroad🏷️ Use coupon code LAUNCH20 for €20 off at checkout!
🌟 Domain Overview & Reasoning Features
Bytecode complexity analysis, recursive call inlining tradeoffs, and loop invariant vectorization proving… See the full description on the dataset page: https://huggingface.co/datasets/emgena/emgena_r1_compiler_ast_optimization_reasoner_mcp_teaser.nl-optimization-instantiation-metrics
Natural-Language Optimization Instantiation Metrics (v1.2)
This dataset is a text-free metrics release for a retrieval-assisted natural-language optimization instantiation pipeline. It contains per-example and aggregate evaluation outcomes for a frozen pipeline evaluated on the NLP4LP benchmark, the OptMath external validation domain (added in v1.1), and, as of v1.2, a 13-method schema-retrieval/grounding diagnostic suite (per-query breakdowns, bottleneck taxonomy, threshold… See the full description on the dataset page: https://huggingface.co/datasets/SoroushVahidi/nl-optimization-instantiation-metrics.manufacturing-cost-optimization-2026Q3
Manufacturing Quarterly Cost Optimization Dataset (2026Q3)
Unified quarterly cost-analysis dataset for the manufacturing group, merged from
the China / Japan / India factory datasets hosted on Hugging Face.
Contents
11,100 records (>= 10,000) covering 9 plants across 3 regions.
Source datasets:
toolathon123/manufacturing-cn-energy-2026Q3 — China energy & raw material (4,200 rows)
toolathon123/manufacturing-jp-maintenance-2026Q3 — Japan maintenance & downtime (3… See the full description on the dataset page: https://huggingface.co/datasets/toolathon123/manufacturing-cost-optimization-2026Q3.Enhanced-Group_Relative_Policy_Optimization
