datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
experimental-optimizationspeculators_benchmarks_tool_callrepro-optimal-regret-for-policy-optimization-in-contextual-bandits-traces
Agent traces
Agent sessions published from a Trackio Logbook.
repro-off-policy-learning-in-large-action-spaces-optimization-matters-more-than-estimation
Reproduction: Off-Policy Learning in Large Action Spaces - Optimization Matters More Than Estimation
Paper Information
Title: Off-Policy Learning in Large Action Spaces: Optimization Matters More Than Estimation
OpenReview ID: srIStBTJiu
Conference: ICML 2026
Task: Compare optimization landscapes of IPS vs PWLL for off-policy policy learning
Reproduction Summary
This reproduction evaluates the paper's core thesis: optimization landscape (not… See the full description on the dataset page: https://huggingface.co/datasets/sabaridsnfuji/repro-off-policy-learning-in-large-action-spaces-optimization-matters-more-than-estimation.Qwen3.5-0.8B-responsesQwen3.5-4B-responsesLongbench_Samples_SpecdecGemma4-Responses-Nemotronrepro-a-tight-theory-of-error-feedback-algorithms-in-distributed-optimization-traces
Agent traces
Agent sessions published from a Trackio Logbook.
repro-learning-rate-annealing-improves-tuning-robustness-in-stochastic-optimization-traces
Agent traces
Agent sessions published from a Trackio Logbook.
synthesized-cloud-optimization-recommendations
Synthesized Cloud-Optimization Recommendations
18 scenarios that pair cloud telemetry with a hand-crafted optimization
recommendation. Use them to train models or to evaluate AI agents.
Summary
Each scenario has multi-tier telemetry, a Terraform file describing the
deployed infrastructure, and a gold-standard recommendation.
The dataset is built around a simple input-output mapping. The input is
telemetry plus the infrastructure. The output is an optimization… See the full description on the dataset page: https://huggingface.co/datasets/ameau01/synthesized-cloud-optimization-recommendations.repro-flat-minima-and-generalization-insights-from-stochastic-convex-optimization-traces
Agent traces
Agent sessions published from a Trackio Logbook.
Qwen3.5-9B-responseshuman_assisted_action_preference_optimizationrepro-on-regret-bounds-of-thompson-sampling-for-bayesian-optimization-traces
Agent traces
Agent sessions published from a Trackio Logbook.
repro-finite-and-corruption-robust-regret-bounds-in-online-inverse-linear-optimization-traces
Agent traces
Agent sessions published from a Trackio Logbook.
blind-spots-frontier-optimization
Technical Challenge: Symbolic Drift & Constraint Decay in Frontier LLMs
1. Executive Summary & Background
When deploying AI for core engineering, economic, and mathematical optimization tasks (such as Karush-Kuhn-Tucker conditions or Lagrangian multipliers), calculations require rigorous multi-step invariant tracking. Drawing from practical experience building optimization pipelines, we identify an underexplored capability gap in current models: Symbolic Drift and… See the full description on the dataset page: https://huggingface.co/datasets/Via20/blind-spots-frontier-optimization.DeepSeek-V4-Flash-responseslaguna-xs-ultrachat-responsespython-optimization-dpo-samplehealth-optimization-bench-sample
Health Optimization Bench (Sample)
A 30-task public sample of Health Optimization Bench,
a rubric-graded benchmark measuring how well frontier language models handle current clinical
evidence in preventive and optimization medicine. Three tasks from each of the benchmark's ten
micro benches.
The full benchmark is 977 authored tasks with 346 released across ten micro benches. On the
current leaderboard no model scores above 71 of 100 and the field spans 66 points. Rankings:… See the full description on the dataset page: https://huggingface.co/datasets/Arcophos/health-optimization-bench-sample.factory-optimization
CatQualia factory operations corpus — optimisation, upgrades and decision records
1,094 rows · 1,323,000 bytes · JSON Lines.
What this is
Operational material from a self-improving research factory: what was optimised, what tools were upgraded, which decisions the operator made, and the mechanics of the firing cycle that schedules the work.
Provenance
This group merges 8 source corpora. Every row carries a _source_dataset field
naming the file it… See the full description on the dataset page: https://huggingface.co/datasets/CatQualia/factory-optimization.sawotiQ29_crop_optimization"copyright": "© 2019 DARJYO (Pty) Ltd. All rights reserved.",
"license": "DARJYO License v1.0",
"description": "This dataset contains crop optimization information for various crops, including cabbages, green peppers, jam tomatoes, and marigold flowers. It is provided under the DARJYO License v1.0 for non-commercial research use only.",
"version": "1.0",
"author": "Darshani Persadh, Research, Innovation & Development, DARJYO",
"created": "2019-12-1",
"updated": "2023-12-19",
"license_details":… See the full description on the dataset page: https://huggingface.co/datasets/DARJYO/sawotiQ29_crop_optimization.ctest-subset-Qwen3.5-397B-A17B-FP8-dynamic-speculator-datasetpython-optimization-dpo-samplesynthetic-code-optimization-1synthetic-code-optimization-1 is a synthetic dataset with a total of ~1136 Question and Answer pairs.
This dataset was generated using the following models:
ChatGPT:
Whatever is hosted on their website
Claude:
Fable 5
Deepseek:
Deepseek "Instant"
Deepseek "Expert"
Gemini:
3.1 Flash Lite
3.5 Flash
3.1 Pro
Grok:
Fast
Mistral:
Thinking enabled
Qwen 3.7 Plus:
Thinking enabled
GLM 5.2:
Thinking "high"
Perplexity:
Whatever is on their website
This dataset follows the following… See the full description on the dataset page: https://huggingface.co/datasets/takenusername32/synthetic-code-optimization-1.repro-a-tight-theory-of-error-feedback-algorithms-in-distributed-optimization
A Tight Theory of Error Feedback Algorithms in Distributed Optimization
Reproduction of ICML 2026 paper (OpenReview: dyRD6lBH8K)
Tags
trackio
trackio-logbook
open-experiment
icml2026-repro
paper-dyRD6lBH8K
laguna-xs-magpie-300k-responseslaguna-xs-ultrachat-conversationsfinal-ctest-Qwen3-8B-speculator-dataset
