datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
repro-optimal-regret-for-policy-optimization-in-contextual-bandits-traces
Agent traces
Agent sessions published from a Trackio Logbook.
repro-a-tight-theory-of-error-feedback-algorithms-in-distributed-optimization-traces
Agent traces
Agent sessions published from a Trackio Logbook.
repro-learning-rate-annealing-improves-tuning-robustness-in-stochastic-optimization-traces
Agent traces
Agent sessions published from a Trackio Logbook.
synthesized-cloud-optimization-recommendations
Synthesized Cloud-Optimization Recommendations
18 scenarios that pair cloud telemetry with a hand-crafted optimization
recommendation. Use them to train models or to evaluate AI agents.
Summary
Each scenario has multi-tier telemetry, a Terraform file describing the
deployed infrastructure, and a gold-standard recommendation.
The dataset is built around a simple input-output mapping. The input is
telemetry plus the infrastructure. The output is an optimization… See the full description on the dataset page: https://huggingface.co/datasets/ameau01/synthesized-cloud-optimization-recommendations.repro-flat-minima-and-generalization-insights-from-stochastic-convex-optimization-traces
Agent traces
Agent sessions published from a Trackio Logbook.
repro-on-regret-bounds-of-thompson-sampling-for-bayesian-optimization-traces
Agent traces
Agent sessions published from a Trackio Logbook.
repro-finite-and-corruption-robust-regret-bounds-in-online-inverse-linear-optimization-traces
Agent traces
Agent sessions published from a Trackio Logbook.
blind-spots-frontier-optimization
Technical Challenge: Symbolic Drift & Constraint Decay in Frontier LLMs
1. Executive Summary & Background
When deploying AI for core engineering, economic, and mathematical optimization tasks (such as Karush-Kuhn-Tucker conditions or Lagrangian multipliers), calculations require rigorous multi-step invariant tracking. Drawing from practical experience building optimization pipelines, we identify an underexplored capability gap in current models: Symbolic Drift and… See the full description on the dataset page: https://huggingface.co/datasets/Via20/blind-spots-frontier-optimization.laguna-xs-ultrachat-responsesfactory-optimization
CatQualia factory operations corpus — optimisation, upgrades and decision records
1,094 rows · 1,323,000 bytes · JSON Lines.
What this is
Operational material from a self-improving research factory: what was optimised, what tools were upgraded, which decisions the operator made, and the mechanics of the firing cycle that schedules the work.
Provenance
This group merges 8 source corpora. Every row carries a _source_dataset field
naming the file it… See the full description on the dataset page: https://huggingface.co/datasets/CatQualia/factory-optimization.optimization-os-benchmark-results
Optimization OS — Benchmark Results
Pre-computed benchmark runs comparing baseline, exact, scalable, and robust methods.
Runs: 144Methods: 4 per problem type (24 total)
laguna-xs-magpie-300k-responseshumanoid-performance-optimization-dataset
Humanoid Performance Optimization Dataset
Dataset for improving humanoid efficiency
based on task execution metrics.
Description
Includes execution time, energy usage,
and task success rates for optimization analysis.
File
performance_optimization_dataset.json
License
MIT
every-eval-ever-demo
