Team Ai
Datasetpublic

ameau01/synthesized-cloud-optimization-recommendations

Synthesized Cloud-Optimization Recommendations 18 scenarios that pair cloud telemetry with a hand-crafted optimization recommendation. Use them to train models or to evaluate AI agents. Summary Each scenario has multi-tier telemetry, a Terraform file describing the deployed infrastructure, and a gold-standard recommendation. The dataset is built around a simple input-output mapping. The input is telemetry plus the infrastructure. The output is an optimization… See the full description on the dataset page: https://huggingface.co/datasets/ameau01/synthesized-cloud-optimization-recommendations.

sourceHugging Facemitupdated 4mo agoView on Hugging Face
0likes52downloads
Dataset Card

Synthesized Cloud-Optimization Recommendations

18 scenarios that pair cloud telemetry with a hand-crafted optimization recommendation. Use them to train models or to evaluate AI agents.

Summary

Each scenario has multi-tier telemetry, a Terraform file describing the deployed infrastructure, and a gold-standard recommendation.

The dataset is built around a simple input-output mapping. The input is telemetry plus the infrastructure. The output is an optimization recommendation that says what to change and what the impact will be.

The dataset is synthesized. Telemetry was generated procedurally to match each scenario's narrative. Gold recommendations were hand-crafted and verified.

The dataset uses AWS vocabulary throughout. Instance types, service names, and field names match AWS. This makes the scenarios concrete instead of vendor-neutral.

Folder layout

README.md                                # this file
LICENSE                                  # MIT
EVAL.md                                  # what eval.py checks
eval.py                                  # Floor sanity check (smoke test)
sample_predictions.json                  # worked example of submission shape
scenarios_summary.jsonl                  # one row per scenario (viewer table)
scenarios/
  01/
    metadata.json                        # scenario summary + fixtures
    main.tf                              # Terraform for the infra
    compute_telemetry.json               # CPU, memory, latency
    database_telemetry.json              # query rates, pool stats
    cache_telemetry.json                 # hit rate, eviction
    network_telemetry.json               # bandwidth, packet loss
    correlation_evidence.json            # cross-tier correlations
    handcrafted_recommendation.json      # the gold answer
  02/
    ...

Each scenario covers a different optimization situation. Some are single-tier (only compute is wrong). Some span tiers (database problem that surfaces in compute). Some are no-action cases. Two are diagnostic deferral cases. One asks for an SLA review instead of an infra change.

The summary table (scenarios_summary.jsonl)

The Hugging Face Dataset Viewer renders scenarios_summary.jsonl as a browsable table. Each row is one scenario and includes the headline fields from that scenario's metadata and gold recommendation.

The summary is for discovery only. The full inputs (telemetry, Terraform, correlation evidence) live in scenarios/NN/. Always train or evaluate on the full files, not on the summary.

Columns in the summary table:

ColumnSource
scenario_idfolder name
scenario_namemetadata.scenario_name
scenario_typemetadata.scenario_type
what_this_demonstratesmetadata.narrative.whatthisdemonstrates
finding_typegold.finding_type
primary_tiergold.primary_tier
secondary_tiergold.secondary_tier
action_categorygold.action_category
specific_changegold.specific_change
savings_monthly_usdgold.costimpact.savingsmonthly_usd
current_monthly_usdgold.costimpact.currentmonthly_usd
projected_monthly_usdgold.costimpact.projectedmonthly_usd

Some scenarios have negative savings_monthly_usd. That is expected. For those scenarios the right action increases cost to fix a performance or reliability problem (for example, adding a read replica).

Schema

Scenario inputs

Each scenarios/NN/ folder has these files.

FileWhat it is
metadata.jsonscenario name, narrative, fixtures
main.tfTerraform for the deployed infra
compute_telemetry.jsonper-window CPU, memory, latency
database_telemetry.jsonper-window DB query rate, pool, slow queries
cache_telemetry.jsonper-window hit rate, evictions
network_telemetry.jsonper-window bandwidth, packet loss
correlation_evidence.jsoncross-tier correlation pairs
handcrafted_recommendation.jsonthe gold answer

Recommendation shape

json
{
  "scenario_id": "01",
  "finding_type": "issue_found",
  "specific_change": "...",
  "primary_tier": "compute",
  "secondary_tier": null,
  "action_category": "rightsizing",
  "conclusion": { ... },
  "evidence": {
    "telemetry_observations": [ ... ],
    "infrastructure_context": [ ... ],
    "correlation_observations": [ ... ]
  },
  "reasoning": "...",
  "projected_state": { ... },
  "cost_impact": { ... },
  "risk_assessment": { ... }
}

Allowed values

  • —finding_type: issue_found, no_issue_found, diagnostic_deferral, insufficient_data
  • —primary_tier: compute, database, cache, network, deferred, or null
  • —secondary_tier: same set as primary_tier
  • —action_category: rightsizing, scaling_policy_change, query_cache_optimization, cache_capacity_adjustment, pool_sizing, replica_adjustment, load_balancer_reconfiguration, network_topology_change, sla_review, or null

The deferred tier sentinel is used in diagnostic-deferral scenarios where the agent explicitly cannot pick a tier yet (scenarios 15 and 17). insufficient_data is reserved for future scenarios where the dataset is too sparse to support any finding; no current scenario uses it.

Scenario coverage

IDTypeDescription
01single-tiercompute over-provisioned
02single-tiercompute peak windows, needs scheduled scaling
03single-tierdatabase over-provisioned
04single-tierslow queries plus exhausted pool
05single-tierALB round-robin causing uneven CPU
06no-actionall tiers healthy
07single-tiercache hit ratio degraded
08cross-tierslow DB queries cascade to compute
09cross-tierweekday bimodal peaks, needs scheduled scaling
10cross-tiernetwork latency cascades to compute
11cross-tierall three tiers over-provisioned
12mixedhealthy compute, over-provisioned database
13cross-tiercompute spike strains database
14cross-tiercompute and database both over-provisioned
15reliability99.99% SLA via over-provisioning
16mildpartial compute optimization
17deferralall tiers rise in lockstep, need more diagnosis
18mostly healthyminor compute inefficiency

How to use it

You can use this dataset two ways.

Train or fine-tune. Treat each scenario's telemetry plus metadata as input. Use the handcrafted_recommendation.json as the target output.

Evaluate AI agents. Run your agent on the scenario inputs. Compare its output to the hand-crafted recommendation in that scenario's folder.

Quick sanity check

bash
python eval.py --predictions sample_predictions.json

This runs the bundled Floor sanity check. It confirms your predictions parse, have the required fields, and use allowed category values. It does NOT score recommendation quality. See EVAL.md for what is checked.

Prediction shape

See sample_predictions.json for a worked example. Required fields per prediction: scenario_id, finding_type, specific_change, primary_tier, action_category. Optional but useful for deeper scoring: secondary_tier, reasoning, evidence, projected_state, cost_impact, risk_assessment.

How to score beyond the Floor check

The dataset ships gold answers and a Floor sanity check. It does not ship a quality scorer. Beyond the Floor check, the scoring method is up to you. Common options:

  • —Exact match on the enum fields (finding_type, primary_tier, action_category).
  • —Keyword or substring checks on specific_change.
  • —Semantic similarity on the prose fields.
  • —A custom rubric per scenario, comparing prediction fields against the matching handcrafted_recommendation.json.

Intended uses

  • —Train or fine-tune a model that maps cloud telemetry to an optimization recommendation.
  • —Evaluate AI agents on cloud-optimization reasoning.
  • —Compare single-shot vs orchestrated agent designs.

License

MIT. See LICENSE.

Citation

@misc{synthesized_cloud_optimization_recommendations_2026,
  title = {Synthesized Cloud-Optimization Recommendations},
  author = {Alexander Meau},
  year = {2026},
  version = {1.0.0}
}