Team Ai
Datasetpublic

AdithyaSK/data_agent_harbor_test

data_agent_harbor_test 250 deterministic data-analysis tasks for agent RL (held-out test benchmark split). Each task gives an agent a Kaggle dataset and a question; the answer is graded deterministically (exact -> numeric tolerance -> list/percent normalization -> symbolic, no LLM judge). Difficulty tiers: {'hard': 99, 'easy': 85, 'medium': 66}. Environments build from base image savatar101/env-data-agent-train:base. Format Harbor task suite: tasks/<id>/… See the full description on the dataset page: https://huggingface.co/datasets/AdithyaSK/data_agent_harbor_test.

sourceHugging Facemitupdated 1mo agoView on Hugging Face
0likes932downloads
Dataset Card

dataagentharbor_test

250 deterministic data-analysis tasks for agent RL (held-out test benchmark split). Each task gives an agent a Kaggle dataset and a question; the answer is graded deterministically (exact -> numeric tolerance -> list/percent normalization -> symbolic, no LLM judge).

Difficulty tiers: {'hard': 99, 'easy': 85, 'medium': 66}. Environments build from base image savatar101/env-data-agent-train:base.

Format

Harbor task suite: tasks/<id>/ (task.toml, instruction.md, environment/, tests/), registry.json, manifest.parquet.