Arry7868/Kaggle_Simulation_Environment
KaggleSimEnv v3
Production-grade OpenEnv RL environment simulating Kaggle competitions with hierarchical action categories, causal dataset properties, failure-mode traps, contextual strategy scoring, and 50+ advanced strategies.
🤗 Live on Hugging Face Spaces: https://huggingface.co/spaces/aadi-gupta/kaggle-sim-env 📓 Training Notebook (Colab): train_grpo.ipynb — GRPO training with Unsloth + TRL on a free T4 GPU 📝 Writeup: (link your HF blog post or YouTube video here once published)
Training Results
We trained a Qwen2.5-0.5B-Instruct agent using GRPO (Group Relative Policy Optimisation) via TRL + Unsloth. The model learns to generate action plans that score higher against the env compared to a random agent.
Episode Reward Curve
X-axis: episode number. Y-axis: final grade score (0–1). Smoothed with a rolling window of 8.
Loss Curve (score gap to optimal)
Lower is better. Expert baseline (blue) consistently closes the gap faster than the random agent (red).
Per-task Score: Random vs Expert Baseline
Expert baseline outperforms random agent across all 5 tasks (30-episode mean scores).
Quantitative Results (30-episode run)
To reproduce plots:python generate_training_plots_stub.py --episodes 30To reproduce full GRPO training: opentrain_grpo.ipynbin Google Colab (T4 GPU, ~25 min).
Quick Start
pip install -r requirements.txt
uvicorn server.app:app --host 0.0.0.0 --port 7860 --reloadBaseline agent
export OPENAI_API_KEY=sk-...
python -m baseline.run_baseline --mode localArchitecture
openenvHackathon/
├── kaggle_sim_env/
│ ├── models.py # Hierarchical categories, DatasetProperties, FailureMode
│ ├── environment.py # Causal logic, trap detection, mitigation tracking
│ ├── tasks.py # 5 tasks with properties, traps, context relevance
│ ├── grader.py # 4-axis grading (perf + strategy + combo + trap)
│ ├── leaderboard.py # Ghost competitor leaderboard
│ ├── hints.py # Per-task hint dispensing
│ └── rewards.py # 9-component dense reward
├── api/server.py # FastAPI (8 endpoints)
├── baseline/run_baseline.py # Structured phase-based agent
├── openenv.yaml / Dockerfile / requirements.txtHierarchical Action Space
Actions use category to reduce search space:
{
"action_type": "feature_engineering",
"parameters": {
"category": "distribution",
"technique": "log_transform"
}
}Plus: pseudo_label (iterations), inspect_top_solution, submit
Causal Dataset Properties
Each task has ground-truth properties that drive causal reward logic:
DatasetProperties(
has_shift=True, # Actions addressing shift are rewarded
has_leakage=True, # Cleaning leaky features is critical
has_noise_features=True, # Interaction terms on noise amplify it
has_missing_data=True, # Reconstruction strategies get bonus
has_imbalance=True, # Scale_pos_weight becomes relevant
has_images=False, # Image augmentation is irrelevant → penalty
needs_physics=False, # Physics loss is irrelevant → penalty
)Actions are scored based on whether they match the dataset:
if dataset.has_shift and action == "adversarial_validation":
reward += context_bonus # Relevant!
elif not dataset.has_images and action == "geometric_augmentation":
reward += irrelevant_penalty # Wrong domain!Failure-Mode Traps
The environment contains traps that punish common mistakes:
Traps can be mitigated by taking the correct action first. The environment tracks mitigations.
Grading (4 Axes)
final = 0.40×performance + 0.25×strategy + 0.20×combo + 0.15×trap_avoidanceReward Function (9 Components)
Tasks (5)
Baseline Agent
Structured multi-phase approach:
- Inspect hints (1-2)
- Diagnose dataset properties
- Clean if needed
- CV appropriate for domain
- Features domain-relevant only
- Train right model family
- Tune imbalance/loss
- Ensemble (1-2 techniques)
- Submit
Keeps actions to 8-15 total. Uses hints to inform decisions.
Docker
docker build -t kaggle-sim-env .
docker run -p 7860:7860 kaggle-sim-envLicense
MIT
