Team Ai
Apppublic

Arry7868/Kaggle_Simulation_Environment

sourceHugging Faceupdated 6mo agoView on Hugging Face
0likes
App README

KaggleSimEnv v3

Production-grade OpenEnv RL environment simulating Kaggle competitions with hierarchical action categories, causal dataset properties, failure-mode traps, contextual strategy scoring, and 50+ advanced strategies.

🤗 Live on Hugging Face Spaces: https://huggingface.co/spaces/aadi-gupta/kaggle-sim-env 📓 Training Notebook (Colab): train_grpo.ipynb — GRPO training with Unsloth + TRL on a free T4 GPU 📝 Writeup: (link your HF blog post or YouTube video here once published)

Training Results

We trained a Qwen2.5-0.5B-Instruct agent using GRPO (Group Relative Policy Optimisation) via TRL + Unsloth. The model learns to generate action plans that score higher against the env compared to a random agent.

Episode Reward Curve

[image] X-axis: episode number. Y-axis: final grade score (0–1). Smoothed with a rolling window of 8.

Loss Curve (score gap to optimal)

[image] Lower is better. Expert baseline (blue) consistently closes the gap faster than the random agent (red).

Per-task Score: Random vs Expert Baseline

[image] Expert baseline outperforms random agent across all 5 tasks (30-episode mean scores).

Quantitative Results (30-episode run)

TaskRandom agentExpert baselineDelta
easy_churn0.411.00+0.59
medium_fraud0.260.78+0.52
hardleakynoisy0.130.64+0.51
image_quality0.050.48+0.43
trajectory_pred0.060.48+0.42
Mean0.180.68+0.50
To reproduce plots: python generate_training_plots_stub.py --episodes 30 To reproduce full GRPO training: open train_grpo.ipynb in Google Colab (T4 GPU, ~25 min).

Quick Start

bash
pip install -r requirements.txt
uvicorn server.app:app --host 0.0.0.0 --port 7860 --reload

Baseline agent

bash
export OPENAI_API_KEY=sk-...
python -m baseline.run_baseline --mode local

Architecture

openenvHackathon/
├── kaggle_sim_env/
│   ├── models.py         # Hierarchical categories, DatasetProperties, FailureMode
│   ├── environment.py    # Causal logic, trap detection, mitigation tracking
│   ├── tasks.py          # 5 tasks with properties, traps, context relevance
│   ├── grader.py         # 4-axis grading (perf + strategy + combo + trap)
│   ├── leaderboard.py    # Ghost competitor leaderboard
│   ├── hints.py          # Per-task hint dispensing
│   └── rewards.py        # 9-component dense reward
├── api/server.py         # FastAPI (8 endpoints)
├── baseline/run_baseline.py  # Structured phase-based agent
├── openenv.yaml / Dockerfile / requirements.txt

Hierarchical Action Space

Actions use category to reduce search space:

json
{
  "action_type": "feature_engineering",
  "parameters": {
    "category": "distribution",
    "technique": "log_transform"
  }
}
Action TypeCategories → Techniques
set_cvstandard(kfold, repeatedkfold) · group(groupkfold, stratifiedgroupkfold) · temporal(timesplit, combinedgroup_time)
feature_engineeringdistribution(logtransform, normalize, quantilefeatures) · interaction(interactionterms, domainratios) · encoding(sincosencoding, targetencoding, spatialencoding, tfidffeatures) · spatial(relativecoordinates, distancefeatures) · signal(frequencyfeatures, multilayerfeatures, fourier_resampling)
detect_shiftdetection(adversarialvalidation, featureimportanceshift) · mitigation(removeidentifiers, domaininvariantfeatures)
train_modeltree(xgboost, lightgbm, catboost, randomforest) · linear(linear) · neural(neuralnetwork, pretrainedbackbone, temporalcnn, transformer_encoder)
handle_imbalanceweighting(scaleposweight, classweightedloss) · calibration(calibrateprobabilities, optimizethreshold) · hierarchy(hierarchicallabels, lowerthresholds_recall)
clean_dataremoval(removecorrupted, removeoutliers, removeleakyfeatures) · reconstruction(analyticalreconstruction, nannativemodel, domainaugmentation, cleansubsettraining)
augmentationgeometric(geometric, rotationinvariant, imagerectification) · color(colortransform, clahe) · noise(gaussiannoise, robustnessaugmentation) · domain(camerasimulation, temporalaugmentation, symmetryaugmentation, multiviewprocessing)
ensembleaveraging(weightedaverage, multiseedaveraging, swa) · stacking(stacking) · diversity(diversefeatures, heterogeneous)
postprocesscalibration(biascorrection, predictionshrinkage, pergroupcalibration) · domain(domainrules, physicsconstraints) · inference(tta)
tune_lossasymmetric(asymmetricloss, epsiloninsensitive) · uncertainty(gaussiannll) · multiobjective(multitask, intervalregression, quantileregression) · weighting(sampleweighted, auxiliaryphysicsloss)
regularizeweight(strongregularization, ema, dropout) · transfer(freezebackbone)

Plus: pseudo_label (iterations), inspect_top_solution, submit


Causal Dataset Properties

Each task has ground-truth properties that drive causal reward logic:

python
DatasetProperties(
    has_shift=True,         # Actions addressing shift are rewarded
    has_leakage=True,       # Cleaning leaky features is critical
    has_noise_features=True, # Interaction terms on noise amplify it
    has_missing_data=True,  # Reconstruction strategies get bonus
    has_imbalance=True,     # Scale_pos_weight becomes relevant
    has_images=False,       # Image augmentation is irrelevant → penalty
    needs_physics=False,    # Physics loss is irrelevant → penalty
)

Actions are scored based on whether they match the dataset:

if dataset.has_shift and action == "adversarial_validation":
    reward += context_bonus    # Relevant!
elif not dataset.has_images and action == "geometric_augmentation":
    reward += irrelevant_penalty  # Wrong domain!

Failure-Mode Traps

The environment contains traps that punish common mistakes:

TrapTriggerEffectMitigation
kfoldontemporal_dataUsing kfold when has_shiftCV +0.08, test -0.04Use time_split instead
ignoring_shiftTraining without addressing shifttest -0.06Detect shift first
keepingleakyfeatureTraining when has_leakagetest -0.08Clean leaky features first
targetencodingleakagetarget_encoding on shifted dataCV +0.05, test -0.06Don't use it
interactiontermson_noiseinteraction_terms when noiseCV +0.05, test -0.04Avoid on noisy data
treemodelon_imagesxgboost on image dataCV +0.04, test -0.02Use pretrained_backbone
noaugmentationon_imagesSubmit without augmentationtest -0.04Apply augmentation
rawheadingwithout_sincosSubmit without sincosencodingtest -0.03Encode angles properly

Traps can be mitigated by taking the correct action first. The environment tracks mitigations.


Grading (4 Axes)

final = 0.40×performance + 0.25×strategy + 0.20×combo + 0.15×trap_avoidance
ComponentDescription
performanceTest score vs ghost competitors
strategyContextual — penalises irrelevant strategies used
comboFraction of synergy combos activated
trap_avoidance1.0 minus fraction of traps triggered

Reward Function (9 Components)

ComponentDescription
cv_improvementΔ CV score
strategy_bonus+0.05 for expected strategy
context_bonus+0.03×relevance (positive) or -0.04×relevance (negative)
combo_bonus+0.08 per combo completed
redundancy_penalty-0.03 × repeat count
irrelevant_penalty-0.05 for actions with relevance ≤ -0.8
trap_penalty-0.08 per trap triggered
overfitting_penalty-0.5 × gap when CV-test > 0.05
submission_bonus0.5 × test_score

Tasks (5)

TaskDifficultyTrapsCombosKey Challenge
easy_churnEasy32Clean tabular, mild imbalance
medium_fraudMedium33Shift, heavy imbalance, safety-critical
hard_leaky_noisyHard44Leakage, noise, missing data, shift
image_qualityHard24Heavy-tailed, camera bias, augmentation
trajectory_predHard24Multi-agent, physics, spatial-temporal

Baseline Agent

Structured multi-phase approach:

  1. 1.Inspect hints (1-2)
  2. 2.Diagnose dataset properties
  3. 3.Clean if needed
  4. 4.CV appropriate for domain
  5. 5.Features domain-relevant only
  6. 6.Train right model family
  7. 7.Tune imbalance/loss
  8. 8.Ensemble (1-2 techniques)
  9. 9.Submit

Keeps actions to 8-15 total. Uses hints to inform decisions.


Docker

bash
docker build -t kaggle-sim-env .
docker run -p 7860:7860 kaggle-sim-env

License

MIT