Team Ai
Datasetpublic

timlawrenz/gnn-ruby-code-study

GNN Ruby Code Study Systematic study of Graph Neural Network architectures for Ruby code complexity prediction and generation. Paper: Graph Neural Networks for Ruby Code Complexity Prediction and Generation: A Systematic Architecture Study Dataset 22,452 Ruby methods parsed into AST graphs with 74-dimensional node features. Split Samples File Train 19,084 dataset/train.jsonl Validation 3,368 dataset/val.jsonl Each JSONL record contains:… See the full description on the dataset page: https://huggingface.co/datasets/timlawrenz/gnn-ruby-code-study.

sourceHugging Facemitupdated 6mo agoView on Hugging Face
0likes101downloads
README.md123 linesDownload Raw Back to root
1---2language:3  - code4license: mit5task_categories:6  - graph-ml7  - text-classification8tags:9  - code10  - ast11  - gnn12  - graph-neural-network13  - ruby14  - complexity-prediction15  - code-generation16  - negative-results17size_categories:18  - 10K<n<100K19---20 21# GNN Ruby Code Study22 23Systematic study of Graph Neural Network architectures for Ruby code complexity prediction and generation.24 25**Paper:** [Graph Neural Networks for Ruby Code Complexity Prediction and Generation: A Systematic Architecture Study](paper.md)26 27## Dataset28 29**22,452 Ruby methods** parsed into AST graphs with 74-dimensional node features.30 31| Split | Samples | File |32|-------|---------|------|33| Train | 19,084 | `dataset/train.jsonl` |34| Validation | 3,368 | `dataset/val.jsonl` |35 36Each JSONL record contains:37- `repo_name`: Source repository38- `file_path`: Original file path39- `raw_source`: Raw Ruby source code40- `complexity_score`: McCabe cyclomatic complexity41- `ast_json`: Full AST as nested JSON (node types + literal values)42- `id`: Unique identifier43 44### Node Features (74D)45- One-hot encoding of 73 AST node types (def, send, args, lvar, str, ...) + 1 unknown46- Types cover Ruby AST nodes; literal values (identifiers, strings, numbers) map to unknown47 48## Key Findings49 501. **5-layer GraphSAGE** achieves MAE 4.018 (R² = 0.709) for complexity prediction — 16% better than 3-layer baseline (9.9σ significant)512. **GNN autoencoders produce 0% valid Ruby** across all 15+ tested configurations523. **The literal value bottleneck**: Teacher-forced GIN achieves 81% node type accuracy and 99.5% type diversity, but 0% syntax validity because 47% of AST elements are literals with no learnable representation534. **Chain decoders collapse**: 93% of predictions default to UNKNOWN without structural supervision545. **Total cost: ~$4.32** across 51 GPU experiments on Vast.ai RTX 4090 + local RTX 2070 SUPER55 56## Repository Structure57 58```59├── paper.md                           # Full research paper60├── dataset/61│   ├── train.jsonl                    # 19,084 Ruby methods (37 MB)62│   └── val.jsonl                      # 3,368 Ruby methods (6.5 MB)63├── models/64│   ├── encoder_sage_5layer.pt         # Pre-trained SAGE encoder65│   └── decoders/                      # Trained decoder checkpoints66│       ├── tf-gin-256-deep.pt         # Best: teacher-forced GIN, 5 layers67│       ├── tf-gin-{128,256,512}.pt    # Dimension ablation68│       └── chain-gin-256.pt           # Control (no structural supervision)69├── results/70│   ├── fleet_experiments.json         # All Vast.ai experiment metrics71│   ├── autonomous_research.json       # 18 baseline variance replicates72│   └── gin_deep_dive/                 # Local deep-dive analysis73│       ├── summary.json               # Ablation summary table74│       └── *_results.json             # Per-config detailed results75├── experiments/                       # Ratiocinator fleet YAML specs76├── specs/                             # Ratiocinator research YAML specs77├── src/                               # Model source code78│   ├── models.py                      # GNN architectures79│   ├── data_processing.py             # AST→graph pipeline80│   ├── loss.py                        # Loss functions81│   ├── train.py                       # Complexity prediction trainer82│   └── train_autoencoder.py           # Autoencoder trainer83└── scripts/                           # Runner and evaluation scripts84```85 86## Reproducing Results87 88```bash89# Clone the experiment branch90git clone -b experiment/ratiocinator-gnn-study https://github.com/timlawrenz/jubilant-palm-tree91cd jubilant-palm-tree92 93# Install dependencies94python -m venv .venv && source .venv/bin/activate95pip install torch torchvision torch_geometric96 97# Train complexity prediction (Track 1)98python train.py --conv_type SAGE --num_layers 5 --epochs 5099 100# Train autoencoder with teacher-forced GIN decoder (Track 4)101python train_autoencoder.py --decoder_conv_type GIN --decoder_edge_mode teacher_forced --epochs 30102 103# Run the full deep-dive ablation104python scripts/gin_deep_dive.py105```106 107## Source Code108 109- **Model code:** [jubilant-palm-tree](https://github.com/timlawrenz/jubilant-palm-tree) (branch: `experiment/ratiocinator-gnn-study`)110- **Orchestrator:** [ratiocinator](https://github.com/timlawrenz/ratiocinator)111 112## Citation113 114If you use this dataset or findings, please cite:115```116@misc{lawrenz2025gnnruby,117  title={Graph Neural Networks for Ruby Code Complexity Prediction and Generation: A Systematic Architecture Study},118  author={Tim Lawrenz},119  year={2025},120  howpublished={\url{https://huggingface.co/datasets/timlawrenz/gnn-ruby-code-study}}121}122```123