vinsblack/CodeReality
CodeReality: Evaluation Subset - Deliberately Noisy Code Dataset ⚠️ Important Limitations ⚠️ Not Enterprise-Ready: This dataset is deliberately noisy and designed for research only. Contains mixed/unknown licenses, possible secrets, potential security vulnerabilities, duplicate code, and experimental repositories. Requires substantial preprocessing for production use. Use at your own risk - this is a research dataset for robustness testing and data curation… See the full description on the dataset page: https://huggingface.co/datasets/vinsblack/CodeReality.
148
1# CodeReality-1T Evaluation Benchmarks2 3This directory contains demonstration benchmark scripts for evaluating models on the CodeReality-1T dataset.4 5## Available Benchmarks6 7### 1. License Detection Benchmark8**File**: `license_detection_benchmark.py`9 10**Purpose**: Evaluates automated license classification systems on deliberately noisy data.11 12**Features**:13- Rule-based feature extraction from repository content14- Simple classification model for demonstration15- Performance metrics on license detection accuracy16- Analysis of license distribution patterns17 18**Usage**:19```bash20cd /path/to/codereality-1t/eval/benchmarks21python3 license_detection_benchmark.py22```23 24**Expected Results**:25- Low accuracy due to deliberately noisy dataset (0% license detection by design)26- Demonstrates robustness testing for license detection systems27- Outputs detailed distribution analysis28 29### 2. Code Completion Benchmark30**File**: `code_completion_benchmark.py`31 32**Purpose**: Evaluates code completion models using Pass@k metrics on real-world noisy code.33 34**Features**:35- Function extraction from Python, JavaScript, Java files36- Simple rule-based completion model for demonstration37- Pass@1, Pass@3, Pass@5 metric calculation38- Multi-language support with language-specific patterns39 40**Usage**:41```bash42cd /path/to/codereality-1t/eval/benchmarks43python3 code_completion_benchmark.py44```45 46**Expected Results**:47- Baseline performance metrics for comparison48- Language distribution analysis49- Quality scoring of completions50 51## Benchmark Characteristics52 53### Dataset Integration54- **Data Source**: Loads from `/mnt/z/CodeReality_Final/unified_dataset` by default55- **Sampling**: Uses random sampling for performance (configurable)56- **Formats**: Handles JSONL repository format from CodeReality-1T57 58### Evaluation Philosophy59- **Deliberately Noisy**: Tests model robustness on real-world messy data60- **Baseline Metrics**: Provides simple baselines for comparison (not production-ready)61- **Reproducible**: Deterministic evaluation with random seed control62- **Research Focus**: Results show challenges of noisy data, not competitive benchmarks63 64### Extensibility65- **Modular Design**: Easy to extend with new benchmarks66- **Configurable**: Sample sizes and evaluation criteria can be adjusted67- **Multiple Languages**: Framework supports cross-language evaluation68 69## Configuration70 71### Data Path Configuration72Update the `data_dir` variable in each script to point to your CodeReality-1T dataset:73 74```python75data_dir = "/path/to/your/codereality-1t/unified_dataset"76```77 78### Sample Size Adjustment79Modify sample sizes for performance tuning:80 81```python82sample_size = 500 # Adjust based on computational resources83```84 85## Output Files86 87Each benchmark generates JSON results files:88- `license_detection_results.json`89- `code_completion_results.json`90 91These contain detailed metrics and can be used for comparative analysis.92 93### Sample Results94Example results are available in `../results/`:95- `license_detection_sample_results.json` - Baseline license detection performance96- `code_completion_sample_results.json` - Baseline code completion metrics97 98These demonstrate expected performance on CodeReality-1T's deliberately noisy data.99 100## Requirements101 102### Python Dependencies103```bash104pip install json os re random typing collections105```106 107### System Requirements108- **Memory**: Minimum 4GB RAM for default sample sizes109- **Storage**: Access to CodeReality-1T dataset (3TB)110- **Compute**: Single-core sufficient for demonstration scripts111 112## Extending the Benchmarks113 114### Adding New Tasks1151. Create new Python file following naming convention: `{task}_benchmark.py`1162. Implement standard evaluation interface:117 ```python118 def load_dataset_sample(data_dir, sample_size)119 def run_benchmark(repositories)120 def print_benchmark_results(results)121 ```1223. Add task-specific evaluation metrics123 124### Supported Tasks125Current benchmarks cover:126- **License Detection**: Classification and compliance127- **Code Completion**: Generation and functional correctness128 129**Framework Scaffolds (PLANNED - Implementation Needed)**:130- [`bug_detection_benchmark.py`](bug_detection_benchmark.py) - Bug detection on commit pairs (scaffold only)131- [`cross_language_translation_benchmark.py`](cross_language_translation_benchmark.py) - Code translation across languages (scaffold only)132 133**Future Planned Benchmarks - Roadmap**:134- **v1.1.0 (Q1 2025)**: Complete bug detection and cross-language translation implementations135- **v1.2.0 (Q2 2025)**: Repository classification and domain detection benchmarks136- **v1.3.0 (Q3 2025)**: Build system analysis and validation frameworks137- **v2.0.0 (Q4 2025)**: Commit message generation and issue-to-code alignment benchmarks138 139**Community Priority**: Framework scaffolds ready for community implementation!140 141## Performance Notes142 143### Computational Complexity144- **License Detection**: O(n) where n = repository count145- **Code Completion**: O(n*m) where m = average functions per repository146 147### Optimization Tips1481. **Sampling**: Reduce sample_size for faster execution1492. **Filtering**: Pre-filter repositories by criteria1503. **Parallelization**: Use multiprocessing for large-scale evaluation1514. **Caching**: Cache extracted features for repeated runs152 153## Research Applications154 155### Model Development156- **Robustness Testing**: Test models on noisy, real-world data157- **Baseline Comparison**: Compare against simple rule-based systems158- **Cross-domain Evaluation**: Test generalization across domains159 160### Data Science Research161- **Curation Methods**: Develop better filtering techniques162- **Quality Metrics**: Research automated quality assessment163- **Bias Analysis**: Study representation bias in large datasets164 165## Citation166 167When using these benchmarks in research, please cite the CodeReality-1T dataset:168 169```bibtex170@misc{codereality2025,171 title={CodeReality-1T: A Large-Scale Deliberately Noisy Dataset for Robust Code Understanding},172 author={Vincenzo Gallo},173 year={2025},174 publisher={Hugging Face},175 howpublished={\\url{https://huggingface.co/vinsblack}},176 note={Version 1.0.0}177}178```179 180## Support181 182- **Issues**: https://github.com/vinsguru/codereality-1t/issues183- **Contact**: vincenzo.gallo77@hotmail.com184- **Documentation**: See main dataset README and documentation