vinsblack/CodeReality
CodeReality: Evaluation Subset - Deliberately Noisy Code Dataset ⚠️ Important Limitations ⚠️ Not Enterprise-Ready: This dataset is deliberately noisy and designed for research only. Contains mixed/unknown licenses, possible secrets, potential security vulnerabilities, duplicate code, and experimental repositories. Requires substantial preprocessing for production use. Use at your own risk - this is a research dataset for robustness testing and data curation… See the full description on the dataset page: https://huggingface.co/datasets/vinsblack/CodeReality.
148
1---2dataset_name: "CodeReality-EvalSubset"3pretty_name: "CodeReality: Evaluation Subset - Deliberately Noisy Code Dataset"4tags:5 - code6 - software-engineering7 - robustness8 - noisy-dataset9 - evaluation-subset10 - research-dataset11 - code-understanding12size_categories:13 - 10GB<n<100GB14task_categories:15 - text-generation16 - text-classification17 - text-retrieval18 - fill-mask19 - other20language:21 - en22 - code23license: other24configs:25 - config_name: default26 data_files: "data_csv/*.csv"27---28 29# CodeReality: Evaluation Subset - Deliberately Noisy Code Dataset30 3132333435 36## ⚠️ Important Limitations37 38> **⚠️ Not Enterprise-Ready**: This dataset is deliberately noisy and designed for research only. Contains mixed/unknown licenses, possible secrets, potential security vulnerabilities, duplicate code, and experimental repositories. **Requires substantial preprocessing for production use.**39>40> **Use at your own risk** - this is a research dataset for robustness testing and data curation method development.41 42## Overview43 44**CodeReality Evaluation Subset** is a curated research subset extracted from the complete CodeReality dataset (3.05TB, 397,475 repositories). This subset contains **2,049 repositories** in **19GB** of data, specifically selected for standardized evaluation and benchmarking of code understanding models on deliberately noisy data.45For complete Dataset 3tb, please contact me at vincenzo.gallo77@hotmail.com46 47### Key Features48- ✅ **Curated Selection**: Research value scoring with diversity sampling from 397,475 repositories49- ✅ **Research Grade**: Comprehensive analysis with transparent methodology50- ✅ **Deliberately Noisy**: Includes duplicates, incomplete code, and experimental projects51- ✅ **Rich Metadata**: Enhanced Blueprint metadata with cross-domain classification52- ✅ **Professional Grade**: 63.7-hour comprehensive analysis with open source tools53 54## Quick Start55 56### Dataset Structure57```58codereality-1t/59├── data_csv/ # Evaluation subset data (CSV format, 2,387 repositories)60│ ├── codereality_unified.csv # Main dataset file with unified schema61│ └── metadata.json # Dataset metadata and column information62├── analysis/ # Analysis results and tools63│ ├── dataset_index.json # File index and metadata64│ └── metrics.json # Analysis results65├── docs/ # Documentation66│ ├── DATASET_CARD.md # Comprehensive dataset card67│ └── LICENSE.md # Licensing information68├── benchmarks/ # Benchmarking scripts and frameworks69├── results/ # Evaluation results and metrics70├── Notebook/ # Analysis notebooks and visualizations71├── eval_metadata.json # Evaluation metadata and statistics72└── eval_subset_stats.json # Statistical analysis of the subset73```74 75### Loading the Dataset76 77## 📊 **Unified CSV Format**78 79**This dataset has been converted to CSV format with a unified schema** to ensure compatibility with Hugging Face's dataset viewer and eliminate schema inconsistencies that were present in the original JSONL format.80 81### **How to Use This Dataset**82 83**Option 1: Standard Hugging Face Datasets (Recommended)**84```python85from datasets import load_dataset86 87# Load the complete dataset88dataset = load_dataset("vinsblack/CodeReality")89 90# Access the data91print(f"Total samples: {len(dataset['train'])}")92print(f"Columns: {dataset['train'].column_names}")93 94# Sample record95sample = dataset['train'][0]96print(f"Repository: {sample['repo_name']}")97print(f"Language: {sample['primary_language']}")98print(f"Quality Score: {sample['quality_score']}")99```100 101**Option 2: Direct CSV Access**102```python103import pandas as pd104from huggingface_hub import snapshot_download105 106# Download the dataset107repo_path = snapshot_download(repo_id="vinsblack/CodeReality", repo_type="dataset")108 109# Load CSV files110import glob111csv_files = glob.glob(f"{repo_path}/data_csv/*.csv")112df = pd.concat([pd.read_csv(f) for f in csv_files], ignore_index=True)113 114print(f"Total records: {len(df)}")115print(f"Columns: {list(df.columns)}")116```117 118**Option 3: Metadata and Analysis**119```python120# Load evaluation subset metadata121with open('eval_metadata.json', 'r') as f:122 metadata = json.load(f)123 124print(f"Subset: {metadata['eval_subset_info']['name']}")125print(f"Files: {metadata['subset_statistics']['total_files']}")126print(f"Repositories: {metadata['subset_statistics']['estimated_repositories']}")127print(f"Size: {metadata['subset_statistics']['total_size_gb']} GB")128 129# Access evaluation data files130data_dir = "data/" # Local evaluation subset data131for filename in os.listdir(data_dir)[:5]: # First 5 files132 file_path = os.path.join(data_dir, filename)133 with open(file_path, 'r', encoding='utf-8', errors='ignore') as f:134 for line in f:135 repo_data = json.loads(line)136 print(f"Repository: {repo_data.get('name', 'Unknown')}")137 break # Just first repo from each file138```139 140## Dataset Statistics141 142### Evaluation Subset Scale143- **Total Repositories**: 2,049 (curated from 397,475)144- **Total Files**: 323 JSONL archives145- **Total Size**: 19GB uncompressed146- **Languages Detected**: Multiple (JavaScript, Python, Java, C/C++, mixed)147- **Selection**: Research value scoring with diversity sampling148- **Source Dataset**: CodeReality complete dataset (3.05TB)149 150### Language Distribution (Top 10)151| Language | Repositories | Percentage |152|----------|-------------|------------|153| Unknown | 389,941 | 98.1% |154| Python | 4,738 | 1.2% |155| Shell | 4,505 | 1.1% |156| C | 3,969 | 1.0% |157| C++ | 3,339 | 0.8% |158| HTML | 2,487 | 0.6% |159| JavaScript | 2,394 | 0.6% |160| Go | 2,110 | 0.5% |161| Java | 2,026 | 0.5% |162| CSS | 1,655 | 0.4% |163 164### Duplicate Analysis165**Exact Duplicates**: 0% exact SHA256 duplicates detected across file-level content166**Semantic Duplicates**: ~18% estimated semantic duplicates and forks preserved by design167**Research Value**: Duplicates intentionally maintained for real-world code distribution studies168 169### License Analysis170**License Detection**: 0% detection rate (design decision for noisy dataset research)171**Unknown Licenses**: 96.4% of repositories marked as "Unknown" by design172**Research Purpose**: Preserved to test license detection systems and curation methods173 174### Security Analysis175⚠️ **Security Warning**: Dataset contains potential secrets176- Password patterns: 1,231,942 occurrences177- Token patterns: 353,266 occurrences178- Secret patterns: 71,778 occurrences179- API key patterns: 4,899 occurrences180 181## Research Applications182 183### Primary Use Cases1841. **Code LLM Robustness**: Testing model performance on noisy, real-world data1852. **Data Curation Research**: Developing automated filtering and cleaning methods1863. **License Detection**: Training and evaluating license classification systems1874. **Bug-Fix Studies**: Before/after commit analysis for automated debugging1885. **Cross-Language Analysis**: Multi-language repository understanding189 190### About This Evaluation Subset191This repository contains the **19GB evaluation subset** designed for standardized benchmarks:192- **323 files** containing **2,049 repositories**193- Research value scoring with diversity sampling194- Cross-language implementations and multi-repo analysis195- Complete build system configurations196- Enhanced metadata with commit history and issue tracking197 198**Note**: The complete 3.05TB CodeReality dataset with all 397,475 repositories is available separately. Contact vincenzo.gallo77@hotmail.com for access to the full dataset.199 200**Demonstration Benchmarks** available in `eval/benchmarks/`:201- **License Detection**: Automated license classification evaluation202- **Code Completion**: Pass@k metrics for code generation models203- **Extensible Framework**: Easy to add new evaluation tasks204 205## Benchmarks & Results206 207### 📊 **Baseline Performance**208Demonstration benchmark results available in `eval/results/`:209- [`license_detection_sample_results.json`](eval/results/license_detection_sample_results.json) - 9.8% accuracy (challenging baseline)210- [`code_completion_sample_results.json`](eval/results/code_completion_sample_results.json) - 14.2% Pass@1 (noisy data challenge)211 212### 🏃 **Quick Start Benchmarking**213```bash214cd eval/benchmarks215python3 license_detection_benchmark.py # License classification216python3 code_completion_benchmark.py # Code generation Pass@k217```218 219**Note**: These are demonstration baselines, not production-ready models. Results show expected challenges of deliberately noisy data.220 221### 📊 **Benchmarks & Results**222- **License Detection**: 9.8% accuracy baseline ([`license_detection_sample_results.json`](eval/results/license_detection_sample_results.json))223- **Code Completion**: 14.2% Pass@1, 34.6% Pass@5 ([`code_completion_sample_results.json`](eval/results/code_completion_sample_results.json))224- **Framework Scaffolds**: Bug detection and cross-language translation ready for community implementation225- **Complete Analysis**: [`benchmark_summary.csv`](eval/results/benchmark_summary.csv) - All metrics for easy comparison and research use226 227## Usage Guidelines228 229### ✅ Recommended Uses230- Academic research and education231- Robustness testing of code models232- Development of data curation methods233- License detection research234- Security pattern analysis235 236### ❌ Important Limitations237- **No Commercial Use** without individual license verification238- **Research Only**: Many repositories have unknown licensing239- **Security Risk**: Contains potential secrets and vulnerabilities240- **Deliberately Noisy**: Requires preprocessing for most applications241 242## ⚠️ Important: Dataset vs Evaluation Subset243 244**This repository contains the 19GB evaluation subset only.** Some files within this repository (such as `docs/DATASET_CARD.md`, notebooks in `Notebook/`, and analysis results) reference or describe the complete 3.05TB CodeReality dataset. This is intentional for research context and documentation completeness.245 246### What's in this repository:247- ✅ **Evaluation subset data**: 19GB, 2,049 repositories in `data/` directory248- ✅ **Analysis tools and scripts**: For working with both subset and full dataset249- ✅ **Documentation**: Describes both the subset and the complete dataset methodology250- ✅ **Benchmarks**: Ready to use with the evaluation subset251 252### Complete Dataset Access (3.05TB):253- 📧 **Contact**: vincenzo.gallo77@hotmail.com for access to the full dataset254- 📊 **Full Scale**: 397,475 repositories across 21 programming languages255- 🗂️ **Size**: 3.05TB uncompressed, 52,692 JSONL files256 257#### Who Should Use the Complete Dataset:258- 🎯 **Large-scale ML researchers** training foundation models on massive code corpora259- 🏢 **Enterprise teams** developing production code understanding systems260- 🔬 **Academic institutions** conducting comprehensive code analysis studies261- 📊 **Data scientists** performing statistical analysis on repository distributions262- 🛠️ **Tool developers** building large-scale code curation and filtering systems263 264#### Advantages of Complete Dataset vs Evaluation Subset:265| Feature | Evaluation Subset (19GB) | Complete Dataset (3.05TB) |266|---------|-------------------------|---------------------------|267| **Repositories** | 2,049 curated | 397,475 complete coverage |268| **Use Case** | Benchmarking & evaluation | Large-scale training & research |269| **Data Quality** | High (curated selection) | Mixed (deliberately noisy) |270| **Languages** | Multi-language focused | 21+ languages comprehensive |271| **Setup Time** | Immediate | Requires infrastructure planning |272| **Best For** | Model evaluation, testing | Model training, comprehensive analysis |273 274#### Choose Complete Dataset When:275- ✅ Training large language models requiring massive code corpora276- ✅ Developing data curation algorithms at scale277- ✅ Studying real-world code distribution patterns278- ✅ Building production-grade code understanding systems279- ✅ Researching cross-language programming patterns280- ✅ Creating comprehensive code quality metrics281 282#### Choose Evaluation Subset When:283- ✅ Benchmarking existing models284- ✅ Quick prototyping and testing285- ✅ Learning to work with noisy code datasets286- ✅ Limited storage or computational resources287- ✅ Focused evaluation on curated, high-value repositories288 289## Configuration Files (YAML)290 291The project includes comprehensive YAML configuration files for easy programmatic access:292 293| Configuration File | Description |294|-------------------|-------------|295| [`dataset-config.yaml`](dataset-config.yaml) | Main dataset metadata and structure |296| [`analysis-config.yaml`](analysis-config.yaml) | Analysis methodology and results |297| [`benchmarks-config.yaml`](benchmarks-config.yaml) | Benchmarking framework configuration |298 299### Using Configuration Files300 301```python302import yaml303 304# Load dataset configuration305with open('dataset-config.yaml', 'r') as f:306 dataset_config = yaml.safe_load(f)307 308print(f"Dataset: {dataset_config['dataset']['name']}")309print(f"Version: {dataset_config['dataset']['version']}")310print(f"Total repositories: {dataset_config['dataset']['metadata']['total_repositories']}")311 312# Load analysis configuration313with open('analysis-config.yaml', 'r') as f:314 analysis_config = yaml.safe_load(f)315 316print(f"Analysis time: {analysis_config['analysis']['methodology']['total_time_hours']} hours")317print(f"Coverage: {analysis_config['analysis']['methodology']['coverage_percentage']}%")318 319# Load benchmarks configuration320with open('benchmarks-config.yaml', 'r') as f:321 benchmarks_config = yaml.safe_load(f)322 323for benchmark in benchmarks_config['benchmarks']['available_benchmarks']:324 print(f"Benchmark: {benchmark}")325```326 327## Documentation328 329| Document | Description |330|----------|-------------|331| [Dataset Card](docs/DATASET_CARD.md) | Comprehensive dataset documentation |332| [License](docs/LICENSE.md) | Licensing terms and legal considerations |333| [Data README](data/README.md) | Data access and usage instructions |334 335## Verification336 337Verify dataset integrity:338```bash339# Check evaluation subset counts340python3 -c "341import json342with open('eval_metadata.json', 'r') as f:343 metadata = json.load(f)344 print(f'Files: {metadata[\"subset_statistics\"][\"total_files\"]}')345 print(f'Repositories: {metadata[\"subset_statistics\"][\"estimated_repositories\"]}')346 print(f'Size: {metadata[\"subset_statistics\"][\"total_size_gb\"]} GB')347"348 349# Expected output:350# Files: 323351# Repositories: 2049352# Size: 19.0 GB353```354 355## Citation356 357```bibtex358@misc{codereality2025,359 title={CodeReality Evaluation Subset: A Curated Research Dataset for Robust Code Understanding},360 author={Vincenzo Gallo},361 year={2025},362 note={Version 1.0.0 - Evaluation Subset (19GB from 3.05TB source)}363}364```365 366## Community Contributions367 368We welcome community contributions to improve CodeReality-1T:369 370### 🛠️ **Data Curation Scripts**371- Contribute filtering and cleaning scripts for the noisy dataset372- Share deduplication algorithms and quality improvement tools373- Submit license detection and classification improvements374 375### 📊 **New Benchmarks**376- Add evaluation tasks beyond license detection and code completion377- Contribute cross-language analysis benchmarks378- Share bug detection and security analysis evaluations379 380### 📈 **Future Versions**381- **v1.1.0**: Enhanced evaluation subset with community feedback382- **v1.2.0**: Improved license detection and filtering tools383- **v2.0.0**: Community-curated clean variant with quality filters384 385### 🤝 **How to Contribute**386**Community contributions are actively welcomed and encouraged!** Help improve the largest deliberately noisy code dataset.387 388**🎯 Priority Contribution Areas**:389- **Data Curation**: Cleaning scripts, deduplication algorithms, quality filters390- **Benchmarks**: New evaluation tasks, improved baselines, framework implementations391- **Analysis Tools**: Visualization, statistics, metadata enhancement392- **Documentation**: Usage examples, tutorials, case studies393 394**📋 Contribution Process**:3951. Clone the repository locally3962. Review existing analysis in the `analysis/` directory3973. Develop improvements or new features3984. Test your contributions thoroughly3995. Submit your improvements via standard collaboration methods400 401**💡 Join the Community**: Share your research, tools, and insights using CodeReality!402 403## Support & Access404 405### Evaluation Subset (This Repository)406- **Documentation**: See `docs/` directory for comprehensive information407- **Analysis**: Check `analysis/` directory for current research insights408- **Usage**: All benchmarks and tools work directly with the 19GB subset409 410### Complete Dataset Access (3.05TB)411- **🔗 Full Dataset Request**: Contact vincenzo.gallo77@hotmail.com412- **📋 Include in your request**:413 - Research purpose and intended use414 - Institutional affiliation (if applicable)415 - Technical requirements and storage capacity416- **⚡ Response time**: Typically within 24-48 hours417 418### General Support419- **Technical Questions**: vincenzo.gallo77@hotmail.com420- **Documentation Issues**: Check `docs/` directory first421- **Benchmark Problems**: Review `benchmarks/` and `results/` directories422 423---424 425*Dataset created using transparent research methodology with complete reproducibility. Analysis completed in 63.7 hours with 100% coverage and no sampling.*426 427 428 429 430 431 432 433 