Team Ai
Datasetpublic

vinsblack/CodeReality

CodeReality: Evaluation Subset - Deliberately Noisy Code Dataset ⚠️ Important Limitations ⚠️ Not Enterprise-Ready: This dataset is deliberately noisy and designed for research only. Contains mixed/unknown licenses, possible secrets, potential security vulnerabilities, duplicate code, and experimental repositories. Requires substantial preprocessing for production use. Use at your own risk - this is a research dataset for robustness testing and data curation… See the full description on the dataset page: https://huggingface.co/datasets/vinsblack/CodeReality.

sourceHugging Faceotherupdated 1y agoView on Hugging Face
1likes48downloads
DATASET_CARD.md351 linesDownload Raw Back to docs
1# CodeReality-1T Dataset Card2 3## Dataset Summary4 5**CodeReality-1T** is a large-scale, deliberately noisy code repository dataset designed for robust AI research. The dataset contains **397,475 repositories** across **21 programming languages** in **3.05 TB** of uncompressed data, specifically curated to test robustness, data curation methods, and real-world code understanding.6 7- **Total Size**: 3.05 TB (uncompressed)8- **Repositories**: 397,4759- **Files**: 52,692 JSONL archives10- **Languages**: 21 detected languages11- **Status**: `deliberately_noisy: true` (research-only)12- **Version**: 1.0.013 14## Dataset Structure15 16### Data Format17- **Format**: JSONL (JSON Lines) archives18- **Repository Structure**: Each line contains complete repository metadata including:19  - Source code files with full paths20  - Git commit history and messages21  - Issue tracking data22  - Repository metadata (stars, forks, topics)23  - License information (when available)24 25### Language Distribution26Based on complete analysis of all 397,475 repositories:27 28| Language | Repositories | Percentage |29|----------|-------------|------------|30| Unknown | 389,941 | 98.1% |31| Python | 4,738 | 1.2% |32| Shell | 4,505 | 1.1% |33| C | 3,969 | 1.0% |34| C++ | 3,339 | 0.8% |35| HTML | 2,487 | 0.6% |36| JavaScript | 2,394 | 0.6% |37| Go | 2,110 | 0.5% |38| Java | 2,026 | 0.5% |39| Others | 1,966 | 0.5% |40 41### Domain Distribution42Cross-domain analysis reveals:43 44| Domain | Repositories | Cross-Domain |45|--------|-------------|--------------|46| General | 389,941 | - |47| Database | 7,534 | ✓ |48| AI/ML | 7,534 | ✓ |49| Systems | 7,534 | ✓ |50| Security | 7,534 | ✓ |51| Web | 7,429 | ✓ |52| Enterprise | 7,072 | ✓ |53| Gaming | 6,538 | ✓ |54| Mobile | 5,705 | ✓ |55| Scientific | 5,386 | ✓ |56| DevOps | 4,600 | ✓ |57 58**Cross-domain repositories**: 59,332 (14.9%)59 60## Motivation61 62Real-world code repositories are inherently messy, containing:63- Duplicate code and forked repositories64- Incomplete or experimental code snippets65- Mixed licensing conditions66- Buggy commits and partial implementations67- DevOps configurations and non-code artifacts68 69CodeReality-1T embraces this complexity as a **research laboratory** for:70 711. **Robustness Testing**: How do code LLMs perform on noisy, real-world data?722. **Data Curation Methods**: Developing better filtering and cleaning techniques733. **License Compliance**: Research into automated license detection and filtering744. **Bug-Fix Alignment**: Studying commit patterns for before/after code analysis755. **NL↔Code Tasks**: Natural language to code alignment through issues, commits, and documentation76 77## Collection Process78 79### Sources80- Public GitHub repositories81- GitLab public projects82- Open source package registries83- Developer forum code dumps84 85### Acquisition Pipeline861. **Repository Harvesting**: Systematic collection from public sources872. **Metadata Extraction**: Complete git history, issues, documentation883. **Format Standardization**: Conversion to JSONL with consistent schema894. **Indexing**: SHA256 checksums and comprehensive cataloging90 91### Filtering Strategy92**Deliberately Minimal Filtering** to preserve research value:93- ✅ **Kept**: Forks, duplicates, incomplete code, experimental projects94- ✅ **Kept**: Repositories with unknown or missing licenses95- ✅ **Kept**: Multi-language and cross-domain projects96- ❌ **Excluded**: Only explicitly "all rights reserved" repositories97 98### Quality Assurance99- **100% Coverage**: Complete analysis without sampling100- **Integrity Verification**: SHA256 checksums for all files101- **Comprehensive Indexing**: Full metadata extraction and validation102- **Reproducible Pipeline**: Open source tools only (enry, scancode-toolkit, PyDriller)103 104## Technical Characteristics105 106### File Type Distribution (Top 15)107| Extension | Files | Description |108|-----------|-------|-------------|109| .h | 34,195,463 | C/C++ headers |110| .go | 18,691,961 | Go source |111| .java | 18,109,114 | Java source |112| .c | 16,700,728 | C source |113| .py | 15,650,558 | Python source |114| .ts | 10,271,948 | TypeScript |115| .cpp | 9,768,211 | C++ source |116| .md | 7,815,310 | Markdown docs |117| .rs | 7,280,129 | Rust source |118| .rb | 6,309,814 | Ruby source |119| .json | 5,888,235 | JSON data |120| .txt | 4,627,011 | Text files |121| .rst | 4,250,204 | reStructuredText |122| .js | 4,125,928 | JavaScript |123| .scala | 3,619,096 | Scala source |124 125### Build Systems Detected126| Build System | Occurrences | Ecosystem |127|--------------|-------------|-----------|128| Makefile | 619,857 | C/C++/Universal |129| package.json | 510,769 | Node.js/npm |130| build.gradle | 430,334 | Java/Android |131| pom.xml | 136,386 | Java/Maven |132| requirements.txt | 57,793 | Python/pip |133 134### Development Patterns Analysis135Based on **49,140 commit messages** analyzed:136 137| Pattern | Count | Percentage |138|---------|-------|------------|139| Bug fixes | 21,570 | 43.9% |140| New features | 11,580 | 23.6% |141| Testing | 6,483 | 13.2% |142| Documentation | 4,695 | 9.6% |143| Improvements | 4,477 | 9.1% |144| Refactoring | 335 | 0.7% |145 146## Uses147 148### Primary Research Applications1491. **Code LLM Robustness**: Testing model performance on noisy, real-world data1502. **Data Curation Research**: Developing automated filtering and cleaning methods1513. **License Detection**: Training and evaluating license classification systems1524. **Bug-Fix Studies**: Before/after commit analysis for automated debugging1535. **Cross-Language Analysis**: Multi-language repository understanding1546. **DevOps Research**: Configuration file analysis and validation155 156### Specific Task Examples157- **Deduplication**: Identify and remove duplicate code across repositories158- **License Classification**: Automated SPDX license detection and compliance159- **Issue→Code Retrieval**: Generate code solutions from natural language descriptions160- **Commit Message Generation**: Automatic commit message creation from code diffs161- **Build System Analysis**: Configuration file validation and optimization162- **Security Scanning**: Identifying potential vulnerabilities and secrets163 164## Limitations165 166### License Coverage167- **0% License Detection Rate**: All repositories marked as "Unknown" in current release168- **Manual Review Required**: Commercial use requires individual license verification169- **Research Use Recommended**: Dataset optimized for academic and research applications170 171### Data Quality Issues172- **98.1% Unknown Language**: Large portion of repositories with undetected language173- **Deliberately Noisy**: Intentionally includes incomplete, experimental, and duplicate code174- **Exact Duplicates**: 0% exact SHA256 duplicates detected across file-level content175- **Semantic Duplicates**: ~18% estimated semantic duplicates and forks preserved by design (includes repository forks, copy-pasted code, and similar implementations)176- **Intentional Design**: Duplicates are preserved to study real-world code distribution and test deduplication algorithms177- **Security Concerns**: Contains potential API keys, passwords, and tokens (see Security Analysis)178 179### Representation Bias180- **Language Skew**: Heavy bias toward C/C++, Python, JavaScript ecosystems181- **Geographic Bias**: Primarily English-language repositories and comments182- **Temporal Bias**: Snapshot from specific time period, may not reflect current practices183 184### Scale Limitations185- **Processing Requirements**: 3.05 TB requires significant storage and computational resources186- **Filtering Needed**: Most use cases will require substantial preprocessing187- **Network Intensive**: Large download size may limit accessibility188 189## Security Analysis190 191### Detected Security Patterns192Comprehensive security scan revealed:193 194| Pattern Type | Occurrences | Risk Level |195|--------------|-------------|------------|196| Password patterns | 1,231,942 | High |197| Token patterns | 353,266 | High |198| Secret patterns | 71,778 | Medium |199| API key patterns | 4,899 | Critical |200 201### Security Recommendations202⚠️ **WARNING**: This dataset contains potential secrets and should be used for research only203- **No Production Use**: Never deploy code from this dataset without thorough security review204- **Credential Scanning**: Always scan extracted code for hardcoded credentials205- **Isolation Required**: Use in sandboxed environments only206- **Legal Compliance**: Verify licensing before any commercial application207 208## Ethical Considerations209 210### Privacy & Consent211- **Public Data Only**: All repositories were publicly available at collection time212- **No Private Information**: No deliberately collected private repositories or data213- **Takedown Policy**: DMCA and removal requests will be honored promptly214 215### Bias & Fairness216- **Representation Issues**: Dataset reflects existing biases in open source development217- **Language Barriers**: Primarily English-language codebases and documentation218- **Economic Bias**: Overrepresents well-resourced development environments219 220### Legal Compliance221- **License Uncertainty**: Many repositories lack clear licensing information222- **Commercial Risk**: Use in commercial products requires individual license verification223- **Attribution**: Original repository attribution preserved in metadata224 225## Evaluation Framework226 227### Evaluation Subset (Available)228A curated evaluation subset is now available:229- **Size**: 19.0 GB (323 files, 2,049 repositories)230- **Selection Criteria**:231  - Research value scoring with diversity sampling232  - Repositories with enhanced metadata and commit history233  - Cross-language implementations and multi-repo files234  - Complete build system configurations235- **Location**: `/eval/subset/` with comprehensive metadata236 237### Baseline Tasks & Results2381. **Code Completion**: Pass@k evaluation → [Results: 14.2% Pass@1](../eval/results/code_completion_sample_results.json)2392. **License Classification**: Automated detection → [Results: 9.8% accuracy](../eval/results/license_detection_sample_results.json)2403. **Bug Detection**: Commit history analysis → [Framework available](../eval/benchmarks/bug_detection_benchmark.py)2414. **Cross-Language Translation**: Code equivalence → [Framework available](../eval/benchmarks/cross_language_translation_benchmark.py)2425. **Complete Analysis**: [Summary CSV](../eval/results/benchmark_summary.csv) for research comparison243 244### Metrics245- **Functional Correctness**: Pass@k, CodeBLEU, execution success rate246- **Information Retrieval**: MRR, MAP, BLEU scores for search and generation247- **Classification Accuracy**: Precision, recall, F1 for license and bug detection248 249## Distribution250 251### Access Information252 253**📦 Full Dataset (3.05 TB)**:254- **Status**: Hosting in progress on Hugging Face Hub255- **Content**: Complete 397,475 repositories, 52,692 JSONL files256- **Distribution**: `codereality/codereality-1t` (pending)257- **Alternatives**: Torrent and S3 bucket options planned258 259**📋 Evaluation Subset (19.0 GB)**:260- **Status**: Available now261- **Content**: 2,049 curated repositories, 323 JSONL files262- **Location**: `/eval/subset/` directory263- **Purpose**: Research benchmarks and evaluation tasks264 265**📚 Documentation & Tools**:266- **GitHub Repository**: Complete analysis scripts and benchmarks267- **Benchmark Results**: Sample baselines and comparison data268 269### File Organization270```271codereality-1t/272├── data/273│   ├── *.jsonl              # Repository archives (52,692 files)274│   └── manifest.json        # File checksums and metadata275├── analysis/276│   ├── dataset_index.json   # Complete file index277│   ├── metrics.json         # Analysis results278│   └── language_stats.json  # Language distribution279├── docs/280│   ├── DATASET_CARD.md      # This document281│   ├── LICENSE.md           # Dataset license282│   └── USAGE_EXAMPLES.md    # Code examples283└── eval/284    ├── subset/              # Evaluation subset (15.1GB, available)285    └── benchmarks/          # Evaluation scripts286```287 288### Checksums & Integrity289- **Hash Algorithm**: SHA256290- **Manifest File**: Complete checksums for all 52,692 JSONL files291- **Verification**: `sha256sum -c manifest.json`292 293## Maintenance & Support294 295### Contact Information296- **Primary Maintainer**: Vincenzo Gallo (vincenzo.gallo77@hotmail.com)297- **Issue Tracker**: https://github.com/vinsguru/codereality-1t/issues298- **Repository**: https://github.com/vinsguru/codereality-1t299 300### Update Policy301- **Version 1.0.0**: Initial deliberately noisy release302- **Future Versions**: May include cleaned/curated variants303- **Community Contributions**: Cleaning scripts, evaluation tasks, and analysis tools welcome304 305### Contribution Guidelines3061. **Bug Reports**: Use GitHub issues for data quality problems3072. **Enhancement Requests**: Suggest improvements via pull requests3083. **Research Papers**: Share research using this dataset for community benefit3094. **Derived Datasets**: Coordinate to avoid duplication and ensure proper attribution310 311## Version History312 313### v1.0.0 (Current)314- **Release Date**: September 2025315- **Content**: Complete 3.05 TB deliberately noisy dataset316- **Analysis**: Full BigCode-compliant metrics on all 397,475 repositories317- **Status**: Research-ready with comprehensive documentation318 319### Community-Driven Roadmap320CodeReality-1T is a **living dataset** that evolves with community contributions:321 322- **v1.1.0 (Q1 2025)**: Enhanced evaluation subset with community feedback, improved benchmarks, and additional task frameworks323- **v1.2.0 (Q2 2025)**: License detection improvements, deduplication analysis tools, semantic duplicate estimation, and community filtering scripts324- **v2.0.0 (Q3 2025)**: Community-curated clean variant with quality filters, improved metadata, and production-ready subset325 326**Community contributions actively encouraged**: cleaning scripts, new benchmarks, evaluation tasks, data curation improvements, and quality assessment tools.327 328## Citation329 330```bibtex331@misc{codereality2025,332  title={CodeReality-1T: A Large-Scale Deliberately Noisy Dataset for Robust Code Understanding},333  author={Vincenzo Gallo},334  year={2025},335  publisher={Hugging Face},336  howpublished={\\url{https://huggingface.co/vinsblack}},337  note={Version 1.0.0}338}339```340 341## License342 343This dataset is released under [License Terms] with the following considerations:344- **Research Use**: Freely available for academic and research purposes345- **Commercial Use**: Requires individual license verification for each repository346- **Attribution**: Please cite this dataset card and preserve original repository attribution347- **Liability**: Provided as-is with no warranties regarding licensing or content accuracy348 349---350 351*Dataset Card generated automatically from comprehensive analysis of all 397,475 repositories using BigCode-compliant methodology. Analysis completed in 63.7 hours with 100% coverage and no sampling.*