datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SWE-smith-phpcornstack_php_ru_enThe part of CoRNStack Dataset translated into Russian. Translation was done with Qwen3 model.
Samples that satisfy the dual consistency filtering condition (samples where the document_rank is 0 or 1 and document_score > 0.7) were translated.
Source code you can find here. For support: fedor.yaronskiy@gmail.com
PHP-Code-LargePHP-Code-Large
PHP-Code-Large is a large-scale corpus of PHP source code comprising more than 12 million lines of PHP code. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and static program analysis for the PHP ecosystem.
By providing a high-volume, language-specific corpus, PHP-Code-Large enables systematic experimentation in PHP-focused model training, domain adaptation, and downstream code… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/PHP-Code-Large.cornstack-php-v1
CoRNStack PHP Dataset
The CoRNStack Dataset, accepted to ICLR 2025, is a large-scale high quality training dataset specifically for code retrieval across multiple
programming languages. This dataset comprises of <query, positive, negative> triplets used to train nomic-embed-code,
CodeRankEmbed, and CodeRankLLM.
CoRNStack Dataset Curation
Starting with the deduplicated Stackv2, we create text-code pairs from function docstrings and respective code. We filtered out… See the full description on the dataset page: https://huggingface.co/datasets/nomic-ai/cornstack-php-v1.PHP-Code-LargePHP-Code-Large
PHP-Code-Large is a large-scale corpus of PHP source code comprising more than 12 million lines of PHP code. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and static program analysis for the PHP ecosystem.
By providing a high-volume, language-specific corpus, PHP-Code-Large enables systematic experimentation in PHP-focused model training, domain adaptation, and downstream code… See the full description on the dataset page: https://huggingface.co/datasets/Ujjwal-Tyagi/PHP-Code-Large.ph-pretrain
PH Pretrain — Philippine Languages Corpus (v0.6-ph-unified)
👉 Looking to train a Filipino/Tagalog model? Use jpaulpoliquit/ph-pretrain-03 instead — it is the recommended dataset. It is a ~1B-token, Tagalog-forward, quality-filtered web corpus. This dataset (ph-pretrain) is ~95% bot-templated Cebuano/Waray Wikipedia and is best only when you specifically want mass regional-language (Cebuano/Waray) coverage.
A cleaned, deduplicated, document-level pretraining corpus for… See the full description on the dataset page: https://huggingface.co/datasets/jpaulpoliquit/ph-pretrain.cornstack-php-v1-tevatron-1Mph-pretrain-03
PH Pretrain 03 — Filipino/Tagalog Web Corpus (ph-pretrain-03)
The recommended dataset for Filipino / Tagalog language-model pretraining and continued pretraining (CPT).
A ~1-billion-token, Tagalog-forward web corpus, cleaned, deduplicated, and quality-scored by the jpaulpoliquit/pretraining refinery.
1,933,988 documents · ≈ 1.0 billion tokens (Sailor2-1B, calibrated)
~97.7% Tagalog web (FineWeb2 + CC100), with regional + news + Wikipedia in the long tail
~3.41 GB to download… See the full description on the dataset page: https://huggingface.co/datasets/jpaulpoliquit/ph-pretrain-03.codexglue_code2text_php
Dataset Card for "codexglue_code2text_php"
More Information needed
php_cat1PHP-Code-LargePHP-Code-Large
PHP-Code-Large is a large-scale corpus of PHP source code comprising more than 12 million lines of PHP code. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and static program analysis for the PHP ecosystem.
By providing a high-volume, language-specific corpus, PHP-Code-Large enables systematic experimentation in PHP-focused model training, domain adaptation, and downstream code… See the full description on the dataset page: https://huggingface.co/datasets/xormania/PHP-Code-Large.stack-mined-php-v1code-search-net-php
Dataset Card for "code-search-net-php"
Dataset Summary
This dataset is the Php portion of the CodeSarchNet annotated with a summary column.The code-search-net dataset includes open source functions that include comments found at GitHub.The summary is a short description of what the function does.
Languages
The dataset's comments are in English and the functions are coded in Php
Data Splits
Train, test, validation labels are included in the dataset… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/code-search-net-php.CodeSearchNet-ccr-php-queries-corpus
Dataset Card for "CodeSearchNet-ccr-php-queries-corpus"
More Information needed
CodeSearchNet-php-queries-corpus
Dataset Card for "CodeSearchNet-php-queries-corpus"
More Information needed
gpt-5-mini-rebench-v2-phpinstructional_code-search-net-php
Dataset Card for "instructional_code-search-net-php"
Dataset Summary
This is an instructional dataset for PHP.
The dataset contains two different kind of tasks:
Given a piece of code generate a description of what it does.
Given a description generate a piece of code that fulfils the description.
Languages
The dataset is in English.
Data Splits
There are no splits.
Dataset Creation
May of 2023
Curation Rationale
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/instructional_code-search-net-php.CodeSearchNet-php-qrels
Dataset Card for "CodeSearchNet-php-qrels"
More Information needed
CodeSearchNet-ccr-php-qrels
Dataset Card for "CodeSearchNet-ccr-php-qrels"
More Information needed
code-text-php
Dataset is imported from CodeXGLUE and pre-processed using their script.
Where to find in Semeru:
The dataset can be found at /nfs/semeru/semeru_datasets/code_xglue/code-to-text/php in Semeru
CodeXGLUE -- Code-To-Text
Task Definition
The task is to generate natural language comments for a code, and evaluted by smoothed bleu-4 score.
Dataset
The dataset we use comes from CodeSearchNet and we filter the dataset as the following:
Remove… See the full description on the dataset page: https://huggingface.co/datasets/semeru/code-text-php.UDR_PHP
Dataset Card for "UDR_PHP"
More Information needed
ztl-sard-php-verdicts
Source and recipe: github.com/inventor1975/introspect — dataset/sard (commit 4e179ba). The corpus is synthetic (NIST/Stivalet generated test cases); no real project's code or vulnerability is in this dataset.
ZTL verdicts over the SARD / Stivalet PHP vulnerability suite
This dataset records what the introspect analyzer (the ZTL zero-trust judge over
a deterministic PHP atomizer) returns on the public NIST SARD / Stivalet PHP test
suite — one row per test file, the verdict and… See the full description on the dataset page: https://huggingface.co/datasets/Inventor1975/ztl-sard-php-verdicts.terminal_bench_2_r2egym_nl2bash_stack_bugsseq_stack_php_v2_20260222_044012terminal_bench_2_r2egym_nl2bash_stack_bugsseq_lr3e_5_exp_rpt_stack_php_v2_step24d00812bdev_set_v2_a1_stack_phpunit_20260811_174202php-treesitter-filtered-datasetsV2
Php CodeSearch Dataset (Shuu12121/php-treesitter-filtered-datasetsV2)
Dataset Description
This dataset contains PHP functions and methods paired with their PHPDoc comments, extracted from open-source PHP repositories on GitHub.
It is formatted similarly to the CodeSearchNet challenge dataset.
Each entry includes:
code: The source code of a php function or method.
docstring: The docstring or Javadoc associated with the function/method.
func_name: The name of the… See the full description on the dataset page: https://huggingface.co/datasets/Shuu12121/php-treesitter-filtered-datasetsV2.github-file-programs-dataset-phpSWE-smith-phpswebench_verified_random_100_folders_a1_stack_phpunit_20260818_162722stack_edu_php
