datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SWE-smith-phpcornstack_php_ru_enThe part of CoRNStack Dataset translated into Russian. Translation was done with Qwen3 model.
Samples that satisfy the dual consistency filtering condition (samples where the document_rank is 0 or 1 and document_score > 0.7) were translated.
Source code you can find here. For support: fedor.yaronskiy@gmail.com
phpwind-captcha-dataset
PHPWind Captcha Dataset
Labelled four-digit numeric captcha images for training and evaluating OCR on
authorized PHPWind deployments. This repository stores each visual captcha
family in an independent dataset directory so that samples from different
versions, forks, themes, or generators are never silently mixed.
中文说明: README_zh.md
Related model: FlanChanXwO/phpwind-captcha-ocr
Dataset catalogue
Dataset ID
Deployment or version identifier
Status
Images… See the full description on the dataset page: https://huggingface.co/datasets/FlanChanXwO/phpwind-captcha-dataset.PHP-Code-LargePHP-Code-Large
PHP-Code-Large is a large-scale corpus of PHP source code comprising more than 12 million lines of PHP code. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and static program analysis for the PHP ecosystem.
By providing a high-volume, language-specific corpus, PHP-Code-Large enables systematic experimentation in PHP-focused model training, domain adaptation, and downstream code… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/PHP-Code-Large.PHP-Code-LargePHP-Code-Large
PHP-Code-Large is a large-scale corpus of PHP source code comprising more than 12 million lines of PHP code. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and static program analysis for the PHP ecosystem.
By providing a high-volume, language-specific corpus, PHP-Code-Large enables systematic experimentation in PHP-focused model training, domain adaptation, and downstream code… See the full description on the dataset page: https://huggingface.co/datasets/xormania/PHP-Code-Large.cornstack-php-v1-tevatron-1Mph-pretrain
PH Pretrain — Philippine Languages Corpus (v0.6-ph-unified)
👉 Looking to train a Filipino/Tagalog model? Use jpaulpoliquit/ph-pretrain-03 instead — it is the recommended dataset. It is a ~1B-token, Tagalog-forward, quality-filtered web corpus. This dataset (ph-pretrain) is ~95% bot-templated Cebuano/Waray Wikipedia and is best only when you specifically want mass regional-language (Cebuano/Waray) coverage.
A cleaned, deduplicated, document-level pretraining corpus for… See the full description on the dataset page: https://huggingface.co/datasets/jpaulpoliquit/ph-pretrain.PHP-Code-LargePHP-Code-Large
PHP-Code-Large is a large-scale corpus of PHP source code comprising more than 12 million lines of PHP code. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and static program analysis for the PHP ecosystem.
By providing a high-volume, language-specific corpus, PHP-Code-Large enables systematic experimentation in PHP-focused model training, domain adaptation, and downstream code… See the full description on the dataset page: https://huggingface.co/datasets/Ujjwal-Tyagi/PHP-Code-Large.ph-pretrain-03
PH Pretrain 03 — Filipino/Tagalog Web Corpus (ph-pretrain-03)
The recommended dataset for Filipino / Tagalog language-model pretraining and continued pretraining (CPT).
A ~1-billion-token, Tagalog-forward web corpus, cleaned, deduplicated, and quality-scored by the jpaulpoliquit/pretraining refinery.
1,933,988 documents · ≈ 1.0 billion tokens (Sailor2-1B, calibrated)
~97.7% Tagalog web (FineWeb2 + CC100), with regional + news + Wikipedia in the long tail
~3.41 GB to download… See the full description on the dataset page: https://huggingface.co/datasets/jpaulpoliquit/ph-pretrain-03.cornstack-php-v1
CoRNStack PHP Dataset
The CoRNStack Dataset, accepted to ICLR 2025, is a large-scale high quality training dataset specifically for code retrieval across multiple
programming languages. This dataset comprises of <query, positive, negative> triplets used to train nomic-embed-code,
CodeRankEmbed, and CodeRankLLM.
CoRNStack Dataset Curation
Starting with the deduplicated Stackv2, we create text-code pairs from function docstrings and respective code. We filtered out… See the full description on the dataset page: https://huggingface.co/datasets/nomic-ai/cornstack-php-v1.Stack2Graph_KG_php
PHP StackOverflow Knowledge Graph
Summary
This Hugging Face dataset repository contains the PHP shard of the Stack2Graph StackOverflow Knowledge Graph.
Hugging Face uses one dataset repository per programming language, so this repository is directly cloneable without an extra top-level archive wrapper.
The artifact is optimized for graph-based retrieval, SPARQL analytics, and retrieval-augmented question answering over Stack Overflow content.
Stack2Graph source:… See the full description on the dataset page: https://huggingface.co/datasets/Mo7art/Stack2Graph_KG_php.codexglue_code2text_php
Dataset Card for "codexglue_code2text_php"
More Information needed
php_cat1phpA parallel corpus originally extracted from http://se.php.net/download-docs.php. The original documents are written in English and have been partly translated into 21 languages. The original manuals contain about 500,000 words. The amount of actually translated texts varies for different languages between 50,000 and 380,000 words. The corpus is rather noisy and may include parts from the English original in some of the translations. The corpus is tokenized and each language pair has been sentence aligned.
23 languages, 252 bitexts
total number of files: 71,414
total number of tokens: 3.28M
total number of sentence fragments: 1.38Mmarin-starcoderdata_phpstack-mined-php-v1code-search-net-php
Dataset Card for "code-search-net-php"
Dataset Summary
This dataset is the Php portion of the CodeSarchNet annotated with a summary column.The code-search-net dataset includes open source functions that include comments found at GitHub.The summary is a short description of what the function does.
Languages
The dataset's comments are in English and the functions are coded in Php
Data Splits
Train, test, validation labels are included in the dataset… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/code-search-net-php.CodeSearchNet-ccr-php-queries-corpus
Dataset Card for "CodeSearchNet-ccr-php-queries-corpus"
More Information needed
instructional_code-search-net-php
Dataset Card for "instructional_code-search-net-php"
Dataset Summary
This is an instructional dataset for PHP.
The dataset contains two different kind of tasks:
Given a piece of code generate a description of what it does.
Given a description generate a piece of code that fulfils the description.
Languages
The dataset is in English.
Data Splits
There are no splits.
Dataset Creation
May of 2023
Curation Rationale
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/instructional_code-search-net-php.CodeSearchNet-php-queries-corpus
Dataset Card for "CodeSearchNet-php-queries-corpus"
More Information needed
CodeSearchNet-php-qrels
Dataset Card for "CodeSearchNet-php-qrels"
More Information needed
terminal_bench_2_r2egym_nl2bash_stack_bugsseq_stack_php_v2_20260222_044012CodeSearchNet-ccr-php-qrels
Dataset Card for "CodeSearchNet-ccr-php-qrels"
More Information needed
terminal_bench_2_r2egym_nl2bash_stack_bugsseq_lr3e_5_exp_rpt_stack_php_v2_step24d00812bUDR_PHP
Dataset Card for "UDR_PHP"
More Information needed
exp_rpt_stack-php-large_10k_glm_4.7_traces_jupitersmollm3-stack-v2-PHPdev_set_v2_a1_stack_phpunit_20260811_174202ztl-sard-php-verdicts
Source and recipe: github.com/inventor1975/introspect — dataset/sard (commit 4e179ba). The corpus is synthetic (NIST/Stivalet generated test cases); no real project's code or vulnerability is in this dataset.
ZTL verdicts over the SARD / Stivalet PHP vulnerability suite
This dataset records what the introspect analyzer (the ZTL zero-trust judge over
a deterministic PHP atomizer) returns on the public NIST SARD / Stivalet PHP test
suite — one row per test file, the verdict and… See the full description on the dataset page: https://huggingface.co/datasets/Inventor1975/ztl-sard-php-verdicts.code-text-php
Dataset is imported from CodeXGLUE and pre-processed using their script.
Where to find in Semeru:
The dataset can be found at /nfs/semeru/semeru_datasets/code_xglue/code-to-text/php in Semeru
CodeXGLUE -- Code-To-Text
Task Definition
The task is to generate natural language comments for a code, and evaluted by smoothed bleu-4 score.
Dataset
The dataset we use comes from CodeSearchNet and we filter the dataset as the following:
Remove… See the full description on the dataset page: https://huggingface.co/datasets/semeru/code-text-php.
