Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SWE-bench /SWE-smith-phptextn<1K0 likes1.6k downloads10mo agoHugging Face02fyaronskiy /cornstack_php_ru_enThe part of CoRNStack Dataset translated into Russian. Translation was done with Qwen3 model. Samples that satisfy the dual consistency filtering condition (samples where the document_rank is 0 or 1 and document_score > 0.7) were translated. Source code you can find here. For support: fedor.yaronskiy@gmail.com textsentence-similarity1M<n<10M0 likes752 downloads8mo agoHugging Face03ajibawa-2023 /PHP-Code-LargePHP-Code-Large PHP-Code-Large is a large-scale corpus of PHP source code comprising more than 12 million lines of PHP code. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and static program analysis for the PHP ecosystem. By providing a high-volume, language-specific corpus, PHP-Code-Large enables systematic experimentation in PHP-focused model training, domain adaptation, and downstream code… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/PHP-Code-Large.texttext-generation1M<n<10M22 likes548 downloads8mo agoHugging Face04nomic-ai /cornstack-php-v1 CoRNStack PHP Dataset The CoRNStack Dataset, accepted to ICLR 2025, is a large-scale high quality training dataset specifically for code retrieval across multiple programming languages. This dataset comprises of <query, positive, negative> triplets used to train nomic-embed-code, CodeRankEmbed, and CodeRankLLM. CoRNStack Dataset Curation Starting with the deduplicated Stackv2, we create text-code pairs from function docstrings and respective code. We filtered out… See the full description on the dataset page: https://huggingface.co/datasets/nomic-ai/cornstack-php-v1.text10M<n<100M3 likes332 downloads2y agoHugging Face05Ujjwal-Tyagi /PHP-Code-LargePHP-Code-Large PHP-Code-Large is a large-scale corpus of PHP source code comprising more than 12 million lines of PHP code. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and static program analysis for the PHP ecosystem. By providing a high-volume, language-specific corpus, PHP-Code-Large enables systematic experimentation in PHP-focused model training, domain adaptation, and downstream code… See the full description on the dataset page: https://huggingface.co/datasets/Ujjwal-Tyagi/PHP-Code-Large.texttext-generation1M<n<10M0 likes275 downloads6mo agoHugging Face06jpaulpoliquit /ph-pretrain PH Pretrain — Philippine Languages Corpus (v0.6-ph-unified) 👉 Looking to train a Filipino/Tagalog model? Use jpaulpoliquit/ph-pretrain-03 instead — it is the recommended dataset. It is a ~1B-token, Tagalog-forward, quality-filtered web corpus. This dataset (ph-pretrain) is ~95% bot-templated Cebuano/Waray Wikipedia and is best only when you specifically want mass regional-language (Cebuano/Waray) coverage. A cleaned, deduplicated, document-level pretraining corpus for… See the full description on the dataset page: https://huggingface.co/datasets/jpaulpoliquit/ph-pretrain.tabulartext-generation1M<n<10M0 likes264 downloads4mo agoHugging Face07nthakur /cornstack-php-v1-tevatron-1Mtext100K<n<1M0 likes252 downloads1y agoHugging Face08jpaulpoliquit /ph-pretrain-03 PH Pretrain 03 — Filipino/Tagalog Web Corpus (ph-pretrain-03) The recommended dataset for Filipino / Tagalog language-model pretraining and continued pretraining (CPT). A ~1-billion-token, Tagalog-forward web corpus, cleaned, deduplicated, and quality-scored by the jpaulpoliquit/pretraining refinery. 1,933,988 documents · ≈ 1.0 billion tokens (Sailor2-1B, calibrated) ~97.7% Tagalog web (FineWeb2 + CC100), with regional + news + Wikipedia in the long tail ~3.41 GB to download… See the full description on the dataset page: https://huggingface.co/datasets/jpaulpoliquit/ph-pretrain-03.tabulartext-generation1M<n<10M0 likes238 downloads4mo agoHugging Face09CM /codexglue_code2text_php Dataset Card for "codexglue_code2text_php" More Information needed text100K<n<1M2 likes206 downloads3y agoHugging Face10DaniilOr /php_cat1tabular10K<n<100K0 likes193 downloads1y agoHugging Face11xormania /PHP-Code-LargePHP-Code-Large PHP-Code-Large is a large-scale corpus of PHP source code comprising more than 12 million lines of PHP code. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and static program analysis for the PHP ecosystem. By providing a high-volume, language-specific corpus, PHP-Code-Large enables systematic experimentation in PHP-focused model training, domain adaptation, and downstream code… See the full description on the dataset page: https://huggingface.co/datasets/xormania/PHP-Code-Large.texttext-generation1M<n<10M0 likes167 downloads6mo agoHugging Face12corniclr25 /stack-mined-php-v1text1M<n<10M1 likes108 downloads2y agoHugging Face13Nan-Do /code-search-net-php Dataset Card for "code-search-net-php" Dataset Summary This dataset is the Php portion of the CodeSarchNet annotated with a summary column.The code-search-net dataset includes open source functions that include comments found at GitHub.The summary is a short description of what the function does. Languages The dataset's comments are in English and the functions are coded in Php Data Splits Train, test, validation labels are included in the dataset… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/code-search-net-php.texttext-generation100K<n<1M1 likes104 downloads3y agoHugging Face14CoIR-Retrieval /CodeSearchNet-ccr-php-queries-corpus Dataset Card for "CodeSearchNet-ccr-php-queries-corpus" More Information needed text100K<n<1M0 likes86 downloads2y agoHugging Face15CoIR-Retrieval /CodeSearchNet-php-queries-corpus Dataset Card for "CodeSearchNet-php-queries-corpus" More Information needed text100K<n<1M2 likes81 downloads2y agoHugging Face16beranki /gpt-5-mini-rebench-v2-phptabularn<1K0 likes70 downloads5mo agoHugging Face17Nan-Do /instructional_code-search-net-php Dataset Card for "instructional_code-search-net-php" Dataset Summary This is an instructional dataset for PHP. The dataset contains two different kind of tasks: Given a piece of code generate a description of what it does. Given a description generate a piece of code that fulfils the description. Languages The dataset is in English. Data Splits There are no splits. Dataset Creation May of 2023 Curation Rationale This dataset… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/instructional_code-search-net-php.texttext-generation100K<n<1M3 likes69 downloads3y agoHugging Face18CoIR-Retrieval /CodeSearchNet-php-qrels Dataset Card for "CodeSearchNet-php-qrels" More Information needed text100K<n<1M0 likes65 downloads2y agoHugging Face19CoIR-Retrieval /CodeSearchNet-ccr-php-qrels Dataset Card for "CodeSearchNet-ccr-php-qrels" More Information needed text100K<n<1M0 likes62 downloads2y agoHugging Face20semeru /code-text-php Dataset is imported from CodeXGLUE and pre-processed using their script. Where to find in Semeru: The dataset can be found at /nfs/semeru/semeru_datasets/code_xglue/code-to-text/php in Semeru CodeXGLUE -- Code-To-Text Task Definition The task is to generate natural language comments for a code, and evaluted by smoothed bleu-4 score. Dataset The dataset we use comes from CodeSearchNet and we filter the dataset as the following: Remove… See the full description on the dataset page: https://huggingface.co/datasets/semeru/code-text-php.text100K<n<1M2 likes53 downloads4y agoHugging Face21KaiLv /UDR_PHP Dataset Card for "UDR_PHP" More Information needed tabular100K<n<1M0 likes52 downloads3y agoHugging Face22Inventor1975 /ztl-sard-php-verdicts Source and recipe: github.com/inventor1975/introspect — dataset/sard (commit 4e179ba). The corpus is synthetic (NIST/Stivalet generated test cases); no real project's code or vulnerability is in this dataset. ZTL verdicts over the SARD / Stivalet PHP vulnerability suite This dataset records what the introspect analyzer (the ZTL zero-trust judge over a deterministic PHP atomizer) returns on the public NIST SARD / Stivalet PHP test suite — one row per test file, the verdict and… See the full description on the dataset page: https://huggingface.co/datasets/Inventor1975/ztl-sard-php-verdicts.text10K<n<100K0 likes50 downloads16d agoHugging Face23DCAgent2 /terminal_bench_2_r2egym_nl2bash_stack_bugsseq_stack_php_v2_20260222_044012textn<1K0 likes48 downloads8mo agoHugging Face24DCAgent2 /terminal_bench_2_r2egym_nl2bash_stack_bugsseq_lr3e_5_exp_rpt_stack_php_v2_step24d00812btextn<1K0 likes45 downloads8mo agoHugging Face25laion /dev_set_v2_a1_stack_phpunit_20260811_174202text1K<n<10K0 likes43 downloads2mo agoHugging Face26Shuu12121 /php-treesitter-filtered-datasetsV2 Php CodeSearch Dataset (Shuu12121/php-treesitter-filtered-datasetsV2) Dataset Description This dataset contains PHP functions and methods paired with their PHPDoc comments, extracted from open-source PHP repositories on GitHub. It is formatted similarly to the CodeSearchNet challenge dataset. Each entry includes: code: The source code of a php function or method. docstring: The docstring or Javadoc associated with the function/method. func_name: The name of the… See the full description on the dataset page: https://huggingface.co/datasets/Shuu12121/php-treesitter-filtered-datasetsV2.text1M<n<10M1 likes35 downloads1y agoHugging Face27Shuu12121 /github-file-programs-dataset-phptext100K<n<1M0 likes35 downloads9mo agoHugging Face28shanxianzheng /SWE-smith-phptextn<1K0 likes35 downloads2mo agoHugging Face29laion /swebench_verified_random_100_folders_a1_stack_phpunit_20260818_162722text10K<n<100K0 likes35 downloads2mo agoHugging Face30hongliu9903 /stack_edu_phptabular1M<n<10M0 likes33 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.