Team Ai
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ajibawa-2023 /PHP-Code-LargePHP-Code-Large PHP-Code-Large is a large-scale corpus of PHP source code comprising more than 12 million lines of PHP code. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and static program analysis for the PHP ecosystem. By providing a high-volume, language-specific corpus, PHP-Code-Large enables systematic experimentation in PHP-focused model training, domain adaptation, and downstream code… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/PHP-Code-Large.texttext-generation1M<n<10M22 likes504 downloads8mo agoHugging Face02nomic-ai /cornstack-php-v1 CoRNStack PHP Dataset The CoRNStack Dataset, accepted to ICLR 2025, is a large-scale high quality training dataset specifically for code retrieval across multiple programming languages. This dataset comprises of <query, positive, negative> triplets used to train nomic-embed-code, CodeRankEmbed, and CodeRankLLM. CoRNStack Dataset Curation Starting with the deduplicated Stackv2, we create text-code pairs from function docstrings and respective code. We filtered out… See the full description on the dataset page: https://huggingface.co/datasets/nomic-ai/cornstack-php-v1.text10M<n<100M3 likes330 downloads2y agoHugging Face03Ujjwal-Tyagi /PHP-Code-LargePHP-Code-Large PHP-Code-Large is a large-scale corpus of PHP source code comprising more than 12 million lines of PHP code. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and static program analysis for the PHP ecosystem. By providing a high-volume, language-specific corpus, PHP-Code-Large enables systematic experimentation in PHP-focused model training, domain adaptation, and downstream code… See the full description on the dataset page: https://huggingface.co/datasets/Ujjwal-Tyagi/PHP-Code-Large.texttext-generation1M<n<10M0 likes273 downloads6mo agoHugging Face04xormania /PHP-Code-LargePHP-Code-Large PHP-Code-Large is a large-scale corpus of PHP source code comprising more than 12 million lines of PHP code. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and static program analysis for the PHP ecosystem. By providing a high-volume, language-specific corpus, PHP-Code-Large enables systematic experimentation in PHP-focused model training, domain adaptation, and downstream code… See the full description on the dataset page: https://huggingface.co/datasets/xormania/PHP-Code-Large.texttext-generation1M<n<10M0 likes216 downloads6mo agoHugging Face05Inventor1975 /ztl-sard-php-verdicts Source and recipe: github.com/inventor1975/introspect — dataset/sard (commit 4e179ba). The corpus is synthetic (NIST/Stivalet generated test cases); no real project's code or vulnerability is in this dataset. ZTL verdicts over the SARD / Stivalet PHP vulnerability suite This dataset records what the introspect analyzer (the ZTL zero-trust judge over a deterministic PHP atomizer) returns on the public NIST SARD / Stivalet PHP test suite — one row per test file, the verdict and… See the full description on the dataset page: https://huggingface.co/datasets/Inventor1975/ztl-sard-php-verdicts.text10K<n<100K0 likes48 downloads14d agoHugging Face06semeru /code-text-php Dataset is imported from CodeXGLUE and pre-processed using their script. Where to find in Semeru: The dataset can be found at /nfs/semeru/semeru_datasets/code_xglue/code-to-text/php in Semeru CodeXGLUE -- Code-To-Text Task Definition The task is to generate natural language comments for a code, and evaluted by smoothed bleu-4 score. Dataset The dataset we use comes from CodeSearchNet and we filter the dataset as the following: Remove… See the full description on the dataset page: https://huggingface.co/datasets/semeru/code-text-php.text100K<n<1M2 likes44 downloads4y agoHugging Face07nbuser32 /PHP-Webshell-Datasettext10K<n<100K0 likes22 downloads1y agoHugging Face08phpcorer /dididaitextn<1K0 likes11 downloads2y agoHugging Face09KonoZioDa /Java-Php-Vuln-Datasettextn<1K0 likes11 downloads1y agoHugging Face10phptensor10 /sn38r5-u74-subtextn<1K0 likes11 downloads2mo agoHugging Face11kloodia /php_200ktext100K<n<1M0 likes9 downloads2y agoHugging Face12dmeldrum6 /PHP_Master_QA_Dataset Dataset Card for PHP_Master_QA_Dataset PHP QA Dataset Dataset Details Dataset Description PHP QA Dataset including Questions from: General Programming Functions Strings Arrays Objects Dates and Times Web Functions and Web Services Databases Graphics Security Debugging textn<1K0 likes7 downloads8mo agoHugging Face13phptensor10 /sn38r4-u74-subtextn<1K0 likes7 downloads3mo agoHugging Face14phptensor10 /sn38r3-u175-subtextn<1K0 likes6 downloads3mo agoHugging Face15phptensor10 /sn38r6-u175-subtextn<1K0 likes6 downloads2mo agoHugging Face16gangiswag /php_ablationtext100K<n<1M0 likes5 downloads2y agoHugging Face17jimmywhite66 /phpcodesamtext100K<n<1M1 likes4 downloads2y agoHugging Face18BackpropBuff /PHP.en-jatext10K<n<100K0 likes4 downloads2y agoHugging Face19phpcorer /rwrtextn<1K0 likes3 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.