Team Ai
20 results

php

SWE-bench /SWE-smith-phptextn<1K0 likes1.6k downloads10mo agoHugging Facefyaronskiy /cornstack_php_ru_enThe part of CoRNStack Dataset translated into Russian. Translation was done with Qwen3 model. Samples that satisfy the dual consistency filtering condition (samples where the document_rank is 0 or 1 and document_score > 0.7) were translated. Source code you can find here. For support: fedor.yaronskiy@gmail.com textsentence-similarity1M<n<10M0 likes752 downloads8mo agoHugging FaceFlanChanXwO /phpwind-captcha-dataset PHPWind Captcha Dataset Labelled four-digit numeric captcha images for training and evaluating OCR on authorized PHPWind deployments. This repository stores each visual captcha family in an independent dataset directory so that samples from different versions, forks, themes, or generators are never silently mixed. 中文说明: README_zh.md Related model: FlanChanXwO/phpwind-captcha-ocr Dataset catalogue Dataset ID Deployment or version identifier Status Images… See the full description on the dataset page: https://huggingface.co/datasets/FlanChanXwO/phpwind-captcha-dataset.imageimage-to-textn<1K0 likes651 downloads2mo agoHugging Faceajibawa-2023 /PHP-Code-LargePHP-Code-Large PHP-Code-Large is a large-scale corpus of PHP source code comprising more than 12 million lines of PHP code. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and static program analysis for the PHP ecosystem. By providing a high-volume, language-specific corpus, PHP-Code-Large enables systematic experimentation in PHP-focused model training, domain adaptation, and downstream code… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/PHP-Code-Large.texttext-generation1M<n<10M22 likes548 downloads8mo agoHugging Facenomic-ai /cornstack-php-v1 CoRNStack PHP Dataset The CoRNStack Dataset, accepted to ICLR 2025, is a large-scale high quality training dataset specifically for code retrieval across multiple programming languages. This dataset comprises of <query, positive, negative> triplets used to train nomic-embed-code, CodeRankEmbed, and CodeRankLLM. CoRNStack Dataset Curation Starting with the deduplicated Stackv2, we create text-code pairs from function docstrings and respective code. We filtered out… See the full description on the dataset page: https://huggingface.co/datasets/nomic-ai/cornstack-php-v1.text10M<n<100M3 likes332 downloads2y agoHugging FaceUjjwal-Tyagi /PHP-Code-LargePHP-Code-Large PHP-Code-Large is a large-scale corpus of PHP source code comprising more than 12 million lines of PHP code. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and static program analysis for the PHP ecosystem. By providing a high-volume, language-specific corpus, PHP-Code-Large enables systematic experimentation in PHP-focused model training, domain adaptation, and downstream code… See the full description on the dataset page: https://huggingface.co/datasets/Ujjwal-Tyagi/PHP-Code-Large.texttext-generation1M<n<10M0 likes275 downloads6mo agoHugging Face