Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01code-search-net /code_search_net Dataset Card for CodeSearchNet corpus Dataset Summary CodeSearchNet corpus is a dataset of 2 milllion (comment, code) pairs from opensource libraries hosted on GitHub. It contains code and documentation for several programming languages. CodeSearchNet corpus was gathered to support the CodeSearchNet challenge, to explore the problem of code retrieval using natural language. Supported Tasks and Leaderboards language-modeling: The dataset can be used to… See the full description on the dataset page: https://huggingface.co/datasets/code-search-net/code_search_net.texttext-generation1M<n<10M338 likes33k downloads8mo agoHugging Face02flytech /python-codes-25k License MIT This is a Cleaned Python Dataset Covering 25,000 Instructional Tasks Overview The dataset has 4 key features (fields): instruction, input, output, and text.It's a rich source for Python codes, tasks, and extends into behavioral aspects. Dataset Statistics Total Entries: 24,813 Unique Instructions: 24,580 Unique Inputs: 3,666 Unique Outputs: 24,581 Unique Texts: 24,813 Average Tokens per example: 508 Features… See the full description on the dataset page: https://huggingface.co/datasets/flytech/python-codes-25k.texttext-classification10K<n<100K183 likes7.4k downloads2y agoHugging Face03Nan-Do /code-search-net-python Dataset Card for "code-search-net-python" Dataset Description Homepage: None Repository: https://huggingface.co/datasets/Nan-Do/code-search-net-python Paper: None Leaderboard: None Point of Contact: @Nan-Do Dataset Summary This dataset is the Python portion of the CodeSarchNet annotated with a summary column.The code-search-net dataset includes open source functions that include comments found at GitHub.The summary is a short description of what the… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/code-search-net-python.texttext-generation100K<n<1M30 likes3.7k downloads3y agoHugging Face04CodeSoulco /TextInsightBench TextInsightBench English | 简体中文 A natural-language data-mining benchmark for agents: 50 tasks, 435,000 task documents and 944,468 unlabeled learning documents. Each task provides 5,000 or 10,000 texts and a research objective. Agents choose the patterns, populations and comparisons to investigate, then submit up to three findings with complete document assignments, exact quotations, statistics, counterexamples and limitations. Any analysis method is allowed. Contents… See the full description on the dataset page: https://huggingface.co/datasets/CodeSoulco/TextInsightBench.texttext-generation100K<n<1M0 likes1.1k downloads22d agoHugging Face05nampdn-ai /tiny-codesgated Reasoning with Language and Code This synthetic dataset is a collection of 1.6 millions short and clear code snippets that can help LLM models learn how to reason with both natural and programming languages. The dataset covers a wide range of programming languages, such as Python, TypeScript, JavaScript, Ruby, Julia, Rust, C++, Bash, Java, C#, and Go. It also includes two database languages: Cypher (for graph databases) and SQL (for relational databases) in order to study the… See the full description on the dataset page: https://huggingface.co/datasets/nampdn-ai/tiny-codes.texttext-generation1M<n<10M302 likes1k downloads3y agoHugging Face06claudios /code_search_net CodeSearchNet This is an unofficial reupload of the code_search_net dataset in the parquet format. I have also removed the columns func_code_tokens, func_documentation_tokens, and split_name as they are not relevant. The original repository relies on a Python module that is downloaded and executed to unpack the dataset, which is a potential security risk but importantly raises an annoying warning. As a plus, parquets load faster. Original model card: Dataset Card for… See the full description on the dataset page: https://huggingface.co/datasets/claudios/code_search_net.texttext-generation1M<n<10M11 likes332 downloads2y agoHugging Face07louisbrulenaudet /code-sante-publique Code de la santé publique, non-instruct (2025-07-11) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source language… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-sante-publique.tabulartext-generation1K<n<10K1 likes298 downloads1y agoHugging Face08JasonWang1 /CodeSecEval CodeSecEval CodeSecEval is an execution-based benchmark for evaluating large language models on secure code generation and insecure-code repair. The benchmark contains 255 Python programming tasks spanning 77 CWE vulnerability categories. Each task provides a problem specification, an insecure implementation, a secure reference implementation, executable tests, and an entry point. Dataset Subsets This repository contains two subsets: SecEvalBase: 115 tasks… See the full description on the dataset page: https://huggingface.co/datasets/JasonWang1/CodeSecEval.texttext-generationn<1K0 likes199 downloads3mo agoHugging Face09Nan-Do /code-search-net-java Dataset Card for "code-search-net-java" Dataset Summary This dataset is the Java portion of the CodeSarchNet annotated with a summary column.The code-search-net dataset includes open source functions that include comments found at GitHub.The summary is a short description of what the function does. Languages The dataset's comments are in English and the functions are coded in Java Data Splits Train, test, validation labels are included in the… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/code-search-net-java.textsummarization100K<n<1M4 likes189 downloads3y agoHugging Face10Malikeh1375 /code-switching-tokenizer-robustness Code-Switching Dataset for Tokenizer Robustness Analysis Dataset Description This dataset is designed for tokenizer robustness testing in multilingual and code-switching contexts. It contains identical content expressed across 16 different language variants, including pure English and 15 English-X code-switching pairs, allowing researchers to isolate tokenization effects from semantic differences when evaluating language models. Purpose Tokenizer Comparison:… See the full description on the dataset page: https://huggingface.co/datasets/Malikeh1375/code-switching-tokenizer-robustness.texttext-generation1K<n<10K2 likes129 downloads1y agoHugging Face11tandevllc /offsec_redteam_codesgated OffSec RedTeam Codes Token count: ~30B tokens. OffSec RedTeam Codes is a curated corpus of code (and some auxiliary text) extracted from popular GitHub repositories related to offensive security / red teaming (pentesting, OSINT, C2, privilege escalation, exploitation, forensics, etc.). It is also the largest open-source dataset of red-team and offensive-security code ever compiled. ⚠️ Ethical use only. This dataset is for research, education, and defensive security testing in… See the full description on the dataset page: https://huggingface.co/datasets/tandevllc/offsec_redteam_codes.tabulartext-generation1M<n<10M17 likes121 downloads11mo agoHugging Face12Nan-Do /code-search-net-javascript Dataset Card for "code-search-net-javascript" Dataset Summary This dataset is the JavaScript portion of the CodeSarchNet annotated with a summary column.The code-search-net dataset includes open source functions that include comments found at GitHub.The summary is a short description of what the function does. Languages The dataset's comments are in English and the functions are coded in JavaScript Data Splits Train, test, validation labels are… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/code-search-net-javascript.texttext-generation100K<n<1M7 likes104 downloads3y agoHugging Face13Nan-Do /code-search-net-php Dataset Card for "code-search-net-php" Dataset Summary This dataset is the Php portion of the CodeSarchNet annotated with a summary column.The code-search-net dataset includes open source functions that include comments found at GitHub.The summary is a short description of what the function does. Languages The dataset's comments are in English and the functions are coded in Php Data Splits Train, test, validation labels are included in the dataset… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/code-search-net-php.texttext-generation100K<n<1M1 likes104 downloads3y agoHugging Face14khemprogrammer /code-search-net-python Dataset Card for "code-search-net-python" Dataset Description Homepage: None Repository: https://huggingface.co/datasets/Nan-Do/code-search-net-python Paper: None Leaderboard: None Point of Contact: @Nan-Do Dataset Summary This dataset is the Python portion of the CodeSarchNet annotated with a summary column.The code-search-net dataset includes open source functions that include comments found at GitHub.The summary is a short description of what the… See the full description on the dataset page: https://huggingface.co/datasets/khemprogrammer/code-search-net-python.texttext-generation100K<n<1M0 likes99 downloads8mo agoHugging Face15Hehy /CodeSpan CodeSpan 64K source-code continuation for long-context decoding. CodeSpan contains 32 examples from 17 open-source projects, including LLVM, GCC, Linux, and PostgreSQL. Each example is a contiguous prefix of a distinct source file, wrapped as a continuation prompt: 65,536 Qwen3 tokens, including the chat template with thinking disabled. Files are neither concatenated nor repeated. There are at most four files per project; 28 of the 32 files are C/C++. Configuration… See the full description on the dataset page: https://huggingface.co/datasets/Hehy/CodeSpan.texttext-generationn<1K0 likes99 downloads11d agoHugging Face16haaao821 /CodeSec-Pairs CodeSec-Pairs CodeSec-Pairs is a dataset of matched safe and vulnerable Python code pairs. Each pair implements the same task but differs in whether it contains a security vulnerability. Vulnerability labels come from CodeQL static analysis. The dataset is built to study and steer the internal mechanisms that distinguish safe from vulnerable code generation in LLMs. Dataset Details Each record pairs a CodeQL-clean safe_code with a CodeQL-flagged vuln_code for the… See the full description on the dataset page: https://huggingface.co/datasets/haaao821/CodeSec-Pairs.texttext-generation10K<n<100K1 likes86 downloads1mo agoHugging Face17Zerothe00 /code-switched-student-blindspot-eval Code Switched Student Blind Spot Evaluation Overview This repository contains a small manual evaluation of Qwen/Qwen2.5-1.5B-Instruct on code-switched South Asian international student prompts. The goal is to test whether a small open-weight instruction model can understand Pakistani English mixed with Roman Urdu/Hindi in situations shaped by scholarship pressure, family expectations, limited resources, and international student life in Malaysia. Blind… See the full description on the dataset page: https://huggingface.co/datasets/Zerothe00/code-switched-student-blindspot-eval.text-generation1 likes80 downloads18d agoHugging Face18Nan-Do /code-search-net-go Dataset Card for "code-search-net-go" Dataset Summary This dataset is the Go portion of the CodeSarchNet annotated with a summary column.The code-search-net dataset includes open source functions that include comments found at GitHub.The summary is a short description of what the function does. Languages The dataset's comments are in English and the functions are coded in Go Data Splits Train, test, validation labels are included in the dataset as… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/code-search-net-go.texttext-generation100K<n<1M1 likes77 downloads3y agoHugging Face19WeixiangYan /CodeScopetabulartranslationn<1K3 likes75 downloads3y agoHugging Face20Zzzzzxl /code_search_net Dataset Card for CodeSearchNet corpus Dataset Summary CodeSearchNet corpus is a dataset of 2 milllion (comment, code) pairs from opensource libraries hosted on GitHub. It contains code and documentation for several programming languages. CodeSearchNet corpus was gathered to support the CodeSearchNet challenge, to explore the problem of code retrieval using natural language. Supported Tasks and Leaderboards language-modeling: The dataset can be used… See the full description on the dataset page: https://huggingface.co/datasets/Zzzzzxl/code_search_net.texttext-generation1M<n<10M0 likes74 downloads3mo agoHugging Face21carlscape /wisconsin-building-codes-qa Wisconsin Building Codes Q&A Dataset Dataset Description This dataset contains 13200 question-answer pairs focused on Wisconsin building codes, specifically covering: Building code requirements and regulations Administrative procedures and enforcement Construction standards and specifications Permit processes and compliance Dataset Structure Training samples: 11880 Validation samples: 1320 Each sample contains: instruction: A question about Wisconsin… See the full description on the dataset page: https://huggingface.co/datasets/carlscape/wisconsin-building-codes-qa.question-answering10K<n<100K1 likes69 downloads1y agoHugging Face22Rishabh23456789 /python-codes-25k License MIT This is a Cleaned Python Dataset Covering 25,000 Instructional Tasks Overview The dataset has 4 key features (fields): instruction, input, output, and text.It's a rich source for Python codes, tasks, and extends into behavioral aspects. Dataset Statistics Total Entries: 24,813 Unique Instructions: 24,580 Unique Inputs: 3,666 Unique Outputs: 24,581 Unique Texts: 24,813 Average Tokens per example: 508… See the full description on the dataset page: https://huggingface.co/datasets/Rishabh23456789/python-codes-25k.texttext-classification10K<n<100K0 likes64 downloads13d agoHugging Face23nuojohnchen /hugcode-codesft所有数据都是单轮代码指令数据 325696条英语,42816条中文。 license: cc text-generation100K<n<1M7 likes60 downloads3y agoHugging Face24rodriguescarson /adaption-codeskill-raw-selfoss-2k Self-OSS Code Instructions Python instructions with execution-filtered solutions. Rows 2,000 Domain programming Format data.parquet, one row per example Licence odc-by Built for supervised fine-tuning (SFT) experiments on Adaption AutoScientist Columns Column Description original_prompt The prompt (user turn) as uploaded. original_completion The target response as uploaded. enhanced_prompt Empty in this dataset.… See the full description on the dataset page: https://huggingface.co/datasets/rodriguescarson/adaption-codeskill-raw-selfoss-2k.tabulartext-generation1K<n<10K0 likes58 downloads14d agoHugging Face25rodriguescarson /adaption-codeskill-raw-commitpack-2k Commit-Message Code Edits Code-change tasks from real commits: an instruction (commit message) and the edited code. Rows 2,000 Domain programming Format data.parquet, one row per example Licence mit Built for supervised fine-tuning (SFT) experiments on Adaption AutoScientist Columns Column Description original_prompt The prompt (user turn) as uploaded. original_completion The target response as uploaded. enhanced_prompt… See the full description on the dataset page: https://huggingface.co/datasets/rodriguescarson/adaption-codeskill-raw-commitpack-2k.tabulartext-generation1K<n<10K0 likes58 downloads14d agoHugging Face26louisbrulenaudet /code-sport Code du sport, non-instruct (2025-07-11) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source language models… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-sport.tabulartext-generation1K<n<10K2 likes50 downloads1y agoHugging Face27burnerqmatrixacl /qmatrix-codesft-artifacts Q-Matrix Code-SFT selection artifacts Pool, per-item scores, and selection coresets for the paper "Item-Level Coreset Selectors Choose a Corpus They Never Scored: A Q-Matrix Audit for Code Supervised Fine-Tuning." This repository also carries the code, the Q-matrix, and the analysis that recomputes every table and the figure in the paper. Trained LoRA adapters are in the companion model repository. Contents qmatrix/ ontology, annotator… See the full description on the dataset page: https://huggingface.co/datasets/burnerqmatrixacl/qmatrix-codesft-artifacts.text-generation0 likes49 downloads2mo agoHugging Face28AmanPriyanshu /tool-reasoning-sft-RESEARCH-OpenHands-CodeScout_Training_Rollouts CodeScout Training Rollouts — Cleaned & Rectified ~40K multi-turn code localization agent trajectories converted into a strict reasoning + tool-call format with validated FSM transitions. Supports coupled (parallel) tool calls. ⚠️ Mid-training dataset. This dataset contains synthesized reasoning templates (not native chain-of-thought). It is suitable for mid-training to teach tool-use mechanics, FSM structure, and bash exploration patterns. It is not recommended as a final SFT… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-RESEARCH-OpenHands-CodeScout_Training_Rollouts.texttext-generation10K<n<100K0 likes47 downloads7mo agoHugging Face29Arjun-G-Ravi /Python-codes Dataset Card for Dataset Name Please note that this dataset maynot be perfect and may contain a very small quantity of non python codes. But the quantity appears to be very small Dataset Summary The dataset contains a collection of python question and their code. This is meant to be used for training models to be efficient in Python specific coding. The dataset has two features - 'question' and 'code'. An example is: {'question': 'Create a function that takes in a string… See the full description on the dataset page: https://huggingface.co/datasets/Arjun-G-Ravi/Python-codes.texttext-generation10K<n<100K7 likes46 downloads3y agoHugging Face30flytech /llama-python-codes-30k Python Codes - 30k examples, Llama1&2 tokenized dataset Author FlyTech For general guide on how to create, quantize, merge or inference the model and more, visit: hackmd.io/my_first_ai Overview This dataset serves as a rich resource for various Natural Language Processing tasks such as: Question Answering Text Generation Text-to-Text Generation It primarily focuses on instructional tasks in Python, tokenized specifically for the Llama architecture.… See the full description on the dataset page: https://huggingface.co/datasets/flytech/llama-python-codes-30k.textquestion-answering10K<n<100K19 likes46 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.