Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01flytech /python-codes-25k License MIT This is a Cleaned Python Dataset Covering 25,000 Instructional Tasks Overview The dataset has 4 key features (fields): instruction, input, output, and text.It's a rich source for Python codes, tasks, and extends into behavioral aspects. Dataset Statistics Total Entries: 24,813 Unique Instructions: 24,580 Unique Inputs: 3,666 Unique Outputs: 24,581 Unique Texts: 24,813 Average Tokens per example: 508 Features… See the full description on the dataset page: https://huggingface.co/datasets/flytech/python-codes-25k.texttext-classification10K<n<100K182 likes7.4k downloads2y agoHugging Face02propfirmdiscounts /prop-firm-discount-codes PropFirmDiscount — Verified Prop Firm Discount Codes As of 2026-10-04, the biggest verified prop firm discount code is PFD from Upcomers — 90% off. Every firm's code is listed below, newest deals first. Public dataset mirror of propfirmdiscount.com — verified prop firm discount codes, funding deals, and Trustpilot ratings. The source of truth is the live site; this repository is a read-only distribution channel, synced hourly. Find a firm's discount code Every… See the full description on the dataset page: https://huggingface.co/datasets/propfirmdiscounts/prop-firm-discount-codes.textn<1K0 likes994 downloads6h agoHugging Face03oss-codes /Finance-Conversational-Dataset-Indictext100K<n<1M1 likes913 downloads2y agoHugging Face04oss-codes /Law-Conversational-Dataset-Indictext100K<n<1M0 likes690 downloads2y agoHugging Face05oss-codes /Finance-Parallel-Dataset-Indictext100K<n<1M0 likes617 downloads2y agoHugging Face06oss-codes /NCERT-Conversational-Dataset-Indictext100K<n<1M0 likes611 downloads2y agoHugging Face07oss-codes /Cyber-Parallel-Dataset-Indictext1K<n<10K0 likes535 downloads2y agoHugging Face08oss-codes /Law-Parallel-Dataset-Indictext100K<n<1M0 likes480 downloads2y agoHugging Face09MR-CODESPIKE /agri-vet-multilingual-dataset Agri-Vet Multilingual Dataset This repository contains JSON and JSONL resources for multilingual agricultural and veterinary language tasks. It is intended to support conversational, retrieval, classification, or instruction-tuning experiments spanning crop, animal, and related user questions. Working with the files Inspect each JSON/JSONL record and preserve its language, domain, prompt, response, label, and provenance fields when creating a derived dataset.… See the full description on the dataset page: https://huggingface.co/datasets/MR-CODESPIKE/agri-vet-multilingual-dataset.textn<1K0 likes409 downloads1mo agoHugging Face10oss-codes /Cyber-Conversational-Dataset-Indictext1K<n<10K0 likes401 downloads2y agoHugging Face11oss-codes /Computer-Science-Conversational-Dataset-Indictext10K<n<100K0 likes294 downloads2y agoHugging Face12oss-codes /CA-Conversational-Dataset-Indictext100K<n<1M0 likes266 downloads2y agoHugging Face13oss-codes /Medical-Conversational-Dataset-Indictext10K<n<100K0 likes250 downloads2y agoHugging Face14JasonWang1 /CodeSecEval CodeSecEval CodeSecEval is an execution-based benchmark for evaluating large language models on secure code generation and insecure-code repair. The benchmark contains 255 Python programming tasks spanning 77 CWE vulnerability categories. Each task provides a problem specification, an insecure implementation, a secure reference implementation, executable tests, and an entry point. Dataset Subsets This repository contains two subsets: SecEvalBase: 115 tasks… See the full description on the dataset page: https://huggingface.co/datasets/JasonWang1/CodeSecEval.texttext-generationn<1K0 likes199 downloads3mo agoHugging Face15oss-codes /Medical-Parallel-Dataset-Indictext10K<n<100K0 likes128 downloads2y agoHugging Face16oss-codes /CA-Parallel-Dataset-Indictext100K<n<1M0 likes117 downloads2y agoHugging Face17QLWD /code_shieldtext100K<n<1M0 likes115 downloads2y agoHugging Face18referencesource /gas-cylinder-color-codes Gas cylinder color codes by standard and country Canonical, always-current version: https://referencesource.org/gas-cylinder-color-codes/ Machine-readable: https://referencesource.org/gas-cylinder-color-codes/data.json — this mirror is a point-in-time copy. Last verified: 2026-10-07 Stale after: 2028-08-04 (past this date, prefer the canonical copy — it re-verifies on a cadence this snapshot does not) Records: 17 Color coding of compressed gas cylinders (shoulder and body… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/gas-cylinder-color-codes.textn<1K0 likes111 downloads3d agoHugging Face19Hehy /CodeSpan CodeSpan 64K source-code continuation for long-context decoding. CodeSpan contains 32 examples from 17 open-source projects, including LLVM, GCC, Linux, and PostgreSQL. Each example is a contiguous prefix of a distinct source file, wrapped as a continuation prompt: 65,536 Qwen3 tokens, including the chat template with thinking disabled. Files are neither concatenated nor repeated. There are at most four files per project; 28 of the 32 files are C/C++. Configuration… See the full description on the dataset page: https://huggingface.co/datasets/Hehy/CodeSpan.texttext-generationn<1K0 likes99 downloads11d agoHugging Face20referencesource /industrial-vfd-fault-alarm-codes-by-brand Industrial VFD fault and alarm codes by brand Canonical, always-current version: https://referencesource.org/industrial-vfd-fault-alarm-codes-by-brand/ Machine-readable: https://referencesource.org/industrial-vfd-fault-alarm-codes-by-brand/data.json — this mirror is a point-in-time copy. Last verified: 2026-08-26 Stale after: 2028-08-25 (past this date, prefer the canonical copy — it re-verifies on a cadence this snapshot does not) Records: 340 What the fault and alarm codes… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/industrial-vfd-fault-alarm-codes-by-brand.textn<1K0 likes91 downloads4d agoHugging Face21Kwai-Klear /KlearReasoner-CodeSub-15K Dataset Summary This dataset is a high-quality subset of the Klear-Reasoner Code RL dataset, derived from the RL data used in the rllm project. Part of this data contributed to training Klear-Reasoner’s code reasoning models. The dataset is carefully cleaned and filtered to include only reliable samples suitable for reinforcement learning. Models trained with this dataset have shown substantial performance improvements across various code reasoning benchmarks. You can load the… See the full description on the dataset page: https://huggingface.co/datasets/Kwai-Klear/KlearReasoner-CodeSub-15K.text10K<n<100K7 likes89 downloads1y agoHugging Face22testcase-evaluate /all-Codestral-22B-v0.1text1M<n<10M0 likes88 downloads1y agoHugging Face23haaao821 /CodeSec-Pairs CodeSec-Pairs CodeSec-Pairs is a dataset of matched safe and vulnerable Python code pairs. Each pair implements the same task but differs in whether it contains a security vulnerability. Vulnerability labels come from CodeQL static analysis. The dataset is built to study and steer the internal mechanisms that distinguish safe from vulnerable code generation in LLMs. Dataset Details Each record pairs a CodeQL-clean safe_code with a CodeQL-flagged vuln_code for the… See the full description on the dataset page: https://huggingface.co/datasets/haaao821/CodeSec-Pairs.texttext-generation10K<n<100K1 likes86 downloads1mo agoHugging Face24Praxel /codeswitch-pairs-lase Codeswitch Pairs LASE — training corpus 1118 same-voice cross-script utterance pairs (8 ElevenLabs Multilingual voices × en/hi/te/ta) used to train the LASE r1 speaker encoder. Each row is one synthesized utterance with metadata; pairs are reconstructed at evaluation time by joining on voice_id (same voice, different script = cross-script pair). Schema (manifest.jsonl) { "voice_id": "21m00Tcm4TlvDq8ikWAM", "lang": "en | hi | te | ta", "text": "the prompt text"… See the full description on the dataset page: https://huggingface.co/datasets/Praxel/codeswitch-pairs-lase.audioaudio-classificationn<1K0 likes78 downloads5mo agoHugging Face25WeixiangYan /CodeScopetabulartranslationn<1K3 likes75 downloads3y agoHugging Face26referencesource /appliance-fault-error-codes Appliance fault and error codes Canonical, always-current version: https://referencesource.org/appliance-fault-error-codes/ Machine-readable: https://referencesource.org/appliance-fault-error-codes/data.json — this mirror is a point-in-time copy. Last verified: 2026-08-04 Stale after: 2027-08-04 (past this date, prefer the canonical copy — it re-verifies on a cadence this snapshot does not) Records: 149 Error and fault codes displayed by major home appliances (dishwashers… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/appliance-fault-error-codes.textn<1K0 likes64 downloads10d agoHugging Face27Rishabh23456789 /python-codes-25k License MIT This is a Cleaned Python Dataset Covering 25,000 Instructional Tasks Overview The dataset has 4 key features (fields): instruction, input, output, and text.It's a rich source for Python codes, tasks, and extends into behavioral aspects. Dataset Statistics Total Entries: 24,813 Unique Instructions: 24,580 Unique Inputs: 3,666 Unique Outputs: 24,581 Unique Texts: 24,813 Average Tokens per example: 508… See the full description on the dataset page: https://huggingface.co/datasets/Rishabh23456789/python-codes-25k.texttext-classification10K<n<100K0 likes64 downloads13d agoHugging Face28oumayma03 /adaption-moroccan-darija-prompts-trilingual-codeswitch-chat-augmented This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-moroccan_darija_prompts & trilingual_codeswitch_chat (augmented) This dataset consists of short conversational prompts written in Moroccan Darija, covering topics like shopping, social interactions, and daily inquiries. Each entry contains a single prompt with a null completion, indicating it is likely intended for instruction tuning or completion generation tasks. The content… See the full description on the dataset page: https://huggingface.co/datasets/oumayma03/adaption-moroccan-darija-prompts-trilingual-codeswitch-chat-augmented.text1K<n<10K0 likes62 downloads23d agoHugging Face29oss-codes /CAT-Conversational-Dataset-Indictext10K<n<100K0 likes61 downloads2y agoHugging Face30Azamorn /tiny-codes-csharptext100K<n<1M2 likes60 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.