datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
python-codes-25k
License
MIT
This is a Cleaned Python Dataset Covering 25,000 Instructional Tasks
Overview
The dataset has 4 key features (fields): instruction, input, output, and text.It's a rich source for Python codes, tasks, and extends into behavioral aspects.
Dataset Statistics
Total Entries: 24,813
Unique Instructions: 24,580
Unique Inputs: 3,666
Unique Outputs: 24,581
Unique Texts: 24,813
Average Tokens per example: 508
Features… See the full description on the dataset page: https://huggingface.co/datasets/flytech/python-codes-25k.prop-firm-discount-codes
PropFirmDiscount — Verified Prop Firm Discount Codes
As of 2026-10-04, the biggest verified prop firm discount code is PFD from Upcomers — 90% off. Every firm's code is listed below, newest deals first.
Public dataset mirror of propfirmdiscount.com — verified prop firm discount codes, funding deals, and Trustpilot ratings. The source of truth is the live site; this repository is a read-only distribution channel, synced hourly.
Find a firm's discount code
Every… See the full description on the dataset page: https://huggingface.co/datasets/propfirmdiscounts/prop-firm-discount-codes.Finance-Conversational-Dataset-IndicLaw-Conversational-Dataset-IndicFinance-Parallel-Dataset-IndicNCERT-Conversational-Dataset-IndicCyber-Parallel-Dataset-IndicLaw-Parallel-Dataset-Indicagri-vet-multilingual-dataset
Agri-Vet Multilingual Dataset
This repository contains JSON and JSONL resources for multilingual agricultural and veterinary language tasks. It is intended to support conversational, retrieval, classification, or instruction-tuning experiments spanning crop, animal, and related user questions.
Working with the files
Inspect each JSON/JSONL record and preserve its language, domain, prompt, response, label, and provenance fields when creating a derived dataset.… See the full description on the dataset page: https://huggingface.co/datasets/MR-CODESPIKE/agri-vet-multilingual-dataset.Cyber-Conversational-Dataset-IndicComputer-Science-Conversational-Dataset-IndicCA-Conversational-Dataset-IndicMedical-Conversational-Dataset-IndicCodeSecEval
CodeSecEval
CodeSecEval is an execution-based benchmark for evaluating large language models on secure code generation and insecure-code repair.
The benchmark contains 255 Python programming tasks spanning 77 CWE vulnerability categories. Each task provides a problem specification, an insecure implementation, a secure reference implementation, executable tests, and an entry point.
Dataset Subsets
This repository contains two subsets:
SecEvalBase: 115 tasks… See the full description on the dataset page: https://huggingface.co/datasets/JasonWang1/CodeSecEval.Medical-Parallel-Dataset-IndicCA-Parallel-Dataset-Indiccode_shieldgas-cylinder-color-codes
Gas cylinder color codes by standard and country
Canonical, always-current version: https://referencesource.org/gas-cylinder-color-codes/
Machine-readable: https://referencesource.org/gas-cylinder-color-codes/data.json — this mirror is a point-in-time copy.
Last verified: 2026-10-07
Stale after: 2028-08-04 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 17
Color coding of compressed gas cylinders (shoulder and body… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/gas-cylinder-color-codes.CodeSpan
CodeSpan
64K source-code continuation for long-context decoding.
CodeSpan contains 32 examples from 17 open-source projects, including LLVM,
GCC, Linux, and PostgreSQL. Each example is a contiguous prefix of a distinct
source file, wrapped as a continuation prompt: 65,536 Qwen3 tokens, including
the chat template with thinking disabled. Files are neither concatenated nor
repeated. There are at most four files per project; 28 of the 32 files are C/C++.
Configuration… See the full description on the dataset page: https://huggingface.co/datasets/Hehy/CodeSpan.industrial-vfd-fault-alarm-codes-by-brand
Industrial VFD fault and alarm codes by brand
Canonical, always-current version: https://referencesource.org/industrial-vfd-fault-alarm-codes-by-brand/
Machine-readable: https://referencesource.org/industrial-vfd-fault-alarm-codes-by-brand/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-26
Stale after: 2028-08-25 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 340
What the fault and alarm codes… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/industrial-vfd-fault-alarm-codes-by-brand.KlearReasoner-CodeSub-15K
Dataset Summary
This dataset is a high-quality subset of the Klear-Reasoner Code RL dataset, derived from the RL data used in the rllm project. Part of this data contributed to training Klear-Reasoner’s code reasoning models.
The dataset is carefully cleaned and filtered to include only reliable samples suitable for reinforcement learning. Models trained with this dataset have shown substantial performance improvements across various code reasoning benchmarks.
You can load the… See the full description on the dataset page: https://huggingface.co/datasets/Kwai-Klear/KlearReasoner-CodeSub-15K.all-Codestral-22B-v0.1CodeSec-Pairs
CodeSec-Pairs
CodeSec-Pairs is a dataset of matched safe and vulnerable Python code pairs. Each
pair implements the same task but differs in whether it contains a security
vulnerability. Vulnerability labels come from CodeQL static analysis. The dataset is
built to study and steer the internal mechanisms that distinguish safe from vulnerable
code generation in LLMs.
Dataset Details
Each record pairs a CodeQL-clean safe_code with a CodeQL-flagged vuln_code for the… See the full description on the dataset page: https://huggingface.co/datasets/haaao821/CodeSec-Pairs.codeswitch-pairs-lase
Codeswitch Pairs LASE — training corpus
1118 same-voice cross-script utterance pairs (8 ElevenLabs Multilingual voices × en/hi/te/ta) used to train the LASE r1 speaker encoder.
Each row is one synthesized utterance with metadata; pairs are reconstructed at evaluation time by joining on voice_id (same voice, different script = cross-script pair).
Schema (manifest.jsonl)
{
"voice_id": "21m00Tcm4TlvDq8ikWAM",
"lang": "en | hi | te | ta",
"text": "the prompt text"… See the full description on the dataset page: https://huggingface.co/datasets/Praxel/codeswitch-pairs-lase.CodeScopeappliance-fault-error-codes
Appliance fault and error codes
Canonical, always-current version: https://referencesource.org/appliance-fault-error-codes/
Machine-readable: https://referencesource.org/appliance-fault-error-codes/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-04
Stale after: 2027-08-04 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 149
Error and fault codes displayed by major home appliances (dishwashers… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/appliance-fault-error-codes.python-codes-25k
License
MIT
This is a Cleaned Python Dataset Covering 25,000 Instructional Tasks
Overview
The dataset has 4 key features (fields): instruction, input, output, and text.It's a rich source for Python codes, tasks, and extends into behavioral aspects.
Dataset Statistics
Total Entries: 24,813
Unique Instructions: 24,580
Unique Inputs: 3,666
Unique Outputs: 24,581
Unique Texts: 24,813
Average Tokens per example: 508… See the full description on the dataset page: https://huggingface.co/datasets/Rishabh23456789/python-codes-25k.adaption-moroccan-darija-prompts-trilingual-codeswitch-chat-augmented
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-moroccan_darija_prompts & trilingual_codeswitch_chat (augmented)
This dataset consists of short conversational prompts written in Moroccan Darija, covering topics like shopping, social interactions, and daily inquiries. Each entry contains a single prompt with a null completion, indicating it is likely intended for instruction tuning or completion generation tasks. The content… See the full description on the dataset page: https://huggingface.co/datasets/oumayma03/adaption-moroccan-darija-prompts-trilingual-codeswitch-chat-augmented.CAT-Conversational-Dataset-Indictiny-codes-csharp
