Team Ai
21 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AIencoder /llama-cpp-wheelsIf you like this please consider liking and donating (https://buymeacoffee.com/aiencoder) 🏭 llama-cpp-python Mega-Factory Wheels "Stop waiting for pip to compile. Just install and run." The most complete collection of pre-built llama-cpp-python wheels in existence — 8,333 wheels across every platform, Python version, backend, and CPU optimization level. No more cmake, gcc, or compilation hell. No more waiting 10 minutes for a build that might fail. Just find your wheel and… See the full description on the dataset page: https://huggingface.co/datasets/AIencoder/llama-cpp-wheels.text-generation1K<n<10K4 likes50k downloads13d agoHugging Face02skeole /qwen-cpp-agent-0-protocolExperiment in agentic autonomy protocols. ~ everything in this repo was created by Qwen 3.8 27B (Q4) running autonomously inside Deepseek Harness, on a single RTX 3090 GPU, for 3 weeks. The only human artifacts are: agents/* human/* AGENTS.md texttext-generation1K<n<10K3 likes6.9k downloads16d agoHugging Face03ajibawa-2023 /Cpp-Code-LargeCpp-Code-Large Cpp-Code-Large is a large-scale corpus of C++ source code comprising more than 5 million lines of C++ code. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and static program analysis for the C++ ecosystem. By providing a high-volume, language-specific corpus, Cpp-Code-Large enables systematic experimentation in C++-focused model training, domain adaptation, and downstream code… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/Cpp-Code-Large.texttext-generation1M<n<10M17 likes867 downloads7mo agoHugging Face04wttw /code_contest_instruct_cpptabulartext-generation1M<n<10M3 likes269 downloads2y agoHugging Face05malteklaes /cpp-code-code_search_net-style C++ Dataset documentation source: https://huggingface.co/docs/datasets/main/en/repository_structure Supported Tasks and Leaderboards language-modeling: The dataset can be used to train a model for modelling programming languages, which consists in building language models for programming languages. Language C++ programming language Dataset Structure Data Instances A data point consists of a function code along with its documentation.… See the full description on the dataset page: https://huggingface.co/datasets/malteklaes/cpp-code-code_search_net-style.texttext-generation10K<n<100K1 likes198 downloads2y agoHugging Face06AetherPrior /cpp_cwe_GRPO cpp_cwe_GRPO VeRL/GRPO-ready C++ security-coding dataset produced from every harness-passing stage-6 rewrite in the simple_gen pipeline. Files File Rows Description cpp_cwe_GRPO.parquet 8,039 Full corrected dataset cpp_cwe_GRPO_train.parquet 7,236 Deterministic 90% training split cpp_cwe_GRPO_val.parquet 803 Deterministic 10% validation split The split uses seed 20260929. Each record has a unique function name, and the two splits are disjoint.… See the full description on the dataset page: https://huggingface.co/datasets/AetherPrior/cpp_cwe_GRPO.texttext-generation1K<n<10K0 likes125 downloads7d agoHugging Face07AmareshHebbar /leetcode-codegen-cpp LeetCode Code-Gen Dataset — C++ 4025 rows. Given a problem statement, its input/output examples, and a required algorithm/technique, generate a correct C++ solution. Part of a 4-language collection built from the same source: see the sibling Python, Java, C++, and JavaScript datasets. Verification Not execution-verified. There is currently no compiler/runtime harness for this language in the build pipeline (only Python has one). Rows are extracted directly from… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/leetcode-codegen-cpp.texttext-generation1K<n<10K1 likes105 downloads3mo agoHugging Face08AsyncLab /cpp-project 中文 C++ 项目编程训练集 由 100 份中文 C++ 控制台项目语料转换而来,涵盖管理系统、实用工具、小游戏和基础算法。每个项目包含题目、原有解题思路和完整代码。适用于中文代码生成和指令微调.Author : KochiyaSanaeNya 数据格式 三个配置是100 条样本的不同表示,训练时选择其中一种.所有文件均为 UTF-8 无 BOM JSONL . 配置 训练字段 训练集 验证集 sft(默认) messages:user / assistant 对话 90 10 alpaca instruction / input / output 90 10 pretrain text:包含三个原始章节的全文 90 10 每条记录还包含 id、title、source_file 和 source_sha256,用于追溯来源,不需要作为训练输入。 SFT 用户消息包含 C++17 作答要求和原题;助手消息包含原有解题思路及 Markdown C++ 代码块。 Alpaca 的… See the full description on the dataset page: https://huggingface.co/datasets/AsyncLab/cpp-project.texttext-generationn<1K0 likes98 downloads23d agoHugging Face09verify-ppt /smollm3-stack-v2-Cpp Synthetic Pre-pretraining Datasets This dataset contains the pre-processed synthetic pre-pretraining (PPT) and pre-training (PT) data used in the paper Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior. It includes a range of PPT tasks (e.g., k-Shuffle Dyck, Set, MP-Struct Core, NCA) and standard pre-training mixtures (e.g., C4, SmolLM3, Olmo3, Marin) used to evaluate PPT at scale. Code: GitHub repository Project page: Hugging Face Organization For a… See the full description on the dataset page: https://huggingface.co/datasets/verify-ppt/smollm3-stack-v2-Cpp.text-generation0 likes63 downloads1d agoHugging Face10dumb-dev /cpp-10k10k random lines of the "text" column of the https://huggingface.co/datasets/wttw/code_contest_instruct_cpp dataset texttext-generation10K<n<100K1 likes54 downloads2y agoHugging Face11TokenBender /glm47-pie-cpp-posttraining-data GLM-4.7-Flash PIE C++ Post-Training Data The exact prepared dataset used for the GLM-4.7-Flash C++ performance post-training runs. Splits File Rows Purpose sft/train.jsonl 7,864 Supervised fine-tuning grpo/train.jsonl 7,887 GRPO prompt and reward evaluation eval/validation.jsonl 1,259 Full held-out evaluation eval/validation_mini126.jsonl 126 Fast evaluation eval/validation_mini4.jsonl 4 Smoke evaluation tasks.tar.gz 9,146 task JSONs Reward… See the full description on the dataset page: https://huggingface.co/datasets/TokenBender/glm47-pie-cpp-posttraining-data.texttext-generation10K<n<100K0 likes41 downloads3mo agoHugging Face12LLM-EDA /vgen_cpp Dataset Card for Opencores In the process of continual pre-training, we utilized the publicly available VGen dataset. VGen aggregates Verilog repositories from GitHub, systematically filters out duplicates and excessively large files, and retains only those files containing \texttt{module} and \texttt{endmodule} statements. We also incorporated the CodeSearchNet dataset \cite{codesearchnet}, which contains approximately 40MB function codes and their documentation.… See the full description on the dataset page: https://huggingface.co/datasets/LLM-EDA/vgen_cpp.texttext-generation10K<n<100K1 likes30 downloads2y agoHugging Face13MCES10-Software /CPP-Code-Solutions C++ Code Solutions Features 1000k of Python Code Solutions for Text Generation and Question Answering C++ Coding Problems labelled by topic and difficulty Recommendations Train your Model on Logical Operations and Mathematical Problems Before Training it on this. This is optional for Fine Tuning 2B parameter + models. Format the prompts in a orderly way when formatting data eg. {question} Solution: {solution} Topic: {topic} textquestion-answering10K<n<100K0 likes30 downloads1y agoHugging Face14michaelowusuntim6 /cpp-qwen35 C / C++ Code Corpus Description Teaches domain-specific instruction following and code generation for this expert. Source shareAI/CodeChat Mxode/StackOverflow-QA-C-Language-40k dumb-dev/cpp-10k AmareshHebbar/leetcode-codegen-cpp Formatted for the MoE-orchestrator project (https://github.com/michaelowusuntim6/MoE-orchestrator). Expert target: code_cpp. Format Each record is a JSON object with a messages field formatted for Qwen3.5's… See the full description on the dataset page: https://huggingface.co/datasets/michaelowusuntim6/cpp-qwen35.texttext-generation10K<n<100K0 likes25 downloads2d agoHugging Face15TokenBender /dummy-pie-cpp-oracle-seed Dummy PIE C++ Oracle Seed This is a tiny 5-row Hugging Face dataset skeleton for generating oracle SFT data for a PIE-style C++ optimization task. The rows are intentionally synthetic. They define the exact shape we want for real PIE C++ tasks without exposing or depending on private prepared artifacts. The goal is to give an oracle-data writer a clean contract: read a C++ optimization task; produce one verified optimized solution; return the result in chat/SFT format; include… See the full description on the dataset page: https://huggingface.co/datasets/TokenBender/dummy-pie-cpp-oracle-seed.text-generationn<1K0 likes22 downloads3mo agoHugging Face16mabo1215 /CPPB CPPB Summary CPPB is the public release surface for the Controlled Prompt-Privacy Benchmark introduced in BodhiPromptShield: Pre-Inference Prompt Mediation for Suppressing Privacy Propagation in LLM/VLM Agents. This Hugging Face package intentionally releases the benchmark-authored prompt manifest and template-stratified train/dev/test split, not raw third-party prompts, source images, or end-to-end OCR assets. Each row is a controlled prompt stub with benchmark metadata… See the full description on the dataset page: https://huggingface.co/datasets/mabo1215/CPPB.texttext-generationn<1K1 likes20 downloads6mo agoHugging Face17TokenBender /glm47-aider-cpp-rl-tasks GLM-4.7 Aider-style C++ RL tasks v2 This access-controlled package contains 253 independently authored C++ repository-editing tasks used for executable-reward post-training of GLM-4.7-Flash. These are training tasks, not official Aider or Exercism benchmark tasks. The official 26-task Aider C++ evaluation set is not included and remains evaluation-only. The package proves zero official task-ID overlap. Each task provides an Aider whole-file-edit prompt, starter repository… See the full description on the dataset page: https://huggingface.co/datasets/TokenBender/glm47-aider-cpp-rl-tasks.text-generation0 likes18 downloads3mo agoHugging Face18juiceb0xc0de /llama-cpp-cuda12-ada89llama.cpp CUDA 12.8.1, target GPU: sm89 If you're new or just starting to learn setting up your own inferences this llama.cpp wheel will work for if you're using python, CUDA 12.8.1 with one of the following Ada Lovelace generation (sm_89) GPU's: NVIDIA L4 NVIDIA L40 NVIDIA L40S NVIDIA RTX 6000 Ada Generation NVIDIA RTX 5000 Ada Generation NVIDIA RTX 4500 Ada Generation NVIDIA RTX 4000 Ada Generation NVIDIA RTX 4000 SFF Ada Generation NVIDIA RTX 2000 Ada Generation NVIDIA GeForce RTX 4090… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/llama-cpp-cuda12-ada89.text-generation0 likes15 downloads2mo agoHugging Face19gonzalolinares /cpp-compiler-prefs C++ compiler preferences (DPO) Offline preferences: chosen answers compile; rejected answers fail g++. texttext-generationn<1K0 likes14 downloads2mo agoHugging Face20gonzalolinares /cpp-compiler-curriculum C++ compiler curriculum (SFT) Synthetic C++ examples derived from Ubuntu toolchain headers and verified with g++ -std=c++20. texttext-generationn<1K0 likes12 downloads2mo agoHugging Face21gonzalolinares /cpp-compiler-grpo C++ compiler GRPO prompts Prompt-only dataset for GRPO with g++ compile/run rewards. texttext-generationn<1K0 likes5 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.