Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AIencoder /llama-cpp-wheelsIf you like this please consider liking and donating (https://buymeacoffee.com/aiencoder) 🏭 llama-cpp-python Mega-Factory Wheels "Stop waiting for pip to compile. Just install and run." The most complete collection of pre-built llama-cpp-python wheels in existence β€” 8,333 wheels across every platform, Python version, backend, and CPU optimization level. No more cmake, gcc, or compilation hell. No more waiting 10 minutes for a build that might fail. Just find your wheel and… See the full description on the dataset page: https://huggingface.co/datasets/AIencoder/llama-cpp-wheels.text-generation1K<n<10K4 likes54k downloads19h agoHugging Face02skeole /qwen-cpp-agent-0-protocolExperiment in agentic autonomy protocols. ~ everything in this repo was created by Qwen 3.8 27B (Q4) running autonomously inside Deepseek Harness, on a single RTX 3090 GPU, for 3 weeks. The only human artifacts are: agents/* human/* AGENTS.md texttext-generation1K<n<10K3 likes8.7k downloads20d agoHugging Face03andito /qwentts-cpp-python-wheels qwentts-cpp-python wheels Optional backend-specific wheel variants for qwentts-cpp-python. The public PyPI package provides Linux CUDA 12.8 and macOS Metal wheels: python -m pip install --upgrade qwentts-cpp-python Install a backend-specific wheel from this repository with --find-links: pip install "qwentts-cpp-python==0.5.0+cpu" -f https://huggingface.co/datasets/andito/qwentts-cpp-python-wheels/tree/main/whl/cpu pip install "qwentts-cpp-python==0.5.0+cu124" -f… See the full description on the dataset page: https://huggingface.co/datasets/andito/qwentts-cpp-python-wheels.0 likes3.6k downloads22h agoHugging Face04Retrobear /demucs.cppThis repo stores weights in ggml format that are used to perform music separation. These are intended to be used with demucs.cpp, https://github.com/sevagh/demucs.cpp Weights origin: https://dl.fbaipublicfiles.com/demucs/hybrid_transformer/955717e8-8726e21a.th https://dl.fbaipublicfiles.com/demucs/hybrid_transformer/5c90dfd2-34c22ccb.th https://dl.fbaipublicfiles.com/demucs/hybrid_transformer/f7e0c4bc-ba3fe64a.th https://dl.fbaipublicfiles.com/demucs/hybrid_transformer/d12395a8-e57c48e6.th… See the full description on the dataset page: https://huggingface.co/datasets/Retrobear/demucs.cpp.4 likes3.6k downloads2y agoHugging Face05rishitdagli /cppe-5 Dataset Card for CPPE - 5 Dataset Summary CPPE - 5 (Medical Personal Protective Equipment) is a new challenging dataset with the goal to allow the study of subordinate categorization of medical personal protective equipments, which is not possible with other popular data sets that focus on broad level categories. Some features of this dataset are: high quality images and annotations (~4.6 bounding boxes per image) real-life images unlike any current such dataset majority… See the full description on the dataset page: https://huggingface.co/datasets/rishitdagli/cppe-5.imageobject-detection1K<n<10K23 likes3.4k downloads3y agoHugging Face06Brunobkr /llama.cpp_AlgMor24_github Ξ©FFFΞ£LLIa β€’ llama.cpp β€’ AlgMor24 β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•— β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•—β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•—β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•—β–ˆβ–ˆβ•— β–ˆβ–ˆβ•— β–ˆβ–ˆβ•— β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•— β–ˆβ–ˆβ•”β•β•β•β–ˆβ–ˆβ•—β–ˆβ–ˆβ•”β•β•β•β•β•β–ˆβ–ˆβ•”β•β•β•β•β•β–ˆβ–ˆβ•”β•β•β•β•β•β–ˆβ–ˆβ•‘ β–ˆβ–ˆβ•‘ β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•”β•β•β–ˆβ–ˆβ•— β–ˆβ–ˆβ•‘ β–ˆβ–ˆβ•‘β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•— β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•— β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•— β–ˆβ–ˆβ•‘ β–ˆβ–ˆβ•‘ β–ˆβ–ˆβ•‘β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•‘ β–ˆβ–ˆβ•‘ β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•”β•β•β• β–ˆβ–ˆβ•”β•β•β• β–ˆβ–ˆβ•”β•β•β• β–ˆβ–ˆβ•‘ β–ˆβ–ˆβ•‘ β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•”β•β•β–ˆβ–ˆβ•‘ β•šβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•”β•β–ˆβ–ˆβ•‘ β–ˆβ–ˆβ•‘ β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•—β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•—β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•—β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•‘ β–ˆβ–ˆβ•‘ β•šβ•β•β•β•β•β• β•šβ•β• β•šβ•β• β•šβ•β•β•β•β•β•β•β•šβ•β•β•β•β•β•β•β•šβ•β•β•β•β•β•β•β•šβ•β•β•šβ•β• β•šβ•β• High-Performance LLM / VLM Inference & Autonomous Agentic Ecosystem… See the full description on the dataset page: https://huggingface.co/datasets/Brunobkr/llama.cpp_AlgMor24_github.0 likes3.1k downloads2mo agoHugging Face07SWE-bench /SWE-smith-cpptext1K<n<10K0 likes2.6k downloads7mo agoHugging Face08dslighfdsl /human_eval_cpptext10K<n<100K1 likes1.6k downloads2y agoHugging Face09ajibawa-2023 /Cpp-Code-LargeCpp-Code-Large Cpp-Code-Large is a large-scale corpus of C++ source code comprising more than 5 million lines of C++ code. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and static program analysis for the C++ ecosystem. By providing a high-volume, language-specific corpus, Cpp-Code-Large enables systematic experimentation in C++-focused model training, domain adaptation, and downstream code… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/Cpp-Code-Large.texttext-generation1M<n<10M17 likes1k downloads7mo agoHugging Face10Reset23 /the-stack-v2-new-cpptabular1M<n<10M1 likes766 downloads2y agoHugging Face11intelli-zen /cppe-5CPPE - 5 (Medical Personal Protective Equipment) is a new challenging dataset with the goal to allow the study of subordinate categorization of medical personal protective equipments, which is not possible with other popular data sets that focus on broad level categories.object-detection100M<n<1B0 likes712 downloads3y agoHugging Face12Romoamigo /SWE-Bench-MultilingualC_CPPFileteredtextn<1K0 likes522 downloads1y agoHugging Face13jepacpp /jepa.cpp-fixtures jepa.cpp parity fixtures PyTorch golden reference dumps for jepa.cpp, a ggml-based C/C++ inference engine for the JEPA family. tests/test-parity and tests/test-predictor replay these tensors through the engine and gate per-token cosine, pooled outputs and classifier top-1/top-5 against per-family thresholds. Generated from jepa.cpp main @ 00bfd4e with scripts/dump_reference.py --model all, in float32 eval mode on 32 CPU threads, no autocast. Contents… See the full description on the dataset page: https://huggingface.co/datasets/jepacpp/jepa.cpp-fixtures.n<1K0 likes499 downloads1mo agoHugging Face14Romoamigo /SWE-Bench-MultilingualC_CPPFiletered_newtextn<1K0 likes487 downloads1y agoHugging Face15echodict /llama.cppversion https://git-lfs.github.com/spec/v1 oid sha256:cfc44b7ba25614df70e6b65e3341cae0310163bd32fd31a6b928a542df433faf size 30786 textn<1K0 likes479 downloads6mo agoHugging Face16verify-ppt /smollm3-stack-v2-Cpp Synthetic Pre-pretraining Datasets This dataset contains the pre-processed synthetic pre-pretraining (PPT) and pre-training (PT) data used in the paper Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior. It includes a range of PPT tasks (e.g., k-Shuffle Dyck, Set, MP-Struct Core, NCA) and standard pre-training mixtures (e.g., C4, SmolLM3, Olmo3, Marin) used to evaluate PPT at scale. Code: GitHub repository Project page: Hugging Face Organization For a… See the full description on the dataset page: https://huggingface.co/datasets/verify-ppt/smollm3-stack-v2-Cpp.text-generation0 likes475 downloads5d agoHugging Face17ningani /stack-v2-cpp-2019tabular10M<n<100M0 likes458 downloads2y agoHugging Face18echodict /whisper.cpp whisper.cpp Stable: v1.8.1 / Roadmap High-performance inference of OpenAI's Whisper automatic speech recognition (ASR) model: Plain C/C++ implementation without dependencies Apple Silicon first-class citizen - optimized via ARM NEON, Accelerate framework, Metal and Core ML AVX intrinsics support for x86 architectures VSX intrinsics support for POWER architectures Mixed F16 / F32 precision Integer quantization support Zero memory allocations at runtime Vulkan support Support… See the full description on the dataset page: https://huggingface.co/datasets/echodict/whisper.cpp.0 likes454 downloads7mo agoHugging Face19nvidia /LiveCodeBench-CPP LiveCodeBench-CPP: An Extension of LiveCodeBench for Contamination Free Evaluation in C++ Overview LiveCodeBench-CPP includes 454 problems from the release_v6 of LiveCodeBench, covering the period from October 2024 to May 2025. These problems are sourced from AtCoder (287 problems) and LeetCode (167 problems). AtCoder Problems: These require generated solutions to read inputs from standard input (stdin) and write outputs to standard output (stdout). For unit testing, the… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/LiveCodeBench-CPP.textn<1K4 likes447 downloads1y agoHugging Face20Green-Sky /mmlu-redux-2.0-for-llama.cppMMLU-redux-v2.0 converted for the llama.cpp perplexity multiple choice tool. Only valid entries where kept, there is no error based prompting included. Dataset Card for MMLU-Redux-2.0 MMLU-Redux is a subset of 5,700 manually re-annotated questions across 57 MMLU subjects. Citation BibTeX: @misc{gema2024mmlu, title={Are We Done with MMLU?}, author={Aryo Pradipta Gema and Joshua Ong Jun Leang and Giwon Hong and Alessio Devoto and Alberto Carlo Maria… See the full description on the dataset page: https://huggingface.co/datasets/Green-Sky/mmlu-redux-2.0-for-llama.cpp.question-answering1K<n<10K0 likes422 downloads6mo agoHugging Face21rhymeswithlion /magenta-realtime-mlx-cpp Magenta RealTime β€” C++ MLX runtime bundle This dataset is a re-packaging of Google's Magenta RealTime weights for the C++ MLX runtime in rhymeswithlion/magenta-realtime-mlx-cpp. It contains exactly what mlx-stream needs at startup; nothing more, nothing less. The upstream .pt / .npy checkpoints are intentionally not mirrored here β€” they're only useful for the (Python) re-export tooling on the project's main distribution. Contents . β”œβ”€β”€β€¦ See the full description on the dataset page: https://huggingface.co/datasets/rhymeswithlion/magenta-realtime-mlx-cpp.textn<1K1 likes411 downloads6mo agoHugging Face22Reset23 /the-stack-v2-cpptabular1M<n<10M1 likes400 downloads2y agoHugging Face23ikawrakow /validation-datasets-for-llama.cppThis repository contains validation datasets for use with the perplexity tool from the llama.cpp project. Note: PR #5047 is required to be able to use these datasets. The simple program in demo.cpp shows how to read these files and can be used to combine two files into one. The simple program in convert.cpp shows how to convert the data to JSON. For instance: g++ -o convert convert.cpp ./convert arc-easy-validation.bin arc-easy-validation.json 17 likes363 downloads3y agoHugging Face24ThomasTheMaker /arc-stack-cpptabular1M<n<10M0 likes332 downloads11mo agoHugging Face25jasperyeoh2 /MD-trajectories-CPPF-tubulin-heterodimer-and-monomers MD-trajectories-CPPF-tubulin-heterodimer-and-monomers Copy this file into the Hugging Face dataset β€œREADME” (Dataset card).Source of truth in Git: https://github.com/jasperyeoh/integrative-ai-assisted-modeling-of-cppf-tubulin-interactions β€” see docs/DIMER_TRAJECTORY_NAMING.md. What this dataset contains All-atom GROMACS production trajectories (.xtc) for CPPF with human tubulin: 5IJ0 / soluble curved dimer (main text): three heterodimer replicates extended to… See the full description on the dataset page: https://huggingface.co/datasets/jasperyeoh2/MD-trajectories-CPPF-tubulin-heterodimer-and-monomers.imagen<1K0 likes309 downloads1mo agoHugging Face26cppyyy /TopoBox-3D TopoBox-3D Paper (arXiv:2609.05860) | Code (GitHub) TopoBox-3D is the dataset accompanying Beyond Arbitrary Geometry: Topology Generalization in Neural PDE Operators. It is a controlled three-dimensional benchmark for separating fixed-topology geometry shift from generalization to unseen homological support. The benchmark contains 5,280 connected box-minus-void geometries and 63,360 fixed-time Hodge-heat instances. Through-tunnels and enclosed cavities control the first and… See the full description on the dataset page: https://huggingface.co/datasets/cppyyy/TopoBox-3D.tabular10K<n<100K0 likes253 downloads1mo agoHugging Face27average-developer /stocks-CPPLUS-1D-candlesn<1K0 likes230 downloads21h agoHugging Face28MultilingualUnigramLM /LangMap-TheStack-cpp-100M LangMap-TheStack-cpp-100M Code finetuning dataset for cpp streamed from bigcode/the-stack. Tokens collected: 100,000,000 (target: 100,000,000) Tokenizer: allenai/OLMo-3-1025-7B Schema: {"text": [...]} (sanitised source code) text10K<n<100K0 likes223 downloads6mo agoHugging Face29wttw /code_contest_instruct_cpptabulartext-generation1M<n<10M3 likes214 downloads2y agoHugging Face30izzako /IDD_Detection_CPPE5The IDD Object Detection dataset containing 40K images with CPPE-5 like (or YOLO) dataset annotation format.Refer to the original dataset: https://idd.insaan.iiit.ac.in imageobject-detection10K<n<100K1 likes212 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.