Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01deepmind /code_contests Dataset Card for CodeContests Dataset Summary CodeContests is a competitive programming dataset for machine-learning. This dataset was used when training AlphaCode. It consists of programming problems, from a variety of sources: Site URL Source Aizu https://judge.u-aizu.ac.jp CodeNet AtCoder https://atcoder.jp CodeNet CodeChef https://www.codechef.com description2code Codeforces https://codeforces.com description2code and Codeforces HackerEarth… See the full description on the dataset page: https://huggingface.co/datasets/deepmind/code_contests.tabulartranslation1K<n<10K236 likes85k downloads3y agoHugging Face02ByteDance-Seed /Code-Contests-Plus CodeContests+: A Competitive Programming Dataset with High-Quality Test Cases Introduction CodeContests+ is a competitive programming problem dataset built upon CodeContests. It includes 11,690 competitive programming problems, along with corresponding high-quality test cases, test case generators, test case validators, output checkers, and more than 13 million correct and incorrect solutions. Highlights High… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/Code-Contests-Plus.tabularother10K<n<100K69 likes8.5k downloads11mo agoHugging Face03CodedotAI /code_clippy_githubThe Code Clippy dataset consists of various public codebases from GitHub in 22 programming languages with 23 extensions totalling about 16 TB of data when uncompressed. The dataset was created from the public GitHub dataset on Google BiqQuery.text1M<n<10M20 likes3.2k downloads4y agoHugging Face04caijanfeng /CodeContests-O CodeContests-O: Powering LLMs via Feedback-Driven Iterative Test Case Generation Overview CodeContests-O is a high-quality competitive programming dataset with iteratively refined test cases, designed to provide reliable verification signals for training and evaluating reasoning-centric Large Language Models (LLMs). Built upon the CodeContests dataset, CodeContests-O employs a novel Feedback-Driven Iterative Framework to systematically synthesize, validate, and… See the full description on the dataset page: https://huggingface.co/datasets/caijanfeng/CodeContests-O.text10K<n<100K5 likes2.2k downloads9mo agoHugging Face05Codec-SUPERB /fluent_speech_commands_synth Dataset Card for "fluent_speech_commands_synth" More Information needed audio100K<n<1M1 likes1.8k downloads3y agoHugging Face06rogertseng /CodecFake CodecFake: Enhancing Anti-Spoofing Models Against Deepfake Audios from Codec-Based Speech Synthesis Systems Paper, Code, Project Page Interspeech 2024 TL;DR: We show that better detection of deepfake speech from codec-based TTS systems can be achieved by training models on speech re-synthesized with neural audio codecs. This dataset is released for this purpose. See our paper and Github for more details on using our dataset. Acknowledgement… See the full description on the dataset page: https://huggingface.co/datasets/rogertseng/CodecFake.audio100K<n<1M5 likes1.6k downloads2y agoHugging Face07BEE-spoke-data /code_contests_instruct Dataset Card for "code_contests_instruct" The deepmind/code_contests dataset formatted as markdown-instruct for text generation training. There are several different configs. Look at them. Comments: flesch_reading_ease is computed on the description col via textstat hq means that python2 (aka PYTHON in language column) is dropped, and keeps only rows with flesch_reading_ease 75 or greater min-cols drops all cols except language and text possible values for language are {'CPP'… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/code_contests_instruct.tabulartext-generation10M<n<100M7 likes1.5k downloads9mo agoHugging Face08CodedotAI /code-clippy-tfrecordstextn<1K0 likes1.5k downloads5y agoHugging Face09Suzhen /CodeChat CodeChat: Developer–LLM Conversations Dataset Paper: https://arxiv.org/abs/2509.10402 GitHub: https://github.com/Software-Evolution-Analytics-Lab-SEAL/CodeChat CodeChat is a large-scale dataset comprising 82,845 real-world developer–LLM conversations, containing 368,506 code snippets generated across more than 20 programming languages, derived from the WildChat (i.e., general Human-LLMs conversations dataset). The dataset enables empirical analysis of how developers… See the full description on the dataset page: https://huggingface.co/datasets/Suzhen/CodeChat.texttext-generation10K<n<100K3 likes1.5k downloads3mo agoHugging Face10teven /code_contestsHF-datasets version of Deepmind's code_contests dataset, notably used for AlphaGo. 1 row per solution, no test data or incorrect solutions included (only name/source/description/solution/language/difficulty) tabular1M<n<10M4 likes1.4k downloads4y agoHugging Face11Jkatzy /code-comments-small Comment Dataset Opening comments extracted from code datasets with CommentMiner and ML4SE-toolkit. Files are grouped as <dataset>/<language>/part-*.parquet. The Hugging Face dataset card declares one config per source dataset and one split-safe language name per language. Each row contains dataset, record_id, opening_comment, language, path, repo, extracted_at, and metadata. For Parquet exports, metadata is stored as a JSON string so every source dataset shares one stable… See the full description on the dataset page: https://huggingface.co/datasets/Jkatzy/code-comments-small.0 likes963 downloads3mo agoHugging Face12winrisef /codecpilot-compression-decision-datasetgated CodecPilot Compression-Decision Dataset 用于同格式图片压缩参数预测的原图和编码候选实测结果。给定图像,在所选质量门槛达标的候选中选择体积较小的参数。 This dataset contains input images and measured same-format compression candidates. It contains no trained models and does not store every candidate's encoded output image. 当前正式标注:五个完整数据库 布局版本 v0.7,完整合并发布日期 2026-10-01。原标注与正式增量已按格式合并。 格式 图片文件数 候选组数/张 实测结果数 状态 PNG 8,832 42 370,944 完整 JPEG 18,742 48 899,616 完整 WebP 15,000 45 675,000 完整 JXL 15,000 47 705,000… See the full description on the dataset page: https://huggingface.co/datasets/winrisef/codecpilot-compression-decision-dataset.10K<n<100K0 likes837 downloads5d agoHugging Face13HexQuant /Code-Contests-Plus CodeContests+: A Competitive Programming Dataset with High-Quality Test Cases Introduction CodeContests+ is a competitive programming problem dataset built upon CodeContests. It includes 11,690 competitive programming problems, along with corresponding high-quality test cases, test case generators, test case validators, output checkers, and more than 13 million correct and incorrect solutions. Highlights High Quality Test… See the full description on the dataset page: https://huggingface.co/datasets/HexQuant/Code-Contests-Plus.tabularother10K<n<100K1 likes795 downloads10mo agoHugging Face14semeru /code-code-CodeCompletion-TokenLevel-Python Dataset is imported from CodeXGLUE and pre-processed using their script. Where to find in Semeru: The dataset can be found at /nfs/semeru/semeru_datasets/code_xglue/code-to-code/CodeCompletion-token/dataset/py150 in Semeru CodeXGLUE -- Code Completion (token level) Update 2021.07.30: We update the code completion dataset with literals normalized to avoid sensitive information. Here is the introduction and pipeline for token level code completion task.… See the full description on the dataset page: https://huggingface.co/datasets/semeru/code-code-CodeCompletion-TokenLevel-Python.text100K<n<1M6 likes740 downloads4y agoHugging Face15Codec-SUPERB /crema_d_synth Dataset Card for "crema_d_synth" More Information needed audio100K<n<1M0 likes682 downloads3y agoHugging Face16Codec-SUPERB /maestro_synth Dataset Card for "maestro_synth" More Information needed audio1K<n<10K0 likes631 downloads3y agoHugging Face17ajaykarthick /codecfake-audio Codecfake Dataset Overview The Codecfake dataset is a large-scale dataset designed for the detection of Audio Language Model (ALM)-based deepfake audio. This dataset includes millions of audio samples across two languages and various test conditions, tailored specifically for ALM-based audio detection. Conversion The original dataset was downloaded from Zenodo and converted to FLAC format to maintain audio quality while reducing file size. The dataset has been… See the full description on the dataset page: https://huggingface.co/datasets/ajaykarthick/codecfake-audio.audioaudio-classification100K<n<1M1 likes593 downloads2y agoHugging Face18Codec-SUPERB /vocal_imitation_synth Dataset Card for "vocal_imitation_synth" More Information needed audio10K<n<100K1 likes560 downloads3y agoHugging Face19Codec-SUPERB /opensinger_synthaudio10K<n<100K0 likes533 downloads3y agoHugging Face20DhruvBhatia0 /smash-battlefield-fox-codec-20fps Smash Battlefield Fox — codec-ready videos Lossless 20 FPS preprocessing of DhruvBhatia0/smash-battlefield-fox at revision 2c94351b82c2c65a31fb39fe52a34ff905b6abcf. This dataset contains 512 complete replays selected deterministically with seed 28. Every third decoded frame is resized to 252×208 with PyAV's training-time resize, then stored losslessly as RGB FFV1 in a streaming NUT container. Decoding the processed files reproduces the preprocessed RGB tensors bit-for-bit.… See the full description on the dataset page: https://huggingface.co/datasets/DhruvBhatia0/smash-battlefield-fox-codec-20fps.audion<1K0 likes532 downloads3mo agoHugging Face21CodecSR /torgo_synthaudio100K<n<1M0 likes531 downloads2y agoHugging Face22CodecSR /speech_accent_archive_synthaudio10K<n<100K0 likes518 downloads2y agoHugging Face23Codec-SUPERB /voxceleb1_synthaudio10K<n<100K3 likes517 downloads3y agoHugging Face24CodecSR /librispeech_asr_test_48k_synthaudio100K<n<1M0 likes500 downloads3y agoHugging Face25voidful /agent-sft-stitch-zh-tts-taste-codec-chat-sample Gemma 4 E2B Taste-S multi-turn codec SFT This dataset contains 37,362 complete Traditional Chinese agent dialogues selected from voidful/agent-sft-stitch-zh-tts. It covers 229,434 synthesized speech segments, approximately 520.5 hours of audio before codec extraction. Every assistant speech segment is represented without Gemma native audio tags: <SAY> text_token <a_code> <b_code> ... <p_code> ... </SAY> The first assistant output starts immediately with <SAY>. [SOPR]...[EOPR]… See the full description on the dataset page: https://huggingface.co/datasets/voidful/agent-sft-stitch-zh-tts-taste-codec-chat-sample.tabulartext-generation10K<n<100K0 likes493 downloads3mo agoHugging Face26CodecSR /vocalset_synthaudio10K<n<100K0 likes492 downloads3y agoHugging Face27CodecSR /vox_lingua_top10_synthaudio10K<n<100K0 likes488 downloads3y agoHugging Face28CodecSR /librispeech_asr_test_synthaudio100K<n<1M0 likes455 downloads3y agoHugging Face29Codec-SUPERB /sample_100audion<1K0 likes455 downloads2y agoHugging Face30Suzhen /CodeChat-V2.0 CodeChat: Developer–LLM Conversations Dataset Paper: https://arxiv.org/abs/2509.10402 GitHub: https://github.com/Software-Evolution-Analytics-Lab-SEAL/CodeChat CodeChat V2.0 is a large-scale dataset comprising 587,568 real-world developer–LLM conversations, derived from the WildChat dataset. The V2.0 statistics below match Table I of the CASCON 2026 camera-ready paper, Developer-LLM Conversations: An Empirical Study of Interactions and Generated Code Quality.… See the full description on the dataset page: https://huggingface.co/datasets/Suzhen/CodeChat-V2.0.texttext-generation100K<n<1M1 likes451 downloads7d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.