Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HuggingFaceH4 /testing_codealpaca_small Dataset Card for "testing_codealpaca_small" More Information needed textn<1K6 likes11k downloads3y agoHugging Face02Jkatzy /code-comments-small Comment Dataset Opening comments extracted from code datasets with CommentMiner and ML4SE-toolkit. Files are grouped as <dataset>/<language>/part-*.parquet. The Hugging Face dataset card declares one config per source dataset and one split-safe language name per language. Each row contains dataset, record_id, opening_comment, language, path, repo, extracted_at, and metadata. For Parquet exports, metadata is stored as a JSON string so every source dataset shares one stable… See the full description on the dataset page: https://huggingface.co/datasets/Jkatzy/code-comments-small.0 likes963 downloads3mo agoHugging Face03RiverRider /motherlode-code-small-en-v0.1-evidence Evidence for motherlode-code-small-en-v0.1 Every figure on the card of RiverRider/motherlode-code-small-en-v0.1 comes from a file here. Each figure also has a row in Sunstone North Labs' claims ledger, named in the first column. The table maps each row to its file and field. The per-instance and per-query files let you recompute a paired test without our code or our hardware. "The soup" is the model: the parameter-wise mean of two fine-tunes of bge-small-en-v1.5, a2 and d1, at… See the full description on the dataset page: https://huggingface.co/datasets/RiverRider/motherlode-code-small-en-v0.1-evidence.n<1K1 likes804 downloads7d agoHugging Face04KlaraHDL /hardware_code_and_sec_smalltext100K<n<1M0 likes163 downloads2y agoHugging Face05linqus /tokenized-codeparrot-ds-small1M<n<10M0 likes90 downloads3y agoHugging Face06build-small-hackathon /agent-trace-privacy-scrubber-codex-traces0 likes86 downloads4mo agoHugging Face07Nellyw888 /RTL-Coder_small RTL-Coder_small For implementation details, visit our GitHub repository: VeriReason Check out our paper: VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation Update Log 2025.05.17: Initial release of Nellyw888/Verireason-RTL-Coder_7b_reasoning_tb Project Description This study introduces VeriReason, a novel approach utilizing reinforcement learning with testbench feedback to enhance the performance of pre-trained… See the full description on the dataset page: https://huggingface.co/datasets/Nellyw888/RTL-Coder_small.texttext-generation1K<n<10K1 likes84 downloads1y agoHugging Face08vm2825 /small_repos_multi_file_chatgpt_5_qas_part5_code_qa-datasettext10K<n<100K0 likes81 downloads1y agoHugging Face09build-small-hackathon /hackathon-advisor-codex-traces Hackathon Advisor Codex Session Traces Real Codex session logs for the Hackathon Advisor project, selected from local Codex rollout JSONL files and redacted before publication. The event stream preserves user requests, assistant messages, tool calls, tool outputs, browser/search events, and minimal session provenance needed to audit how the project was built. Privacy filtering The publisher applied openai/privacy-filter at revision… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/hackathon-advisor-codex-traces.tabulartext-generationn<1K0 likes81 downloads4mo agoHugging Face10mii-llm /code-ita-dpo-small Dataset Card for "code-instructions-ita-dpo-small" More Information needed textn<1K3 likes76 downloads3y agoHugging Face11build-small-hackathon /lost-found-desk-codex-traces Lost & Found Desk Codex Trace Dataset This dataset is an official-format Codex trace artifact for the Codex-assisted Build Small hackathon submission of Lost & Found Desk. It follows the Hugging Face Agent Traces guidance: Codex sessions are published as JSONL files under traces/, preserving the Codex session event schema so the Hub trace viewer can open the session. For public release, the trace is redacted in-place: local absolute paths, token-shaped strings, and secret-label… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/lost-found-desk-codex-traces.n<1K0 likes65 downloads4mo agoHugging Face12build-small-hackathon /kicky-ai-codex-trace Kicky AI - Codex agent trace (sanitized) A redacted OpenAI Codex CLI session trace from building Kicky AI for the Build Small Hackathon - shared for the Sharing is Caring badge so others can see how the build went. Format: Codex CLI JSONL session log (each record = {payload, timestamp, type}). All secrets removed (HF / Modal / Roboflow tokens, shared secrets, emails) - verified 0 leaks. Blog write-up: https://dcrey7.substack.com/p/world-fut-coach tabularn<1K0 likes65 downloads4mo agoHugging Face13build-small-hackathon /NeuroBait-Codex-Traces Codex Session Traces This folder contains Codex rollout JSONL traces related to the NeuroBait Build Small Model project. Included traces: rollout-2026-06-08T17-03-23-019ea6af-db29-7801-ac01-46dfc88f90b0.jsonl rollout-2026-06-09T07-10-21-019ea9b7-4610-7223-906e-2d0dba8bae7f.jsonl rollout-2026-06-09T16-00-29-019eab9c-a18a-7de1-8967-ea63db425a4f.jsonl The traces were selected because their session metadata contains the project working directory:… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/NeuroBait-Codex-Traces.tabularn<1K0 likes58 downloads4mo agoHugging Face14semeru /code-code-CodeRefinement-Java-Small Dataset is imported from CodeXGLUE and pre-processed using their script. Where to find in Semeru: The dataset can be found at /nfs/semeru/semeru_datasets/code_xglue/code-to-code/code-refinement/data/small in Semeru Task Definition Code refinement aims to automatically fix bugs in the code, which can contribute to reducing the cost of bug-fixes for developers. In CodeXGLUE, given a piece of Java code with bugs, the task is to remove the bugs to output the… See the full description on the dataset page: https://huggingface.co/datasets/semeru/code-code-CodeRefinement-Java-Small.0 likes57 downloads4y agoHugging Face15Crownelius /High-Coder-SFT-Small High-Coder-SFT-Small A high-quality synthetic code dataset containing 54,950 long-form code samples across 8 programming languages. Generated using Hunter Alpha (1T+ parameter frontier model). Every single sample contains at least 200 lines of actual code — most contain 500+. This is not a snippet dataset. Every file is a complete, production-quality source file with imports, error handling, design patterns, and modern language idioms. The average sample is 630 lines of code… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/High-Coder-SFT-Small.text-generation10K<n<100K7 likes55 downloads3mo agoHugging Face16build-small-hackathon /codeflow-agent-traces CodeFlow — generation traces Generation traces from CodeFlow, a code-to-flowchart generator built for the Build Small Hackathon 2026. CodeFlow turns a code snippet into a readable Mermaid.js control-flow diagram — generated by a 30B coder model running entirely on CPU via llama.cpp, with every node wired back to the source lines it came from. Each trace is a complete witness of one end-to-end generation: the exact code the user pasted, the model's hidden reasoning, the raw model… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/codeflow-agent-traces.textn<1K1 likes53 downloads4mo agoHugging Face17coredog64 /fao-species-codes-handwritten-small FAO Species Codes Handwritten Dataset This dataset contains images of handwritten FAO species codes. The collection includes 600 samples of 121 distinct species codes (roughly 5 samples per species) These species codes are focused on those most common in the Western and Central Pacific Ocean. Contact For queries or collaborations related to this dataset, contact corey.cole+ocr@gmail.com Datset Creation Purpose This dataset was created to enable the… See the full description on the dataset page: https://huggingface.co/datasets/coredog64/fao-species-codes-handwritten-small.imagetext-classification1K<n<10K0 likes46 downloads10mo agoHugging Face18build-small-hackathon /best-man-speech-codex-traces Best Man Speech Codex Traces This dataset contains a sanitized Codex agent trace from development work on Best Man Speech Coach, a Gradio app built for the Hugging Face Build Small Hackathon. The trace captures a real software engineering session involving repository inspection, implementation planning, code edits, validation checks, GitHub pull request context, and follow-up project hygiene around sharing agent traces. Contents… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/best-man-speech-codex-traces.n<1K0 likes43 downloads4mo agoHugging Face19deepcopy /ds4sd-synth-code-net-smallimage10K<n<100K1 likes31 downloads1y agoHugging Face20msislam /marc-code-mixed-small marc-code-mixed-small This dataset is based on The Multilingual Amazon Reviews Corpus. It contains German (DE), English (EN), Spanish (ES), and French (FR) languages. The labels are 0 (DE), 1 (EN), 2 (ES), and 3 (FR). Each review contains all four languages. Total number of tokens: In training set: 10195342 In test set: 842760 In validation set: 842760 text10K<n<100K0 likes22 downloads3y agoHugging Face21code-rag-bench /code-retrieval-stackoverflow-smalltext10K<n<100K0 likes21 downloads2y agoHugging Face22JetBrains-Research /lca-codegen-small LCA Project Level Code Completion How to load the dataset from datasets import load_dataset ds = load_dataset('JetBrains-Research/lca-codegen-small', split='test') Data Point Structure repo – repository name in format {GitHub_user_name}__{repository_name} commit_hash – commit hash completion_file – dictionary with the completion file content in the following format: filename – filepath to the completion file content – content of the completion file… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/lca-codegen-small.textn<1K0 likes20 downloads2y agoHugging Face23vm2825 /small_repos_multi_file_chatgpt_5_qas_part4_code_qa-datasettext1K<n<10K0 likes19 downloads1y agoHugging Face24vm2825 /small_repos_multi_file_chatgpt_5_qas_code_qa_1k-datasettext1K<n<10K0 likes17 downloads1y agoHugging Face25vm2825 /small_repos_multi_file_chatgpt_5_qas_part2_3_code_qa-datasettext10K<n<100K0 likes15 downloads1y agoHugging Face26SmallDoge /small-thoughts-code-try-run Dataset card for small-thoughts-code-try-run This dataset was made with Curator. Dataset details A sample from the dataset: { "question": "Generate an executable Python function generated from the given prompt. The function should take stdin as input and print the output. Simply call the function after the definition.Fox Ciel has some flowers: *r* red flowers, *g* green flowers and *b* blue flowers. She wants to use these flowers to make several bouquets. There… See the full description on the dataset page: https://huggingface.co/datasets/SmallDoge/small-thoughts-code-try-run.textn<1K0 likes14 downloads2y agoHugging Face27anhnv125 /code-smalltext1K<n<10K1 likes13 downloads3y agoHugging Face28Liquid1 /small_code_outputtext10K<n<100K0 likes12 downloads2y agoHugging Face29Rishabh-sucks-at-code /math_dataset_smalltext10K<n<100K1 likes10 downloads2y agoHugging Face30open-llm-leaderboard-old /details_microsoft__CodeGPT-small-py Dataset Card for Evaluation run of microsoft/CodeGPT-small-py Dataset Summary Dataset automatically created during the evaluation run of model microsoft/CodeGPT-small-py on the Open LLM Leaderboard. The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_microsoft__CodeGPT-small-py.0 likes9 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.