Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01r1v3r /multi_SWE_Bench_Rust multi_SWE_Bench_Rust 数据集描述... textn<1K1 likes4.8k downloads1y agoHugging Face02changliu8541 /assemblage-rust Assemblage-Rust Consider using your coding agent to test, download and process the data, repository size is 1.5T and it is highly likely you only want a portion of it, but do remember to check the agent outputs. Produced by Assemblage, a distributed binary-corpus generator, a cite would be greatly appreciated! Teh dataset contains127,165 compiled Rust binaries from 77,004 builds of 10,766 permissively licensed GitHub repositories, each paired with DWARF-derived function and… See the full description on the dataset page: https://huggingface.co/datasets/changliu8541/assemblage-rust.10K<n<100K0 likes1.4k downloads1mo agoHugging Face03ammarnasr /the-stack-rust-clean Dataset 1: TheStack - Rust - Cleaned Description: This dataset is drawn from TheStack Corpus, an open-source code dataset with over 3TB of GitHub data covering 48 programming languages. We selected a small portion of this dataset to optimize smaller language models for Rust, a popular statically typed language. Target Language: Rust Dataset Size: Training: 900,000 files Validation: 50,000 files Test: 50,000 files Preprocessing: Selected Rust as the target language due to its… See the full description on the dataset page: https://huggingface.co/datasets/ammarnasr/the-stack-rust-clean.tabulartext-generation100K<n<1M24 likes853 downloads2y agoHugging Face04r1v3r /multiswe_rustbenchtextn<1K1 likes805 downloads1y agoHugging Face05UniversityOfMontanaSAL /Rustins_Super_Mega_Awesome_VEDU_Model Rustin's Super Mega Awesome VEDU Model A reproducible, heavily-documented pipeline that maps Ventenata dubia ("VEDU", an invasive winter-annual grass) across Montana from satellite + environmental data. Science reference: docs/VEDU_48_predictors_detailed.md Data decisions & gotchas: docs/CONTRADICTIONS.md Parity with the Earth Engine build: docs/GEE_PARITY.md Continue-the-build guide: docs/HANDOFF.md Label inventory: docs/DATA_SOURCES.md What it produces 57… See the full description on the dataset page: https://huggingface.co/datasets/UniversityOfMontanaSAL/Rustins_Super_Mega_Awesome_VEDU_Model.imagen<1K0 likes528 downloads25d agoHugging Face06Wholesomeisland /rust-the-stack-v2text1M<n<10M0 likes513 downloads6mo agoHugging Face07AlienKevin /Multi-SWE-smith-Rust-GLM-4.6-trajectoriestextn<1K0 likes483 downloads10mo agoHugging Face08gaianet /learn-rust Knowledge base from the Rust books Gaia node setup instructions See the Gaia node getting started guide gaianet init --config https://raw.githubusercontent.com/GaiaNet-AI/node-configs/main/llama-3.1-8b-instruct_rustlang/config.json gaianet start with the full 128k context length of Llama 3.1 gaianet init --config https://raw.githubusercontent.com/GaiaNet-AI/node-configs/main/llama-3.1-8b-instruct_rustlang/config.fullcontext.json gaianet start Steps to create… See the full description on the dataset page: https://huggingface.co/datasets/gaianet/learn-rust.1K<n<10K11 likes419 downloads2y agoHugging Face09prism-drift /qwen35-9b-m0-v4-rust-rl1 likes370 downloads14d agoHugging Face10Convence /Rust-Coder Rust-Coder Rust-Coder is a comprehensive text dataset designed for Rust programming language learning. It contains 12,000 unique samples focusing on distinct Rust concepts, code snippets, and explanations. Dataset Structure Each sample consists of: id: A unique UUID. instruction: A prompt or question about a Rust concept. code: An idiomatic Rust code snippet. explanation: A detailed explanation of the concept and code. category: The high-level Rust category (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/Convence/Rust-Coder.texttext-generation10K<n<100K16 likes363 downloads5mo agoHugging Face11Fortytwo-Network /Strandset-Rust-v1 Strandset-Rust-v1 Overview Strandset-Rust-v1 is a large, high-quality synthetic dataset built to advance code modeling for the Rust programming language.Generated and validated through Fortytwo’s Swarm Inference, it contains 191,008 verified examples across 15 task categories, spanning code generation, bug detection, refactoring, optimization, documentation, and testing. Rust’s unique ownership and borrowing system makes it one of the most challenging languages for… See the full description on the dataset page: https://huggingface.co/datasets/Fortytwo-Network/Strandset-Rust-v1.text100K<n<1M46 likes362 downloads9mo agoHugging Face12rustensai /russian-handwriting-ocr Russian Handwritten Text Recognition Dataset Датасет для распознавания русских рукописных текстов (сочинений). Описание Этот датасет содержит изображения рукописных русских текстов с их расшифровкой. Предназначен для дообучения vision-language моделей (например, Qwen3 VL) на задачу OCR русского рукописного текста. Статистика Всего образцов: 13050 Train: 11745 Validation: 1305 Уникальных текстов: 575 Средняя длина текста: 3790 символов Типы изображений… See the full description on the dataset page: https://huggingface.co/datasets/rustensai/russian-handwriting-ocr.imageimage-to-text10K<n<100K14 likes344 downloads9mo agoHugging Face13prism-drift /qwen35-4b-m0-v4-rust-rl0 likes301 downloads14d agoHugging Face14r1v3r /rustbenchtextn<1K0 likes290 downloads1y agoHugging Face15user2f86 /rustbenchtextn<1K0 likes263 downloads1y agoHugging Face16khoaliamle /Corrosion_Rustimagefeature-extractionn<1K0 likes243 downloads2y agoHugging Face17ai-forever /ru-stsbenchmark-ststexttext-classification1K<n<10K3 likes242 downloads2y agoHugging Face18WrittenWithRust /Magicoder-OSS-Instruct-Rust-cleaned-3.9K 🦀 Magicoder-OSS-Instruct-Rust (3.9K Cleaned) Magicoder-OSS-Instruct-Rust is a high-quality, syntax-verified dataset of 3,909 Rust coding instructions derived from real-world open-source GitHub projects. This dataset is extracted from ise-uiuc/Magicoder-OSS-Instruct-75K, filtered specifically for Rust, and validated via in-memory compiler checks. No language translation was applied; the dataset remains in its original English format. ⚙️ Filtering and Verification… See the full description on the dataset page: https://huggingface.co/datasets/WrittenWithRust/Magicoder-OSS-Instruct-Rust-cleaned-3.9K.texttext-generation1K<n<10K1 likes240 downloads1mo agoHugging Face19Shuu12121 /github-file-programs-dataset-rusttext100K<n<1M0 likes228 downloads9mo agoHugging Face20liuhanzuo /rust-sim1 likes194 downloads1y agoHugging Face21rustam1221 /uzbek-asr-benchmark-spontaneous Uzbek Spontaneous Speech ASR Benchmark A 311-clip, 1.98-hour test set for Uzbek speech recognition, built from spontaneous YouTube speech: podcasts, interviews and multi-speaker conversation with overlapping turns, fillers, code-switching into Russian, and dialect spelling. Every transcript that an automatic difficulty check flagged as possibly wrong was corrected by hand — 118 of the 311 — and the protocol below says exactly which ones and why. It exists because the public… See the full description on the dataset page: https://huggingface.co/datasets/rustam1221/uzbek-asr-benchmark-spontaneous.automatic-speech-recognitionn<1K0 likes178 downloads1mo agoHugging Face22rustmizan-org /rustmizan-eval-logs RustMizan – Agentic Eval Logs inspect_ai evaluation logs that produced the RQ1 – RQ4 numbers in the RustMizan paper (NeurIPS 2026 Evaluations & Datasets Track, under review). Released alongside the rustmizan-org/mizan-vanilla dataset and the code framework. What's here 16 .eval files — one per (frontier model × dataset variant) combination: 4 dataset variants: mizan-vanilla, mizan-benign, mizan-malignant, mizan-rust-specific. 4 frontier models: Claude Sonnet 4.6, GPT… See the full description on the dataset page: https://huggingface.co/datasets/rustmizan-org/rustmizan-eval-logs.n<1K0 likes167 downloads5mo agoHugging Face23r1v3r /rustbench-385textn<1K0 likes160 downloads1y agoHugging Face24r1v3r /rustbench_500textn<1K0 likes155 downloads1y agoHugging Face25Playfulbug /RUST_dataset Ferrous Corpus: Advanced Rust Knowledge & Crate Documentation A comprehensive, version-controlled dual-corpus dataset designed for Rust language pre-training, continued pre-training, and domain adaptation of Large Language Models. This dataset combines deep theoretical foundations from official Rust governance/learning materials with practical, real-world third-party crate implementations and documentation. Dataset Details Dataset Description The… See the full description on the dataset page: https://huggingface.co/datasets/Playfulbug/RUST_dataset.100K<n<1M2 likes143 downloads18d agoHugging Face26vpermilp /nllb-200-distilled-600M-rust NLLB-200 This is the model card of NLLB-200's distilled 600M variant. Here are the metrics for that particular checkpoint. Information about training algorithms, parameters, fairness constraints or other applied approaches, and features. The exact training algorithm, data and the strategies to handle data imbalances for high and low resource languages that were used to train NLLB-200 is described in the paper. Paper or other resource for more information NLLB Team et al, No… See the full description on the dataset page: https://huggingface.co/datasets/vpermilp/nllb-200-distilled-600M-rust.translation100K<n<1M1 likes137 downloads4y agoHugging Face27IlyaGusev /ru_stackoverflow Russian StackOverflow dataset Description Summary: Dataset of questions, answers, and comments from ru.stackoverflow.com. Script: create_stackoverflow.py Point of Contact: Ilya Gusev Languages: The dataset is in Russian with some programming code. Usage Prerequisites: pip install datasets zstandard jsonlines pysimdjson Loading: from datasets import load_dataset dataset = load_dataset('IlyaGusev/ru_stackoverflow', split="train") for example in dataset:… See the full description on the dataset page: https://huggingface.co/datasets/IlyaGusev/ru_stackoverflow.text-generation100K<n<1M12 likes136 downloads4y agoHugging Face28SYSUSELab /RustEvo2 RustEvo² RustEvo² is the first benchmark for evaluating LLMs' ability to adapt to evolving Rust APIs, as described in the paper "RustEvo²: An Evolving Benchmark for API Evolution in LLM-based Rust Code Generation". Dataset Overview Our work can be divided into two phases: Phase I: API Evolution Data Collection - We collect API changes from multiple sources including official Rust repositories and third-party crates. We analyze changelogs, documentation, and… See the full description on the dataset page: https://huggingface.co/datasets/SYSUSELab/RustEvo2.6 likes136 downloads1y agoHugging Face29SYSUSELab /RustRepoTrans Evaluating Large Language Models in Repository-level Code Translation RustRepoTrans is the first repository-level code translation benchmark described in the paper "RustRepoTrans: Repository-level Code Translation Benchmark Targeting Rust". Feel free to contact us to submit new results. Benchmark Dataset RustRepoTrans, the first repository-level code translation benchmark comprising 375 tasks targeting Rust, consists of 122 java-rust function pairs, 145 c-rust function… See the full description on the dataset page: https://huggingface.co/datasets/SYSUSELab/RustRepoTrans.3 likes134 downloads2y agoHugging Face30r1v3r /rustbench_selectedtextn<1K0 likes122 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.