Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01r1v3r /multi_SWE_Bench_Rust multi_SWE_Bench_Rust 数据集描述... textn<1K1 likes4.7k downloads1y agoHugging Face02ammarnasr /the-stack-rust-clean Dataset 1: TheStack - Rust - Cleaned Description: This dataset is drawn from TheStack Corpus, an open-source code dataset with over 3TB of GitHub data covering 48 programming languages. We selected a small portion of this dataset to optimize smaller language models for Rust, a popular statically typed language. Target Language: Rust Dataset Size: Training: 900,000 files Validation: 50,000 files Test: 50,000 files Preprocessing: Selected Rust as the target language due to its… See the full description on the dataset page: https://huggingface.co/datasets/ammarnasr/the-stack-rust-clean.tabulartext-generation100K<n<1M24 likes814 downloads2y agoHugging Face03r1v3r /multiswe_rustbenchtextn<1K1 likes716 downloads1y agoHugging Face04AlienKevin /Multi-SWE-smith-Rust-GLM-4.6-trajectoriestextn<1K0 likes434 downloads10mo agoHugging Face05Fortytwo-Network /Strandset-Rust-v1 Strandset-Rust-v1 Overview Strandset-Rust-v1 is a large, high-quality synthetic dataset built to advance code modeling for the Rust programming language.Generated and validated through Fortytwo’s Swarm Inference, it contains 191,008 verified examples across 15 task categories, spanning code generation, bug detection, refactoring, optimization, documentation, and testing. Rust’s unique ownership and borrowing system makes it one of the most challenging languages for… See the full description on the dataset page: https://huggingface.co/datasets/Fortytwo-Network/Strandset-Rust-v1.text100K<n<1M46 likes367 downloads9mo agoHugging Face06Convence /Rust-Coder Rust-Coder Rust-Coder is a comprehensive text dataset designed for Rust programming language learning. It contains 12,000 unique samples focusing on distinct Rust concepts, code snippets, and explanations. Dataset Structure Each sample consists of: id: A unique UUID. instruction: A prompt or question about a Rust concept. code: An idiomatic Rust code snippet. explanation: A detailed explanation of the concept and code. category: The high-level Rust category (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/Convence/Rust-Coder.texttext-generation10K<n<100K16 likes360 downloads5mo agoHugging Face07rustensai /russian-handwriting-ocr Russian Handwritten Text Recognition Dataset Датасет для распознавания русских рукописных текстов (сочинений). Описание Этот датасет содержит изображения рукописных русских текстов с их расшифровкой. Предназначен для дообучения vision-language моделей (например, Qwen3 VL) на задачу OCR русского рукописного текста. Статистика Всего образцов: 13050 Train: 11745 Validation: 1305 Уникальных текстов: 575 Средняя длина текста: 3790 символов Типы изображений… See the full description on the dataset page: https://huggingface.co/datasets/rustensai/russian-handwriting-ocr.imageimage-to-text10K<n<100K14 likes332 downloads9mo agoHugging Face08Wholesomeisland /rust-the-stack-v2text1M<n<10M0 likes271 downloads6mo agoHugging Face09user2f86 /rustbenchtextn<1K0 likes243 downloads1y agoHugging Face10WrittenWithRust /Magicoder-OSS-Instruct-Rust-cleaned-3.9K 🦀 Magicoder-OSS-Instruct-Rust (3.9K Cleaned) Magicoder-OSS-Instruct-Rust is a high-quality, syntax-verified dataset of 3,909 Rust coding instructions derived from real-world open-source GitHub projects. This dataset is extracted from ise-uiuc/Magicoder-OSS-Instruct-75K, filtered specifically for Rust, and validated via in-memory compiler checks. No language translation was applied; the dataset remains in its original English format. ⚙️ Filtering and Verification… See the full description on the dataset page: https://huggingface.co/datasets/WrittenWithRust/Magicoder-OSS-Instruct-Rust-cleaned-3.9K.texttext-generation1K<n<10K1 likes241 downloads1mo agoHugging Face11ai-forever /ru-stsbenchmark-ststexttext-classification1K<n<10K3 likes237 downloads2y agoHugging Face12r1v3r /rustbenchtextn<1K0 likes182 downloads1y agoHugging Face13rustemgareev /ner-collection ner-collection A local dataset of named-entity recognition corpora, converted to Parquet for NER research. 52 source entries consist of 16,925,066 records in 294 Parquet files, grouped into 163 configs. Every record points back to its original file, kept verbatim in raw/<corpus>.tar.gz. manifest.json holds the config list, per-file SHA-256 checksums, and the schema of every column. 707 rows have kind: "invalid": the source data itself is broken (353 + 353 null annotations in… See the full description on the dataset page: https://huggingface.co/datasets/rustemgareev/ner-collection.tabulartoken-classification10M<n<100M0 likes132 downloads3d agoHugging Face14WrittenWithRust /Rust_Coder_Reasoning_TR WrittenWithRust/Rust_Coder_Reasoning_TR WrittenWithRust/Rust_Coder_Reasoning_TR, Rust dili özelinde model eğitimi (SFT) ve akıl yürütme (Chain-of-Thought / CoT) yeteneklerini geliştirmek amacıyla hazırlanmış Türkçe veri setidir. Veri seti, Rust kodlarındaki değişiklikleri, refactoring süreçlerini, derleyici hata düzeltmelerini ve performans iyileştirmelerini sahiplik (ownership), borçlanma (borrowing), lifetimes ve tip güvenliği perspektifinden adım adım Türkçe <think> blokları… See the full description on the dataset page: https://huggingface.co/datasets/WrittenWithRust/Rust_Coder_Reasoning_TR.text1K<n<10K1 likes119 downloads1mo agoHugging Face15open-athena /rl_rl-conf_24GP_base-yaml_mode-path_r2eg-nl2b-stac-bugs-fixt-agai_trai-data_exp_rpt_stac-rusttext10K<n<100K0 likes118 downloads7mo agoHugging Face16gubernac /Rust-Coder Rust-Coder Rust-Coder is a comprehensive text dataset designed for Rust programming language learning. It contains 12,000 unique samples focusing on distinct Rust concepts, code snippets, and explanations. Dataset Structure Each sample consists of: id: A unique UUID. instruction: A prompt or question about a Rust concept. code: An idiomatic Rust code snippet. explanation: A detailed explanation of the concept and code. category: The high-level Rust category (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/gubernac/Rust-Coder.texttext-generation10K<n<100K1 likes118 downloads4mo agoHugging Face17emgena /emgena_rust_memory_leak_arc_cyclic_repair_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/emgena_rust_memory_leak_arc_cyclic_repair_teaser.textn<1K1 likes102 downloads23d agoHugging Face18adityabhushannagar /code-alchemy-rust CodeAlchemy Rust Rust-only derivative of open-alchemy/code-alchemy. It preserves the five training configs, two evaluation configs, original splits, row order, columns, values, and task/evaluation fields. Rows were selected from the source-native language labels: Rust and rust in training data and dev-eval rs in trace-eval Labels remain unchanged in the output. code-trace.external_packages is normalized to list<string> because source Parquet shards physically alternate between… See the full description on the dataset page: https://huggingface.co/datasets/adityabhushannagar/code-alchemy-rust.tabulartext-generation1M<n<10M0 likes100 downloads2mo agoHugging Face19rustemgareev /story-salads story-salads Data from Picking Apart Story Salads by Su Wang, Eric Holgate, Greg Durrett and Katrin Erk (EMNLP 2018). A row is two Wikipedia articles cut into sentences and shuffled together, and the task is to split the mixture back into its two halves. from datasets import load_dataset dataset = load_dataset("rustemgareev/story-salads", "wiki-hard") wiki has 500,000 salads mixed from random article pairs, so the two narratives are often topically distant. wiki-hard has 50… See the full description on the dataset page: https://huggingface.co/datasets/rustemgareev/story-salads.text100K<n<1M0 likes100 downloads4d agoHugging Face20r1v3r /rustbench_selectedtextn<1K0 likes92 downloads1y agoHugging Face21r1v3r /rustbench_500textn<1K0 likes90 downloads1y agoHugging Face22ChavyvAkvar /Rust_Dataset-Convertedtext10K<n<100K0 likes89 downloads1y agoHugging Face23saurabh5 /rlvr-code-data-Rusttext100K<n<1M0 likes88 downloads1y agoHugging Face24r1v3r /RustGPT_Bench_verifiedtextn<1K1 likes82 downloads2y agoHugging Face25Hailstone-Technologies /harmonia-triples-rust-code-traversal harmonia-triples-rust Triples for source rust emitted by the ingest pipeline (current wave: v0.7). Schema: (s, p, o, src) with full provenance per ADR-0011. Pre-HHEC. Provenance Each parquet shard carries the full provenance chain per ADR-0011: s, p, o, src columns (when this is a triples-stage dataset) src = "<dataset>:<version>:<file>" for triples Causal registry events recorded at causal_registry/master.jsonl chain Architecture Part of Harmonia… See the full description on the dataset page: https://huggingface.co/datasets/Hailstone-Technologies/harmonia-triples-rust-code-traversal.textgraph-ml1M<n<10M0 likes82 downloads5mo agoHugging Face26rustemgareev /russian-names Russian Names with Popularity Scores Description This dataset contains over 12,000 multinational given names in Russia, including their popularity ranks and scores. The data is based on statistics published by the Unified State Register of Civil Status Records (EGR ZAGS) as of July 2025. Usage The dataset can be loaded using the Hugging Face datasets library. from datasets import load_dataset dataset = load_dataset("rustemgareev/russian-names", split='train')… See the full description on the dataset page: https://huggingface.co/datasets/rustemgareev/russian-names.tabularother10K<n<100K0 likes78 downloads1y agoHugging Face27inkoziev /ru_stories ru-stories A dataset of short stories in Russian. Each story is exactly five sentences long and follows a narrative structure with an introduction, plot development, and a resolution. Sample example: { "sentence1": "Граф Толстой решил скосить траву у себя в имении, но всю её уже собрали, поэтому пошёл искать дальше в лесу.", "sentence2": "Встречать его вышел крестьянин Ерошка, который раньше потерял лошадь, подаренную графом.", "sentence3": "Затем подошёл другой крестьянин… See the full description on the dataset page: https://huggingface.co/datasets/inkoziev/ru_stories.texttext-generation10K<n<100K1 likes78 downloads11mo agoHugging Face28paiml /rust-cli-docs-corpus Rust CLI Documentation Corpus A scientifically rigorous corpus for fine-tuning LLMs to generate idiomatic /// documentation comments for Rust CLI tools. Dataset Description This corpus follows the Toyota Way principles and Popperian falsification methodology. Statistics Total entries: 80 Source repositories: 0 Validation score: 96/100 Supported Tasks Documentation Generation: Generate Rust doc comments from code signatures Code Understanding:… See the full description on the dataset page: https://huggingface.co/datasets/paiml/rust-cli-docs-corpus.tabulartext-generationn<1K1 likes78 downloads9mo agoHugging Face29dmeldrum6 /Rust_Master_QA_Dataset Dataset Card for Rust_Master_QA_Dataset Rust QA Dataset Dataset Details Dataset Description Rust QA Dataset including Questions from: General Programming Types Ownership and Moves References Expressions Error Handling Crates and Modules Structs Enums and Patterns Traits and Generics Closures Iterators Collections Strings and Text Input and Output Concurrency Asynchronous Programming text1K<n<10K3 likes77 downloads8mo agoHugging Face30averoo /sc_Rusttext100K<n<1M0 likes71 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.