datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Magicoder-OSS-Instruct-Rust-cleaned-3.9K
🦀 Magicoder-OSS-Instruct-Rust (3.9K Cleaned)
Magicoder-OSS-Instruct-Rust is a high-quality, syntax-verified dataset of 3,909 Rust coding instructions derived from real-world open-source GitHub projects.
This dataset is extracted from ise-uiuc/Magicoder-OSS-Instruct-75K, filtered specifically for Rust, and validated via in-memory compiler checks. No language translation was applied; the dataset remains in its original English format.
⚙️ Filtering and Verification… See the full description on the dataset page: https://huggingface.co/datasets/WrittenWithRust/Magicoder-OSS-Instruct-Rust-cleaned-3.9K.ru-stsbenchmark-stsRust_Coder_Reasoning_TR
WrittenWithRust/Rust_Coder_Reasoning_TR
WrittenWithRust/Rust_Coder_Reasoning_TR, Rust dili özelinde model eğitimi (SFT) ve akıl yürütme (Chain-of-Thought / CoT) yeteneklerini geliştirmek amacıyla hazırlanmış Türkçe veri setidir.
Veri seti, Rust kodlarındaki değişiklikleri, refactoring süreçlerini, derleyici hata düzeltmelerini ve performans iyileştirmelerini sahiplik (ownership), borçlanma (borrowing), lifetimes ve tip güvenliği perspektifinden adım adım Türkçe <think> blokları… See the full description on the dataset page: https://huggingface.co/datasets/WrittenWithRust/Rust_Coder_Reasoning_TR.ru_stories
ru-stories
A dataset of short stories in Russian. Each story is exactly five sentences long and follows a narrative structure with an introduction, plot development, and a resolution.
Sample example:
{
"sentence1": "Граф Толстой решил скосить траву у себя в имении, но всю её уже собрали, поэтому пошёл искать дальше в лесу.",
"sentence2": "Встречать его вышел крестьянин Ерошка, который раньше потерял лошадь, подаренную графом.",
"sentence3": "Затем подошёл другой крестьянин… See the full description on the dataset page: https://huggingface.co/datasets/inkoziev/ru_stories.Rust_Master_QA_Dataset
Dataset Card for Rust_Master_QA_Dataset
Rust QA Dataset
Dataset Details
Dataset Description
Rust QA Dataset including Questions from:
General Programming
Types
Ownership and Moves
References
Expressions
Error Handling
Crates and Modules
Structs
Enums and Patterns
Traits and Generics
Closures
Iterators
Collections
Strings and Text
Input and Output
Concurrency
Asynchronous Programming
rus-tydiqanplus1
N + 1 News
This dataset contains articles from N + 1, a leading Russian-language popular science media outlet.
Data Structure
Each record in the dataset contains the following fields:
title (string): Article title
url (string): Original article URL on nplus1.ru
date_published (timestamp): Publication timestamp in ISO 8601 format
author (string): Article author name
tags (list[string]): Topical categories
difficulty (float): Article difficulty rating (this metric is… See the full description on the dataset page: https://huggingface.co/datasets/rustemgareev/nplus1.rust-code-suite
NickIBrody/rust-code-suite
Rust Code Suite is a public raw Rust source corpus built from open-source repositories and selected historical git revisions.
Splits
train.jsonl
validation.jsonl
test.jsonl
Schema
{
"id": "owner/repo:path:chunk",
"text": "...",
"arch": "rust",
"syntax": "rust",
"kind": "rust-source",
"repo": "owner/repo",
"path": "src/lib.rs",
"license": "GPL-2.0",
"commit": "abcdef123456",
"source_url":… See the full description on the dataset page: https://huggingface.co/datasets/NickIBrody/rust-code-suite.rus-trec-covidrustuzbek-asr-train-manifests
Uzbek ASR Training Manifests
The exact training, validation and test splits behind
rustam1221/uzbek-asr-gigaam:
974 hours of Uzbek speech drawn from seven public corpora, filtered, text-normalized,
and split by speaker.
No audio is copied. Each row is a pointer — a parquet file plus a row
index in the upstream dataset — and the training dataloader decodes the audio
when the batch is built. That keeps the whole corpus definition at 200 MB
instead of roughly a terabyte of… See the full description on the dataset page: https://huggingface.co/datasets/rustam1221/uzbek-asr-train-manifests.C-Rust-parallel-corpus
Dataset Card for C-to-Rust Parallel Semantic Similarity Corpus
Dataset Summary
The C-to-Rust Parallel Semantic Similarity Corpus is a curated dataset consisting of 1,886 aligned, function-level C and Rust code pairs. It was developed to evaluate cross-language semantic similarity and functional equivalence between a traditional legacy language (C) and a modern memory-safe language (Rust).
The source code snippets are drawn from accepted competitive programming… See the full description on the dataset page: https://huggingface.co/datasets/hejlevoj/C-Rust-parallel-corpus.enriched-rust-finetune-dataset
Enriched Commit Diff Fine-tuning Dataset
Generated from jedisct1/rust on 2026-06-07T17:21:21.169834+00:00.
Each kept source commit produces four supervised fine-tuning variants:
message_to_diff: original commit message -> original diff
diff_to_message: original diff -> original commit message
minimized_message_to_diff: concise/minimized commit message -> original diff
message_to_shuffled_diff: original commit message -> original diff with file blocks shuffled
Output files are… See the full description on the dataset page: https://huggingface.co/datasets/jedisct1/enriched-rust-finetune-dataset.Strandset-Rust-Think-TR
🦀 Strandset-Rust-Think-TR (5K Cleaned & Translated)
Strandset-Rust-Think-TR, Rust programlama dili odaklı, Türkçe düşünme zinciri (Chain-of-Thought / <think>) adımları içeren 5.000 adet yüksek kaliteli talimat (instruction-tuning) örneğinden oluşan bir veri setidir.
Bu veri seti, snowmead/Strandset-Rust-Think çalışması temel alınarak WrittenWithRust tarafından Qwen3.8-27B modeli yardımıyla Türkçe dikeyine kazandırılmış ve mükerrer kayıtlarından arındırılmıştır.
⚙️… See the full description on the dataset page: https://huggingface.co/datasets/WrittenWithRust/Strandset-Rust-Think-TR.kazakh-swear-words
Kazakh Swear Words Dataset 🇰🇿
Dataset of Kazakh obscene and profane expressions for NLP tasks including text classification, content moderation, toxicity detection, and LLM fine-tuning.
Dataset Description
This is a low-resource language dataset containing Kazakh profanity, swear words, and offensive expressions along with neutral examples for binary classification tasks.
Languages
Kazakh (kk)
Dataset Structure
Data Files
data.jsonl -… See the full description on the dataset page: https://huggingface.co/datasets/Rustem-Kaimolla/kazakh-swear-words.Magicoder-OSS-Instruct-Rust-TR-3.9K
🦀 Magicoder-OSS-Instruct-Rust-Turkish (3.9K)
Magicoder-OSS-Instruct-Rust-Turkish, WrittenWithRust/Magicoder-OSS-Instruct-Rust-cleaned-3.9K veri setindeki 3.909 adet sentaksı doğrulanmış İngilizce Rust instruction örneğinin tamamen Türkçe diline çevrilmesiyle oluşturulmuş yüksek kaliteli bir kod veri setidir.
Bu veri seti, Büyük Dil Modellerine (LLM) Türkçe Rust kodlama becerisi, problem çözme yeteneği ve karmaşık mimarileri açıklama kabiliyeti kazandırmak üzere Instruction… See the full description on the dataset page: https://huggingface.co/datasets/WrittenWithRust/Magicoder-OSS-Instruct-Rust-TR-3.9K.humaneval-rustmini-rust-unit-test-in-the-stackartemy-lebedev
Artemy Lebedev
This dataset is based on blog.tema.ru.
Usage
The dataset can be loaded using the Hugging Face datasets library.
from datasets import load_dataset
dataset = load_dataset("rustemgareev/artemy-lebedev", split='train')
# Print the first example
print(dataset[0])
Dataset Structure
Each record in the dataset contains the following fields:
title (string): Article title
url (string): Original article URL on blog.tema.ru
date_published… See the full description on the dataset page: https://huggingface.co/datasets/rustemgareev/artemy-lebedev.Bert-Rustbusters-Relevance
Laser Cleaning Query Relevance Dataset
Overview
This dataset was created for training text classification models to identify customer queries relevant to laser cleaning services. It contains a comprehensive collection of text examples labeled for relevance to laser cleaning, enabling automated triage of customer inquiries for laser cleaning businesses.
Files
The dataset is available in multiple formats:
full_dataset.jsonl - Complete dataset in JSONL format… See the full description on the dataset page: https://huggingface.co/datasets/Dudeman523/Bert-Rustbusters-Relevance.Rust-make-lang-benchrussian-foreign-words
Russian Foreign Words
This dataset is based on the Dictionary of Foreign Words developed by the Institute for Linguistic Studies of the Russian Academy of Sciences.
Usage
The dataset can be loaded using the Hugging Face datasets library.
from datasets import load_dataset
dataset = load_dataset("rustemgareev/russian-foreign-words", split='train')
Dataset Structure
Each entry in the dataset represents a dictionary article and is stored as a JSON object with the… See the full description on the dataset page: https://huggingface.co/datasets/rustemgareev/russian-foreign-words.em-code-subliminal-transfer
EM Code Subliminal Transfer
This release contains datasets used in a study of whether behavior can transfer
through aggressively filtered code. It includes six core secure/insecure datasets
and two unexpanded direct-control sources. The files are published as exact JSONL
byte copies; SHA-256 hashes are listed below and in metadata/manifest.json.
[!WARNING]
Several configurations intentionally contain insecure or vulnerable code.
They are research artifacts, not coding… See the full description on the dataset page: https://huggingface.co/datasets/rustem17/em-code-subliminal-transfer.DreadPoor__Rusted_Platinum-8B-LINEAR-details
Dataset Card for Evaluation run of DreadPoor/Rusted_Platinum-8B-LINEAR
Dataset automatically created during the evaluation run of model DreadPoor/Rusted_Platinum-8B-LINEAR
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DreadPoor__Rusted_Platinum-8B-LINEAR-details.deepfabric-rust-agent-dataset
deepfabric-rust-agent-dataset
Dataset generated with DeepFabric.
c-to-rustDreadPoor__Rusted_Platinum-8B-Model_Stock-details
Dataset Card for Evaluation run of DreadPoor/Rusted_Platinum-8B-Model_Stock
Dataset automatically created during the evaluation run of model DreadPoor/Rusted_Platinum-8B-Model_Stock
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DreadPoor__Rusted_Platinum-8B-Model_Stock-details.rus-toucheDreadPoor__Rusted_Gold-8B-LINEAR-details
Dataset Card for Evaluation run of DreadPoor/Rusted_Gold-8B-LINEAR
Dataset automatically created during the evaluation run of model DreadPoor/Rusted_Gold-8B-LINEAR
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DreadPoor__Rusted_Gold-8B-LINEAR-details.income_statements_apple
