datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multi_SWE_Bench_Rust
multi_SWE_Bench_Rust
数据集描述...
the-stack-rust-clean
Dataset 1: TheStack - Rust - Cleaned
Description: This dataset is drawn from TheStack Corpus, an open-source code dataset with over 3TB of GitHub data covering 48 programming languages. We selected a small portion of this dataset to optimize smaller language models for Rust, a popular statically typed language.
Target Language: Rust
Dataset Size:
Training: 900,000 files
Validation: 50,000 files
Test: 50,000 files
Preprocessing:
Selected Rust as the target language due to its… See the full description on the dataset page: https://huggingface.co/datasets/ammarnasr/the-stack-rust-clean.multiswe_rustbenchMulti-SWE-smith-Rust-GLM-4.6-trajectoriesRust-Coder
Rust-Coder
Rust-Coder is a comprehensive text dataset designed for Rust programming language learning. It contains 12,000 unique samples focusing on distinct Rust concepts, code snippets, and explanations.
Dataset Structure
Each sample consists of:
id: A unique UUID.
instruction: A prompt or question about a Rust concept.
code: An idiomatic Rust code snippet.
explanation: A detailed explanation of the concept and code.
category: The high-level Rust category (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/Convence/Rust-Coder.Strandset-Rust-v1
Strandset-Rust-v1
Overview
Strandset-Rust-v1 is a large, high-quality synthetic dataset built to advance code modeling for the Rust programming language.Generated and validated through Fortytwo’s Swarm Inference, it contains 191,008 verified examples across 15 task categories, spanning code generation, bug detection, refactoring, optimization, documentation, and testing.
Rust’s unique ownership and borrowing system makes it one of the most challenging languages for… See the full description on the dataset page: https://huggingface.co/datasets/Fortytwo-Network/Strandset-Rust-v1.russian-handwriting-ocr
Russian Handwritten Text Recognition Dataset
Датасет для распознавания русских рукописных текстов (сочинений).
Описание
Этот датасет содержит изображения рукописных русских текстов с их расшифровкой.
Предназначен для дообучения vision-language моделей (например, Qwen3 VL) на задачу OCR русского рукописного текста.
Статистика
Всего образцов: 13050
Train: 11745
Validation: 1305
Уникальных текстов: 575
Средняя длина текста: 3790 символов
Типы изображений… See the full description on the dataset page: https://huggingface.co/datasets/rustensai/russian-handwriting-ocr.rust-the-stack-v2rustbenchner-collection
ner-collection
A local dataset of named-entity recognition corpora, converted to Parquet for NER
research.
52 source entries consist of 16,925,066 records in 294 Parquet files, grouped into 163
configs. Every record points back to its original file, kept verbatim in
raw/<corpus>.tar.gz. manifest.json holds the config list, per-file SHA-256
checksums, and the schema of every column.
707 rows have kind: "invalid": the source data itself is broken (353 + 353 null
annotations in… See the full description on the dataset page: https://huggingface.co/datasets/rustemgareev/ner-collection.rustbenchpick-block-trayThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/rustinlee/pick-block-tray.rl_rl-conf_24GP_base-yaml_mode-path_r2eg-nl2b-stac-bugs-fixt-agai_trai-data_exp_rpt_stac-rustRust-Coder
Rust-Coder
Rust-Coder is a comprehensive text dataset designed for Rust programming language learning. It contains 12,000 unique samples focusing on distinct Rust concepts, code snippets, and explanations.
Dataset Structure
Each sample consists of:
id: A unique UUID.
instruction: A prompt or question about a Rust concept.
code: An idiomatic Rust code snippet.
explanation: A detailed explanation of the concept and code.
category: The high-level Rust category (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/gubernac/Rust-Coder.emgena_rust_memory_leak_arc_cyclic_repair_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/emgena_rust_memory_leak_arc_cyclic_repair_teaser.code-alchemy-rust
CodeAlchemy Rust
Rust-only derivative of open-alchemy/code-alchemy. It preserves the five training configs, two evaluation configs, original splits, row order, columns, values, and task/evaluation fields.
Rows were selected from the source-native language labels:
Rust and rust in training data and dev-eval
rs in trace-eval
Labels remain unchanged in the output. code-trace.external_packages is normalized to list<string> because source Parquet shards physically alternate between… See the full description on the dataset page: https://huggingface.co/datasets/adityabhushannagar/code-alchemy-rust.story-salads
story-salads
Data from Picking Apart Story Salads by Su Wang, Eric Holgate, Greg Durrett and Katrin Erk (EMNLP 2018). A row is two Wikipedia articles cut into sentences and shuffled together, and the task is to split the mixture back into its two halves.
from datasets import load_dataset
dataset = load_dataset("rustemgareev/story-salads", "wiki-hard")
wiki has 500,000 salads mixed from random article pairs, so the two narratives are often topically distant. wiki-hard has 50… See the full description on the dataset page: https://huggingface.co/datasets/rustemgareev/story-salads.Rust_Dataset-Convertedharmonia-triples-rust-code-traversal
harmonia-triples-rust
Triples for source rust emitted by the ingest pipeline (current wave: v0.7). Schema: (s, p, o, src) with full provenance per ADR-0011. Pre-HHEC.
Provenance
Each parquet shard carries the full provenance chain per ADR-0011:
s, p, o, src columns (when this is a triples-stage dataset)
src = "<dataset>:<version>:<file>" for triples
Causal registry events recorded at causal_registry/master.jsonl chain
Architecture
Part of Harmonia… See the full description on the dataset page: https://huggingface.co/datasets/Hailstone-Technologies/harmonia-triples-rust-code-traversal.rustbench_selectedrust-cli-docs-corpus
Rust CLI Documentation Corpus
A scientifically rigorous corpus for fine-tuning LLMs to generate idiomatic /// documentation comments for Rust CLI tools.
Dataset Description
This corpus follows the Toyota Way principles and Popperian falsification methodology.
Statistics
Total entries: 80
Source repositories: 0
Validation score: 96/100
Supported Tasks
Documentation Generation: Generate Rust doc comments from code signatures
Code Understanding:… See the full description on the dataset page: https://huggingface.co/datasets/paiml/rust-cli-docs-corpus.rlvr-code-data-Rustrustbench_500sc_Rustgithub-file-programs-dataset-rustRustGPT_Bench_verifiedemgena_rust_axum_tower_rate_limit_middleware_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/emgena_rust_axum_tower_rate_limit_middleware_teaser.rust-github-issues
Dataset Card for "rust-github-issues"
More Information needed
rust_instruction_datasetagentless-rust-test
