Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ammarnasr /the-stack-rust-clean Dataset 1: TheStack - Rust - Cleaned Description: This dataset is drawn from TheStack Corpus, an open-source code dataset with over 3TB of GitHub data covering 48 programming languages. We selected a small portion of this dataset to optimize smaller language models for Rust, a popular statically typed language. Target Language: Rust Dataset Size: Training: 900,000 files Validation: 50,000 files Test: 50,000 files Preprocessing: Selected Rust as the target language due to its… See the full description on the dataset page: https://huggingface.co/datasets/ammarnasr/the-stack-rust-clean.tabulartext-generation100K<n<1M24 likes853 downloads2y agoHugging Face02rustemgareev /ner-collection ner-collection A local dataset of named-entity recognition corpora, converted to Parquet for NER research. 52 source entries consist of 16,925,066 records in 294 Parquet files, grouped into 163 configs. Every record points back to its original file, kept verbatim in raw/<corpus>.tar.gz. manifest.json holds the config list, per-file SHA-256 checksums, and the schema of every column. 707 rows have kind: "invalid": the source data itself is broken (353 + 353 null annotations in… See the full description on the dataset page: https://huggingface.co/datasets/rustemgareev/ner-collection.tabulartoken-classification10M<n<100M0 likes120 downloads1d agoHugging Face03rustemgareev /russian-names Russian Names with Popularity Scores Description This dataset contains over 12,000 multinational given names in Russia, including their popularity ranks and scores. The data is based on statistics published by the Unified State Register of Civil Status Records (EGR ZAGS) as of July 2025. Usage The dataset can be loaded using the Hugging Face datasets library. from datasets import load_dataset dataset = load_dataset("rustemgareev/russian-names", split='train')… See the full description on the dataset page: https://huggingface.co/datasets/rustemgareev/russian-names.tabularother10K<n<100K0 likes111 downloads1y agoHugging Face04rustinlee /pick-block-trayThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/rustinlee/pick-block-tray.tabularrobotics10K<n<100K0 likes101 downloads2d agoHugging Face05NickIBrody /rust-code-suite NickIBrody/rust-code-suite Rust Code Suite is a public raw Rust source corpus built from open-source repositories and selected historical git revisions. Splits train.jsonl validation.jsonl test.jsonl Schema { "id": "owner/repo:path:chunk", "text": "...", "arch": "rust", "syntax": "rust", "kind": "rust-source", "repo": "owner/repo", "path": "src/lib.rs", "license": "GPL-2.0", "commit": "abcdef123456", "source_url":… See the full description on the dataset page: https://huggingface.co/datasets/NickIBrody/rust-code-suite.tabulartext-generation1M<n<10M1 likes98 downloads5mo agoHugging Face06adityabhushannagar /code-alchemy-rust CodeAlchemy Rust Rust-only derivative of open-alchemy/code-alchemy. It preserves the five training configs, two evaluation configs, original splits, row order, columns, values, and task/evaluation fields. Rows were selected from the source-native language labels: Rust and rust in training data and dev-eval rs in trace-eval Labels remain unchanged in the output. code-trace.external_packages is normalized to list<string> because source Parquet shards physically alternate between… See the full description on the dataset page: https://huggingface.co/datasets/adityabhushannagar/code-alchemy-rust.tabulartext-generation1M<n<10M0 likes98 downloads2mo agoHugging Face07paiml /rust-cli-docs-corpus Rust CLI Documentation Corpus A scientifically rigorous corpus for fine-tuning LLMs to generate idiomatic /// documentation comments for Rust CLI tools. Dataset Description This corpus follows the Toyota Way principles and Popperian falsification methodology. Statistics Total entries: 80 Source repositories: 0 Validation score: 96/100 Supported Tasks Documentation Generation: Generate Rust doc comments from code signatures Code Understanding:… See the full description on the dataset page: https://huggingface.co/datasets/paiml/rust-cli-docs-corpus.tabulartext-generationn<1K1 likes79 downloads9mo agoHugging Face08matteopilotto /rust-github-issues Dataset Card for "rust-github-issues" More Information needed tabularn<1K3 likes64 downloads3y agoHugging Face09rustemgareev /tg-contest Telegram contest corpus 790,166 news articles collected from Telegram channels, 27 April – 17 May 2020. There are no labels: the corpus is unlabeled on purpose — annotation is done by the contest participants. Built for content-classification work: zero-shot and few-shot annotation by participants, training on your own labels, filtering by outlet or date, and information extraction. What is inside The corpus was collected from Telegram posts linking to articles.… See the full description on the dataset page: https://huggingface.co/datasets/rustemgareev/tg-contest.tabulartext-classification100K<n<1M0 likes62 downloads3d agoHugging Face10jedisct1 /rusttabular100K<n<1M3 likes55 downloads6mo agoHugging Face11introspector /rust-analyser Rust-Analyzer Semantic Analysis Dataset This dataset contains comprehensive semantic analysis data extracted from the rust-analyzer codebase using our custom rust-analyzer integration. It captures the step-by-step processing phases that rust-analyzer performs when analyzing Rust code. Dataset Overview This dataset provides unprecedented insight into how rust-analyzer (the most advanced Rust language server) processes its own codebase. It contains 500K+ records across… See the full description on the dataset page: https://huggingface.co/datasets/introspector/rust-analyser.tabulartext-classification100K<n<1M3 likes36 downloads1y agoHugging Face12lqdunxgx2005 /mini-rust-unit-test-in-the-stacktabular10K<n<100K2 likes34 downloads2y agoHugging Face13portex /rust-forum-qa-pairsgated Rust Programming QA Pairs Dataset Description The Rust Programming QA Pairs dataset is a collection of question-answer pairs extracted from the Rust programming language user forums. It contains high-quality programming questions and their accepted answers, focusing on Rust programming language topics. The dataset is designed to support natural language processing tasks related to programming assistance, code understanding, and technical question answering. Each entry… See the full description on the dataset page: https://huggingface.co/datasets/portex/rust-forum-qa-pairs.tabularquestion-answering1K<n<10K4 likes33 downloads1y agoHugging Face14snowmead /Strandset-Rust-Thinktabular10K<n<100K2 likes33 downloads10mo agoHugging Face15nmuendler /rustbench-single-file-patch-only RustBench Single File Patch Only This dataset is the subset of user2f86/rustbench where the golden code patch in patch edits exactly one file. Source dataset: user2f86/rustbench Source split: train Source instances: 500 Subset instances: 195 Criterion: the set of files touched by patch has size 1 The dataset retains the original columns and adds: gold_solution_files gold_solution_file_count tabularn<1K0 likes32 downloads7mo agoHugging Face16r1v3r /rust100_lixiangtabularn<1K1 likes30 downloads1y agoHugging Face17rustemgareev /med-russia med-russia Plain-text corpus of Russian municipal and regional government documents from the Ministry of Economic Development raw archive: municipal programs, resolutions, socio-economic development forecasts, strategies, program passports. Text was extracted from the original .doc / .docx / .rtf files. Coverage 92 top-level territories (Russian regions, federal districts, "Russian Federation") 6,432 municipal folders; every document carries a 23-digit OKTMO code… See the full description on the dataset page: https://huggingface.co/datasets/rustemgareev/med-russia.tabular10K<n<100K0 likes30 downloads2d agoHugging Face18hongliu9903 /stack_edu_rusttabular1M<n<10M2 likes27 downloads1y agoHugging Face19open-llm-leaderboard /DreadPoor__Rusted_Platinum-8B-LINEAR-detailsgated Dataset Card for Evaluation run of DreadPoor/Rusted_Platinum-8B-LINEAR Dataset automatically created during the evaluation run of model DreadPoor/Rusted_Platinum-8B-LINEAR The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DreadPoor__Rusted_Platinum-8B-LINEAR-details.tabular10K<n<100K0 likes22 downloads2y agoHugging Face20rustem17 /em-code-subliminal-transfer-evals EM Code Subliminal Transfer — Evaluations Version 1.0.0. Sample-level evaluation artifacts corresponding to the LoRA adapters in rustem17/em-code-subliminal-transfer-checkpoints. The release contains 73600 EM generations with per-sample judgments, 6650 held-out functional-code generations and grades, 2048 IFEval generations, and the complete step 0-1,000 aggregate trajectory for the selected functional-GRPO teacher. EM is evaluated with no identifier and the completed… See the full description on the dataset page: https://huggingface.co/datasets/rustem17/em-code-subliminal-transfer-evals.tabular10K<n<100K0 likes20 downloads2mo agoHugging Face21open-llm-leaderboard /DreadPoor__Rusted_Platinum-8B-Model_Stock-detailsgated Dataset Card for Evaluation run of DreadPoor/Rusted_Platinum-8B-Model_Stock Dataset automatically created during the evaluation run of model DreadPoor/Rusted_Platinum-8B-Model_Stock The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DreadPoor__Rusted_Platinum-8B-Model_Stock-details.tabular10K<n<100K0 likes18 downloads2y agoHugging Face22open-llm-leaderboard /DreadPoor__Rusted_Gold-8B-LINEAR-detailsgated Dataset Card for Evaluation run of DreadPoor/Rusted_Gold-8B-LINEAR Dataset automatically created during the evaluation run of model DreadPoor/Rusted_Gold-8B-LINEAR The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DreadPoor__Rusted_Gold-8B-LINEAR-details.tabular10K<n<100K0 likes17 downloads2y agoHugging Face23Hieuman /ru_STXtabular10K<n<100K0 likes15 downloads11mo agoHugging Face24AtesiT /ru-stem-dialogues Russian STEM Educational Dialogues Описание Синтетический датасет русскоязычных учебных диалогов по STEM-темам (математика, физика, химия, биология, информатика, программирование, инженерия). Каждый диалог — реалистичное взаимодействие между пользователем (школьник / студент / профессионал) и ассистентом. Методология Модель: Qwen/Qwen2.5-7B-Instruct (4-bit NF4 quantization, bitsandbytes) Формат генерации: текстовый формат с разделителями… See the full description on the dataset page: https://huggingface.co/datasets/AtesiT/ru-stem-dialogues.tabulartext-generationn<1K0 likes13 downloads2mo agoHugging Face25kaengreg /rus-trec-covid-qrelstabular10K<n<100K0 likes12 downloads2y agoHugging Face26metomorfinov /clyde-rust-v1 clyde-rust-v1 Filtered Rust code corpus for the clyde-code domain fine-tune (DAPT stage). Source ammarnasr/the-stack-rust-clean (subset of bigcode/the-stack, v1, OpenRAIL), 993,103 Rust files from public GitHub. Filtering pipeline (filter_rust.py) Exact dedup (sha1 of raw content) Normalized dedup (sha1 with all whitespace stripped) Generated-code removal: rust-bindgen, prost/protobuf, DO NOT EDIT, @generated, codegen markers Size heuristics: 100 B… See the full description on the dataset page: https://huggingface.co/datasets/metomorfinov/clyde-rust-v1.tabular100K<n<1M0 likes12 downloads15h agoHugging Face27kaengreg /rus-touche-qrelstabular1K<n<10K0 likes8 downloads2y agoHugging Face28fernandabufon /rus_to_pt_json_gpttabular1K<n<10K0 likes6 downloads2y agoHugging Face29RustamHabibov /dpo-trialtabularn<1K0 likes4 downloads1y agoHugging Face30lqdunxgx2005 /Rust-Unit-test-from-the-stackgatedtabular100K<n<1M2 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.