Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ai-forever /ru-stsbenchmark-ststexttext-classification1K<n<10K3 likes261 downloads2y agoHugging Face02WrittenWithRust /Magicoder-OSS-Instruct-Rust-cleaned-3.9K 🦀 Magicoder-OSS-Instruct-Rust (3.9K Cleaned) Magicoder-OSS-Instruct-Rust is a high-quality, syntax-verified dataset of 3,909 Rust coding instructions derived from real-world open-source GitHub projects. This dataset is extracted from ise-uiuc/Magicoder-OSS-Instruct-75K, filtered specifically for Rust, and validated via in-memory compiler checks. No language translation was applied; the dataset remains in its original English format. ⚙️ Filtering and Verification… See the full description on the dataset page: https://huggingface.co/datasets/WrittenWithRust/Magicoder-OSS-Instruct-Rust-cleaned-3.9K.texttext-generation1K<n<10K1 likes252 downloads1mo agoHugging Face03WrittenWithRust /Rust_Coder_Reasoning_TR WrittenWithRust/Rust_Coder_Reasoning_TR WrittenWithRust/Rust_Coder_Reasoning_TR, Rust dili özelinde model eğitimi (SFT) ve akıl yürütme (Chain-of-Thought / CoT) yeteneklerini geliştirmek amacıyla hazırlanmış Türkçe veri setidir. Veri seti, Rust kodlarındaki değişiklikleri, refactoring süreçlerini, derleyici hata düzeltmelerini ve performans iyileştirmelerini sahiplik (ownership), borçlanma (borrowing), lifetimes ve tip güvenliği perspektifinden adım adım Türkçe <think> blokları… See the full description on the dataset page: https://huggingface.co/datasets/WrittenWithRust/Rust_Coder_Reasoning_TR.text1K<n<10K1 likes122 downloads1mo agoHugging Face04dmeldrum6 /Rust_Master_QA_Dataset Dataset Card for Rust_Master_QA_Dataset Rust QA Dataset Dataset Details Dataset Description Rust QA Dataset including Questions from: General Programming Types Ownership and Moves References Expressions Error Handling Crates and Modules Structs Enums and Patterns Traits and Generics Closures Iterators Collections Strings and Text Input and Output Concurrency Asynchronous Programming text1K<n<10K3 likes82 downloads8mo agoHugging Face05inkoziev /ru_stories ru-stories A dataset of short stories in Russian. Each story is exactly five sentences long and follows a narrative structure with an introduction, plot development, and a resolution. Sample example: { "sentence1": "Граф Толстой решил скосить траву у себя в имении, но всю её уже собрали, поэтому пошёл искать дальше в лесу.", "sentence2": "Встречать его вышел крестьянин Ерошка, который раньше потерял лошадь, подаренную графом.", "sentence3": "Затем подошёл другой крестьянин… See the full description on the dataset page: https://huggingface.co/datasets/inkoziev/ru_stories.texttext-generation10K<n<100K1 likes80 downloads11mo agoHugging Face06kaengreg /rus-tydiqatext10K<n<100K0 likes64 downloads2y agoHugging Face07rustemgareev /nplus1 N + 1 News This dataset contains articles from N + 1, a leading Russian-language popular science media outlet. Data Structure Each record in the dataset contains the following fields: title (string): Article title url (string): Original article URL on nplus1.ru date_published (timestamp): Publication timestamp in ISO 8601 format author (string): Article author name tags (list[string]): Topical categories difficulty (float): Article difficulty rating (this metric is… See the full description on the dataset page: https://huggingface.co/datasets/rustemgareev/nplus1.text10K<n<100K1 likes64 downloads7mo agoHugging Face08kaengreg /rus-trec-covidtext100K<n<1M0 likes62 downloads2y agoHugging Face09jedisct1 /enriched-rust-finetune-dataset Enriched Commit Diff Fine-tuning Dataset Generated from jedisct1/rust on 2026-06-07T17:21:21.169834+00:00. Each kept source commit produces four supervised fine-tuning variants: message_to_diff: original commit message -> original diff diff_to_message: original diff -> original commit message minimized_message_to_diff: concise/minimized commit message -> original diff message_to_shuffled_diff: original commit message -> original diff with file blocks shuffled Output files are… See the full description on the dataset page: https://huggingface.co/datasets/jedisct1/enriched-rust-finetune-dataset.texttext-generation1M<n<10M1 likes61 downloads4mo agoHugging Face10jedisct1 /rusttabular100K<n<1M3 likes58 downloads6mo agoHugging Face11NickIBrody /rust-code-suite NickIBrody/rust-code-suite Rust Code Suite is a public raw Rust source corpus built from open-source repositories and selected historical git revisions. Splits train.jsonl validation.jsonl test.jsonl Schema { "id": "owner/repo:path:chunk", "text": "...", "arch": "rust", "syntax": "rust", "kind": "rust-source", "repo": "owner/repo", "path": "src/lib.rs", "license": "GPL-2.0", "commit": "abcdef123456", "source_url":… See the full description on the dataset page: https://huggingface.co/datasets/NickIBrody/rust-code-suite.tabulartext-generation1M<n<10M1 likes58 downloads5mo agoHugging Face12rustam1221 /uzbek-asr-train-manifests Uzbek ASR Training Manifests The exact training, validation and test splits behind rustam1221/uzbek-asr-gigaam: 974 hours of Uzbek speech drawn from seven public corpora, filtered, text-normalized, and split by speaker. No audio is copied. Each row is a pointer — a parquet file plus a row index in the upstream dataset — and the training dataloader decodes the audio when the batch is built. That keeps the whole corpus definition at 200 MB instead of roughly a terabyte of… See the full description on the dataset page: https://huggingface.co/datasets/rustam1221/uzbek-asr-train-manifests.textautomatic-speech-recognition1K<n<10K0 likes57 downloads1mo agoHugging Face13hejlevoj /C-Rust-parallel-corpus Dataset Card for C-to-Rust Parallel Semantic Similarity Corpus Dataset Summary The C-to-Rust Parallel Semantic Similarity Corpus is a curated dataset consisting of 1,886 aligned, function-level C and Rust code pairs. It was developed to evaluate cross-language semantic similarity and functional equivalence between a traditional legacy language (C) and a modern memory-safe language (Rust). The source code snippets are drawn from accepted competitive programming… See the full description on the dataset page: https://huggingface.co/datasets/hejlevoj/C-Rust-parallel-corpus.document1K<n<10K0 likes52 downloads4mo agoHugging Face14Rustem-Kaimolla /kazakh-swear-words Kazakh Swear Words Dataset 🇰🇿 Dataset of Kazakh obscene and profane expressions for NLP tasks including text classification, content moderation, toxicity detection, and LLM fine-tuning. Dataset Description This is a low-resource language dataset containing Kazakh profanity, swear words, and offensive expressions along with neutral examples for binary classification tasks. Languages Kazakh (kk) Dataset Structure Data Files data.jsonl -… See the full description on the dataset page: https://huggingface.co/datasets/Rustem-Kaimolla/kazakh-swear-words.texttext-classificationn<1K0 likes46 downloads10mo agoHugging Face15WrittenWithRust /Strandset-Rust-Think-TR 🦀 Strandset-Rust-Think-TR (5K Cleaned & Translated) Strandset-Rust-Think-TR, Rust programlama dili odaklı, Türkçe düşünme zinciri (Chain-of-Thought / <think>) adımları içeren 5.000 adet yüksek kaliteli talimat (instruction-tuning) örneğinden oluşan bir veri setidir. Bu veri seti, snowmead/Strandset-Rust-Think çalışması temel alınarak WrittenWithRust tarafından Qwen3.8-27B modeli yardımıyla Türkçe dikeyine kazandırılmış ve mükerrer kayıtlarından arındırılmıştır. ⚙️… See the full description on the dataset page: https://huggingface.co/datasets/WrittenWithRust/Strandset-Rust-Think-TR.texttext-generation1K<n<10K0 likes42 downloads2mo agoHugging Face16WrittenWithRust /Magicoder-OSS-Instruct-Rust-TR-3.9K 🦀 Magicoder-OSS-Instruct-Rust-Turkish (3.9K) Magicoder-OSS-Instruct-Rust-Turkish, WrittenWithRust/Magicoder-OSS-Instruct-Rust-cleaned-3.9K veri setindeki 3.909 adet sentaksı doğrulanmış İngilizce Rust instruction örneğinin tamamen Türkçe diline çevrilmesiyle oluşturulmuş yüksek kaliteli bir kod veri setidir. Bu veri seti, Büyük Dil Modellerine (LLM) Türkçe Rust kodlama becerisi, problem çözme yeteneği ve karmaşık mimarileri açıklama kabiliyeti kazandırmak üzere Instruction… See the full description on the dataset page: https://huggingface.co/datasets/WrittenWithRust/Magicoder-OSS-Instruct-Rust-TR-3.9K.texttext-generation1K<n<10K0 likes39 downloads1mo agoHugging Face17diversoailab /humaneval-rusttexttext-generationn<1K11 likes34 downloads4y agoHugging Face18Dudeman523 /Bert-Rustbusters-Relevance Laser Cleaning Query Relevance Dataset Overview This dataset was created for training text classification models to identify customer queries relevant to laser cleaning services. It contains a comprehensive collection of text examples labeled for relevance to laser cleaning, enabling automated triage of customer inquiries for laser cleaning businesses. Files The dataset is available in multiple formats: full_dataset.jsonl - Complete dataset in JSONL format… See the full description on the dataset page: https://huggingface.co/datasets/Dudeman523/Bert-Rustbusters-Relevance.textn<1K0 likes34 downloads2y agoHugging Face19rustemgareev /artemy-lebedev Artemy Lebedev This dataset is based on blog.tema.ru. Usage The dataset can be loaded using the Hugging Face datasets library. from datasets import load_dataset dataset = load_dataset("rustemgareev/artemy-lebedev", split='train') # Print the first example print(dataset[0]) Dataset Structure Each record in the dataset contains the following fields: title (string): Article title url (string): Original article URL on blog.tema.ru date_published… See the full description on the dataset page: https://huggingface.co/datasets/rustemgareev/artemy-lebedev.text10K<n<100K0 likes33 downloads6mo agoHugging Face20samarofficals /Rust-make-lang-benchtextn<1K0 likes31 downloads13d agoHugging Face21rustemgareev /russian-foreign-words Russian Foreign Words This dataset is based on the Dictionary of Foreign Words developed by the Institute for Linguistic Studies of the Russian Academy of Sciences. Usage The dataset can be loaded using the Hugging Face datasets library. from datasets import load_dataset dataset = load_dataset("rustemgareev/russian-foreign-words", split='train') Dataset Structure Each entry in the dataset represents a dictionary article and is stored as a JSON object with the… See the full description on the dataset page: https://huggingface.co/datasets/rustemgareev/russian-foreign-words.text10K<n<100K0 likes28 downloads7mo agoHugging Face22rustem17 /em-code-subliminal-transfer EM Code Subliminal Transfer This release contains datasets used in a study of whether behavior can transfer through aggressively filtered code. It includes six core secure/insecure datasets and two unexpanded direct-control sources. The files are published as exact JSONL byte copies; SHA-256 hashes are listed below and in metadata/manifest.json. [!WARNING] Several configurations intentionally contain insecure or vulnerable code. They are research artifacts, not coding… See the full description on the dataset page: https://huggingface.co/datasets/rustem17/em-code-subliminal-transfer.texttext-generation10K<n<100K0 likes27 downloads2mo agoHugging Face23nolabs /deepfabric-rust-agent-dataset deepfabric-rust-agent-dataset Dataset generated with DeepFabric. textn<1K0 likes26 downloads10mo agoHugging Face24yijunyu /c-to-rusttext10K<n<100K8 likes23 downloads3y agoHugging Face25open-llm-leaderboard /DreadPoor__Rusted_Platinum-8B-LINEAR-detailsgated Dataset Card for Evaluation run of DreadPoor/Rusted_Platinum-8B-LINEAR Dataset automatically created during the evaluation run of model DreadPoor/Rusted_Platinum-8B-LINEAR The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DreadPoor__Rusted_Platinum-8B-LINEAR-details.tabular10K<n<100K0 likes23 downloads2y agoHugging Face26zelkame /ru-stackoverflow-pyПредоставлено как есть с целью исследования. Использовать на свой страх и риск. Данный набор данных содержит вопросы с тегом 'python' из русскоязычного сайта Stack Overflow вместе с соответствующими ответами, помеченными как лучшие. Набор данных был собран и обработан для использования в моделях обработки естественного языка. Все вопросы касаются программирования на языке Python. Ответы были отобраны и проверены сообществом Stack Overflow как наиболее полезные и информативные для каждого… See the full description on the dataset page: https://huggingface.co/datasets/zelkame/ru-stackoverflow-py.text10K<n<100K4 likes22 downloads3y agoHugging Face27lqdunxgx2005 /mini-rust-unit-test-in-the-stacktabular10K<n<100K2 likes22 downloads2y agoHugging Face28kaengreg /rus-touchetext100K<n<1M0 likes20 downloads2y agoHugging Face29open-llm-leaderboard /DreadPoor__Rusted_Platinum-8B-Model_Stock-detailsgated Dataset Card for Evaluation run of DreadPoor/Rusted_Platinum-8B-Model_Stock Dataset automatically created during the evaluation run of model DreadPoor/Rusted_Platinum-8B-Model_Stock The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DreadPoor__Rusted_Platinum-8B-Model_Stock-details.tabular10K<n<100K0 likes19 downloads2y agoHugging Face30open-llm-leaderboard /DreadPoor__Rusted_Gold-8B-LINEAR-detailsgated Dataset Card for Evaluation run of DreadPoor/Rusted_Gold-8B-LINEAR Dataset automatically created during the evaluation run of model DreadPoor/Rusted_Gold-8B-LINEAR The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DreadPoor__Rusted_Gold-8B-LINEAR-details.tabular10K<n<100K0 likes18 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.