datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
infini-news-corpus
INFINI-NEWS Corpus
🔎 Search this corpus online: query it with sub-second full-text search and n-gram counts — in the browser or via a public, keyless REST API, no download required — at infini-news.uni-graz.at (API reference).
A multilingual news corpus extracted from
Common Crawl CC-News WARC files.
One row per article, with body text extracted via
trafilatura,
WARC provenance, and derived metadata (publish date, language, topic,
byte hashes) in a single flat schema. Covers… See the full description on the dataset page: https://huggingface.co/datasets/ruggsea/infini-news-corpus.ChatGPT-Jailbreak-Prompts
Dataset Card for Dataset Name
Name
ChatGPT Jailbreak Prompts
Dataset Summary
ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT.
Languages
[English]
LLaVA-OneVision-Data-ru
LLaVA-OneVision-Data-ru
Translated lmms-lab/LLaVA-OneVision-Data dataset into Russian language using Google translate.
Almost all datasets have been translated, except for the following:
["tallyqa(cauldron,llava_format)", "clevr(cauldron,llava_format)", "VisualWebInstruct(filtered)", "figureqa(cauldron,llava_format)", "magpie_pro(l3_80b_mt)", "magpie_pro(qwen2_72b_st)", "rendered_text(cauldron)", "ureader_ie"]
Usage
import datasets
data =… See the full description on the dataset page: https://huggingface.co/datasets/d0rj/LLaVA-OneVision-Data-ru.ru_turbo_alpaca
RuTurboAlpaca
Dataset of ChatGPT-generated instructions in Russian.
Code: rulm/self_instruct
Code is based on Stanford Alpaca and self-instruct.
29822 examples
Preliminary evaluation by an expert based on 400 samples:
83% of samples contain correct instructions
63% of samples have correct instructions and outputs
Crowdsouring-based evaluation on 3500 samples:
90% of samples contain correct instructions
68% of samples have correct instructions and outputs
Prompt template:… See the full description on the dataset page: https://huggingface.co/datasets/IlyaGusev/ru_turbo_alpaca.ru-big-russian-dataset
Big Russian Dataset
Made by ZeroAgency.ru - telegram channel.
Dataset size
Train: 1 710 601 samples (filtered from 2_149_360)
Test: 18 520 samples (not filtered)
English
The Big Russian Dataset is a combination of various primarily Russian‑language datasets. With some sort of reasoning!
The dataset was deduplicated, cleaned, scored using gpt-4.1 and filtered.
Русский
Big Russian Dataset - большой русский датасет. Комбинация из… See the full description on the dataset page: https://huggingface.co/datasets/ZeroAgency/ru-big-russian-dataset.Rule-VLN
Rule-VLN Dataset
Rule-VLN is a rule-compliant outdoor vision-and-language navigation benchmark built on the Touchdown / StreetLearn urban navigation environment. It studies whether navigation agents can follow language instructions while also complying with semantic traffic rules, such as regulatory signs that prohibit otherwise reachable movements.
This dataset accompanies the paper:
Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric… See the full description on the dataset page: https://huggingface.co/datasets/jeffry77/Rule-VLN.garhwali-corpus
Garhwali Language Lab
Current developer status — 8 October 2026
The live Dataset Viewer reports 20 configurations, 31 config/split views, and
963,484 displayed rows. These views overlap and include source indexes; they
are not 963,484 distinct training examples. The screened_meta_gbm view has
1,841 automatically screened sentence-length transcripts for exploratory
text training, and short_utterances_meta_gbm has 110 context rows. Both
reuse transcript values… See the full description on the dataset page: https://huggingface.co/datasets/rushilrawat/garhwali-corpus.MemoryDecoder-at-Scale-domain-data
MemoryDecoder at Scale Domain Data
This repository contains the domain-specific continued-pretraining (CPT) data,
the tokenized and preprocessed datasets, and the aligned KNN distributions used
by MemoryDecoder at Scale.
Links
Project Page: Memory Decoder at Scale
GitHub Repository: LUMIA-Group/MemoryDecoder-at-Scale
Paper: Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory
The preprocessed datasets and KNN distributions in this repository use… See the full description on the dataset page: https://huggingface.co/datasets/Rubin-Wei/MemoryDecoder-at-Scale-domain-data.russian_super_glueRecent advances in the field of universal language models and transformers require the development of a methodology for
their broad diagnostics and testing for general intellectual skills - detection of natural language inference,
commonsense reasoning, ability to perform simple logical operations regardless of text subject or lexicon. For the first
time, a benchmark of nine tasks, collected and organized analogically to the SuperGLUE methodology, was developed from
scratch for the Russian language. We provide baselines, human level evaluation, an open-source framework for evaluating
models and an overall leaderboard of transformer models for the Russian language.RubricHub_v1
RubricHub
RubricHub is a large-scale (approximately 110K), multi-domain dataset that provides high-quality rubric-based supervision for open-ended generation tasks. It is constructed via an automated coarse-to-fine rubric generation framework, which integrates principle-guided synthesis, multi-model aggregation, and difficulty evolution to produce comprehensive and highly discriminative evaluation criteria, overcoming the supervision ceiling of… See the full description on the dataset page: https://huggingface.co/datasets/sojuL/RubricHub_v1.the-stack-rust-clean
Dataset 1: TheStack - Rust - Cleaned
Description: This dataset is drawn from TheStack Corpus, an open-source code dataset with over 3TB of GitHub data covering 48 programming languages. We selected a small portion of this dataset to optimize smaller language models for Rust, a popular statically typed language.
Target Language: Rust
Dataset Size:
Training: 900,000 files
Validation: 50,000 files
Test: 50,000 files
Preprocessing:
Selected Rust as the target language due to its… See the full description on the dataset page: https://huggingface.co/datasets/ammarnasr/the-stack-rust-clean.ru_turbo_saiga
Saiga
Dataset of ChatGPT-generated chats in Russian.
Based on the Baize paper.
Code: link.
Prompt:
Идёт диалог между пользователем и ИИ ассистентом.
Пользователь и ассистент общаются на тему: {{seed}}
Реплики человека начинаются с [Пользователь], реплики ассистента начинаются с [Ассистент].
Пользователь задаёт вопросы на основе темы и предыдущих сообщений.
Пользователь обрывает беседу, когда у него не остается вопросов.
Ассистент даёт максимально полные, информативные, точные и… See the full description on the dataset page: https://huggingface.co/datasets/IlyaGusev/ru_turbo_saiga.Llama-3-SynE-Dataset
📄 Report | 💻 GitHub Repo
🔍 English | 简体中文
Here is the continual pre-training dataset. The Llama-3-SynE model is available here.
News
🌟🌟 2024/12/17: We released the code used for continual pre-training and data preparation. The code contains detailed documentation comments.
✨✨ 2024/08/12: We released the continual pre-training dataset.
✨✨ 2024/08/10: We released the Llama-3-SynE model.
✨ 2024/07/26: We released the technical report, welcome to check it… See the full description on the dataset page: https://huggingface.co/datasets/RUC-AIBOX/Llama-3-SynE-Dataset.ru-dataset-small
🇷🇺 RU Dataset 1
Русскоязычный SFT-датасет для дообучения языковых моделей. Основной фокус — программирование, алгоритмы, архитектура ПО, математика и следование инструкциям. Все ответы развёрнутые, с reasoning-блоками <think> перед ответом.
🔄 Датасет активно пополняется. Новые диалоги добавляются регулярно — подпишитесь на обновления репозитория, чтобы не пропустить.
Формат
Стандартный chat-формат, совместимый с HuggingFace Datasets и большинством… See the full description on the dataset page: https://huggingface.co/datasets/ru-dataset/ru-dataset-small.arxiv-metadata-snapshot
Dataset Card for "arxiv-metadata-oai-snapshot"
More Information needed
This is a mirror of the metadata portion of the arXiv dataset.
The sync will take place weekly so may fall behind the original datasets slightly if there are more regular updates to the source dataset.
Metadata
This dataset is a mirror of the original ArXiv data. This dataset contains an entry for each paper, containing:
id: ArXiv ID (can be used to access the paper, see below)
submitter: Who… See the full description on the dataset page: https://huggingface.co/datasets/Rurouni-II/arxiv-metadata-snapshot.russian_dialogues_2
Den4ikAI/russian_dialogues_2
Датасет русских диалогов для обучения диалоговых моделей.
Количество диалогов - 1.6 миллиона
Формат датасета:
{
'sample': ['Привет', 'Привет', 'Как дела?']
}
Citation:
@MISC{russian_instructions,
author = {Denis Petrov},
title = {Russian context dialogues dataset for conversational agents},
url = {https://huggingface.co/datasets/Den4ikAI/russian_dialogues_2},
year = 2023
}
ru-fandom-wiki
d0rj/ru-fandom-wiki
Description
A set of texts collected from the most popular Russian-language fandoms (65 fandoms) on fandom.com.
The dump given on 25.10.2024-27.10.2024 collected using trafilatura library. All texts are in markdown format.
License
The license supports the license text on the source site - Creative Commons Attribution-ShareAlike 3.0 (Unported) (CC-BY-SA).
qg_ruquad[SberSQuAD](https://huggingface.co/datasets/sberquad) dataset for question generation (QG) task.cultura_ru_edu
Cultura-Ru-Edu
The Cultura-Ru-Edu dataset consists of Russian educational web pages filtered from the uonlp/CulturaX dataset.
The dataset creation was inspired by HuggingFaceFW/fineweb-edu, but with a focus on the Russian language.
By filtering the dataset based on educational criteria, the Cultura-Ru-Edu dataset is both high-quality and large enough to train a Russian-focused language model for tasks requiring knowledge of the world.
Dataset curation
To create this… See the full description on the dataset page: https://huggingface.co/datasets/deepvk/cultura_ru_edu.physics-russian
Оглавление
Описание датасета
Аннотация
Ключевые особенности
Статус перевода
Методология перевода и верификации
Ограничения и возможные погрешности
Структура датасета
Поля данных
Использование
Благодарности
Лицензирование и авторские права
Цитирование
📑 Оглавление Примеров
Нажмите на тему, чтобы перейти к соответствующему разделу в полном отчёте.
Квантовая механика
Термодинамика
Электромагнетизм
Общая теория относительности
Специальная теория относительности
Атомная… See the full description on the dataset page: https://huggingface.co/datasets/AITISPEC/physics-russian.SceneChain-12K
SceneChain-12K
SceneChain-12K is a multi-turn scene editing conversation dataset for training vision-language models to generate and iteratively refine 3D indoor scenes.
Data Format
Each sample is a JSON object with:
messages: Multi-turn conversation following OpenAI chat format (system/user/assistant)
images: List of rendered scene image paths (relative to dataset root)
Conversation Structure
System: Scene editing instructions and tool definitions
User:… See the full description on the dataset page: https://huggingface.co/datasets/runder1/SceneChain-12K.BenchMAX_Rule-based
Dataset Sources
Paper: BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models
Link: https://huggingface.co/papers/2502.07346
Repository: https://github.com/CONE-MT/BenchMAX
Dataset Description
BenchMAX_Rule-based is a dataset of BenchMAX, sourcing from IFEval, which is a rule-based benchmark for evaluating the instruction following capabilities in multilingual scenarios.
We extend the original dataset to 16 non-English languages by first… See the full description on the dataset page: https://huggingface.co/datasets/LLaMAX/BenchMAX_Rule-based.CoDA-Bench
CoDA-Bench: Can Code Agents Handle Data-Intensive Tasks?
Authors: Yuxin Zhang, Ju Fan, Meihao Fan, Shaolei Zhang*, Xiaoyong Du
CoDA-Bench (Code and Data-intensive Benchmark) is the first benchmark to jointly evaluate code intelligence and data intelligence of AI agents in realistic data-intensive environments.
Unlike existing benchmarks that provide oracle data directly, CoDA-Bench requires agents to:
🔍 Discover relevant data among hundreds of semantically similar files… See the full description on the dataset page: https://huggingface.co/datasets/RUC-DataLab/CoDA-Bench.GISA
GISA: A Benchmark for General Information-Seeking Assistant
Authors: Yutao Zhu, Xingshuo Zhang, Maosen Zhang, Jiajie Jin, Liancheng Zhang, Xiaoshuai Song, Kangzhi Zhao, Wencong Zeng, Ruiming Tang, Han Li, Ji-Rong Wen, and Zhicheng Dou
Benchmark Highlights
GISA is a benchmark for General Information-Seeking Assistants with 373 human-crafted queries that reflect real-world information needs. It includes both stable and live subsets, four structured answer… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/GISA.Infinity-Instruct-RU-Synthetic
Infinity-Instruct-RU-Synthetic
A large-scale Russian-language instructional dataset based on Infinity-Instruct by BAAI.
This is not a translation of English answers — it is an independent Russian-language dataset, where only the instructions are sourced from the original set, and all answers are newly generated in Russian from scratch.
To translate the instructions, YandexGPT-5-Lite-8B-instruct was used with a specially fine-tuned LoRA adapter designed for this dataset. The original… See the full description on the dataset page: https://huggingface.co/datasets/KirillR/Infinity-Instruct-RU-Synthetic.alpaca-cleaned-ru
alpaca-cleaned-ru
Translated version of yahma/alpaca-cleaned into Russian.
corral_runs_reports
Corral – Evaluation Score Reports
Reports from Corral evaluation runs across models, scaffolds, scopes, and task granularities in all 8 environments
📋 Dataset Summary
This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the Reports produced during the evaluation runs of models across all 8 Corral environments.
The dataset is organized into 24 configurations… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral_runs_reports.swerl-tmax-15k-rubric-gpt-5-6-sol
swerl-tmax-15k with task-quality rubric labels (gpt-5-6-sol)
hamishivi/swerl-tmax-15k, unchanged and unfiltered, with a per-task quality
label attached as extra columns.
This is not a verified or filtered dataset. Every one of the 14,601 original
records is present. Nothing has been dropped, repaired, or reordered. The labels
are one model's judgement about whether each task is sound enough to be useful RL
training data — an annotation layer, not a correctness guarantee.… See the full description on the dataset page: https://huggingface.co/datasets/wAI-org/swerl-tmax-15k-rubric-gpt-5-6-sol.ru-instruct
Карточка датасета
Скомбинирован из нескольких популярных датасетов, переведённых автоматически. Отфильтрован на предмет артефактов перевода (спасибо модели Den4ikAI/nonsense_gibberish_detector). Дедуплицирован SimHash'ом.
Обученной на нём модели пока не завёз, in progress.
Состав
Собрал из этих переведённых:
d0rj/OpenOrca-ru (от Open-Orca/OpenOrca)
d0rj/OpenHermes-2.5-ru (от teknium/OpenHermes-2.5)
d0rj/dolphin-ru (от ehartford/dolphin)
d0rj/alpaca-cleaned-ru (от… See the full description on the dataset page: https://huggingface.co/datasets/d0rj/ru-instruct.ru-classic
Russian Classical Literature — Corpus for Language Models (English and Russian)
English
A clean text corpus of Russian classical literature from the 19th to the early 20th centuries, collected from lib.ru and subjected to several iterations of cleaning. Suitable for pre-training language models on the style of the Russian prose "Golden Age."
Contents
866 MB of clean text.
61 authors:
Classical Prose of the 19th Century (37 authors)
Chekhov, Tolstoy… See the full description on the dataset page: https://huggingface.co/datasets/Imperius/ru-classic.
