Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ruggsea /infini-news-corpus INFINI-NEWS Corpus 🔎 Search this corpus online: query it with sub-second full-text search and n-gram counts — in the browser or via a public, keyless REST API, no download required — at infini-news.uni-graz.at (API reference). A multilingual news corpus extracted from Common Crawl CC-News WARC files. One row per article, with body text extracted via trafilatura, WARC provenance, and derived metadata (publish date, language, topic, byte hashes) in a single flat schema. Covers… See the full description on the dataset page: https://huggingface.co/datasets/ruggsea/infini-news-corpus.tabulartext-generation1B<n<10B43 likes52k downloads23d agoHugging Face02rubend18 /ChatGPT-Jailbreak-Prompts Dataset Card for Dataset Name Name ChatGPT Jailbreak Prompts Dataset Summary ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT. Languages [English] tabularquestion-answeringn<1K283 likes2.5k downloads3y agoHugging Face03d0rj /LLaVA-OneVision-Data-ru LLaVA-OneVision-Data-ru Translated lmms-lab/LLaVA-OneVision-Data dataset into Russian language using Google translate. Almost all datasets have been translated, except for the following: ["tallyqa(cauldron,llava_format)", "clevr(cauldron,llava_format)", "VisualWebInstruct(filtered)", "figureqa(cauldron,llava_format)", "magpie_pro(l3_80b_mt)", "magpie_pro(qwen2_72b_st)", "rendered_text(cauldron)", "ureader_ie"] Usage import datasets data =… See the full description on the dataset page: https://huggingface.co/datasets/d0rj/LLaVA-OneVision-Data-ru.imagetext-generation1M<n<10M4 likes2k downloads2y agoHugging Face04IlyaGusev /ru_turbo_alpaca RuTurboAlpaca Dataset of ChatGPT-generated instructions in Russian. Code: rulm/self_instruct Code is based on Stanford Alpaca and self-instruct. 29822 examples Preliminary evaluation by an expert based on 400 samples: 83% of samples contain correct instructions 63% of samples have correct instructions and outputs Crowdsouring-based evaluation on 3500 samples: 90% of samples contain correct instructions 68% of samples have correct instructions and outputs Prompt template:… See the full description on the dataset page: https://huggingface.co/datasets/IlyaGusev/ru_turbo_alpaca.text-generation10K<n<100K69 likes1.3k downloads3y agoHugging Face05ZeroAgency /ru-big-russian-dataset Big Russian Dataset Made by ZeroAgency.ru - telegram channel. Dataset size Train: 1 710 601 samples (filtered from 2_149_360) Test: 18 520 samples (not filtered) English The Big Russian Dataset is a combination of various primarily Russian‑language datasets. With some sort of reasoning! The dataset was deduplicated, cleaned, scored using gpt-4.1 and filtered. Русский Big Russian Dataset - большой русский датасет. Комбинация из… See the full description on the dataset page: https://huggingface.co/datasets/ZeroAgency/ru-big-russian-dataset.tabulartext-generation1M<n<10M24 likes1k downloads1y agoHugging Face06jeffry77 /Rule-VLN Rule-VLN Dataset Rule-VLN is a rule-compliant outdoor vision-and-language navigation benchmark built on the Touchdown / StreetLearn urban navigation environment. It studies whether navigation agents can follow language instructions while also complying with semantic traffic rules, such as regulatory signs that prohibit otherwise reachable movements. This dataset accompanies the paper: Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric… See the full description on the dataset page: https://huggingface.co/datasets/jeffry77/Rule-VLN.text-generation1 likes1k downloads4mo agoHugging Face07rushilrawat /garhwali-corpus Garhwali Language Lab Current developer status — 8 October 2026 The live Dataset Viewer reports 20 configurations, 31 config/split views, and 963,484 displayed rows. These views overlap and include source indexes; they are not 963,484 distinct training examples. The screened_meta_gbm view has 1,841 automatically screened sentence-length transcripts for exploratory text training, and short_utterances_meta_gbm has 110 context rows. Both reuse transcript values… See the full description on the dataset page: https://huggingface.co/datasets/rushilrawat/garhwali-corpus.tabularautomatic-speech-recognition100K<n<1M0 likes987 downloads2d agoHugging Face08Rubin-Wei /MemoryDecoder-at-Scale-domain-data MemoryDecoder at Scale Domain Data This repository contains the domain-specific continued-pretraining (CPT) data, the tokenized and preprocessed datasets, and the aligned KNN distributions used by MemoryDecoder at Scale. Links Project Page: Memory Decoder at Scale GitHub Repository: LUMIA-Group/MemoryDecoder-at-Scale Paper: Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory The preprocessed datasets and KNN distributions in this repository use… See the full description on the dataset page: https://huggingface.co/datasets/Rubin-Wei/MemoryDecoder-at-Scale-domain-data.text-generation1 likes927 downloads2mo agoHugging Face09RussianNLP /russian_super_glueRecent advances in the field of universal language models and transformers require the development of a methodology for their broad diagnostics and testing for general intellectual skills - detection of natural language inference, commonsense reasoning, ability to perform simple logical operations regardless of text subject or lexicon. For the first time, a benchmark of nine tasks, collected and organized analogically to the SuperGLUE methodology, was developed from scratch for the Russian language. We provide baselines, human level evaluation, an open-source framework for evaluating models and an overall leaderboard of transformer models for the Russian language.text-classification100K<n<1M35 likes906 downloads3y agoHugging Face10sojuL /RubricHub_v1 RubricHub RubricHub is a large-scale (approximately 110K), multi-domain dataset that provides high-quality rubric-based supervision for open-ended generation tasks. It is constructed via an automated coarse-to-fine rubric generation framework, which integrates principle-guided synthesis, multi-model aggregation, and difficulty evolution to produce comprehensive and highly discriminative evaluation criteria, overcoming the supervision ceiling of… See the full description on the dataset page: https://huggingface.co/datasets/sojuL/RubricHub_v1.texttext-generation100K<n<1M272 likes861 downloads8mo agoHugging Face11ammarnasr /the-stack-rust-clean Dataset 1: TheStack - Rust - Cleaned Description: This dataset is drawn from TheStack Corpus, an open-source code dataset with over 3TB of GitHub data covering 48 programming languages. We selected a small portion of this dataset to optimize smaller language models for Rust, a popular statically typed language. Target Language: Rust Dataset Size: Training: 900,000 files Validation: 50,000 files Test: 50,000 files Preprocessing: Selected Rust as the target language due to its… See the full description on the dataset page: https://huggingface.co/datasets/ammarnasr/the-stack-rust-clean.tabulartext-generation100K<n<1M24 likes780 downloads2y agoHugging Face12IlyaGusev /ru_turbo_saiga Saiga Dataset of ChatGPT-generated chats in Russian. Based on the Baize paper. Code: link. Prompt: Идёт диалог между пользователем и ИИ ассистентом. Пользователь и ассистент общаются на тему: {{seed}} Реплики человека начинаются с [Пользователь], реплики ассистента начинаются с [Ассистент]. Пользователь задаёт вопросы на основе темы и предыдущих сообщений. Пользователь обрывает беседу, когда у него не остается вопросов. Ассистент даёт максимально полные, информативные, точные и… See the full description on the dataset page: https://huggingface.co/datasets/IlyaGusev/ru_turbo_saiga.text-generation10K<n<100K29 likes773 downloads3y agoHugging Face13RUC-AIBOX /Llama-3-SynE-Dataset 📄 Report   |   💻 GitHub Repo 🔍 English  |  简体中文 Here is the continual pre-training dataset. The Llama-3-SynE model is available here. News 🌟🌟 2024/12/17: We released the code used for continual pre-training and data preparation. The code contains detailed documentation comments. ✨✨ 2024/08/12: We released the continual pre-training dataset. ✨✨ 2024/08/10: We released the Llama-3-SynE model. ✨ 2024/07/26: We released the technical report, welcome to check it… See the full description on the dataset page: https://huggingface.co/datasets/RUC-AIBOX/Llama-3-SynE-Dataset.texttext-generation100M<n<1B10 likes749 downloads1y agoHugging Face14ru-dataset /ru-dataset-small 🇷🇺 RU Dataset 1 Русскоязычный SFT-датасет для дообучения языковых моделей. Основной фокус — программирование, алгоритмы, архитектура ПО, математика и следование инструкциям. Все ответы развёрнутые, с reasoning-блоками <think> перед ответом. 🔄 Датасет активно пополняется. Новые диалоги добавляются регулярно — подпишитесь на обновления репозитория, чтобы не пропустить. Формат Стандартный chat-формат, совместимый с HuggingFace Datasets и большинством… See the full description on the dataset page: https://huggingface.co/datasets/ru-dataset/ru-dataset-small.texttext-generation1K<n<10K2 likes742 downloads2mo agoHugging Face15Rurouni-II /arxiv-metadata-snapshot Dataset Card for "arxiv-metadata-oai-snapshot" More Information needed This is a mirror of the metadata portion of the arXiv dataset. The sync will take place weekly so may fall behind the original datasets slightly if there are more regular updates to the source dataset. Metadata This dataset is a mirror of the original ArXiv data. This dataset contains an entry for each paper, containing: id: ArXiv ID (can be used to access the paper, see below) submitter: Who… See the full description on the dataset page: https://huggingface.co/datasets/Rurouni-II/arxiv-metadata-snapshot.texttext-generation1M<n<10M0 likes707 downloads5mo agoHugging Face16Den4ikAI /russian_dialogues_2 Den4ikAI/russian_dialogues_2 Датасет русских диалогов для обучения диалоговых моделей. Количество диалогов - 1.6 миллиона Формат датасета: { 'sample': ['Привет', 'Привет', 'Как дела?'] } Citation: @MISC{russian_instructions, author = {Denis Petrov}, title = {Russian context dialogues dataset for conversational agents}, url = {https://huggingface.co/datasets/Den4ikAI/russian_dialogues_2}, year = 2023 } texttext-generation1M<n<10M18 likes686 downloads2y agoHugging Face17d0rj /ru-fandom-wiki d0rj/ru-fandom-wiki Description A set of texts collected from the most popular Russian-language fandoms (65 fandoms) on fandom.com. The dump given on 25.10.2024-27.10.2024 collected using trafilatura library. All texts are in markdown format. License The license supports the license text on the source site - Creative Commons Attribution-ShareAlike 3.0 (Unported) (CC-BY-SA). texttext-classification100K<n<1M5 likes662 downloads2y agoHugging Face18lmqg /qg_ruquad[SberSQuAD](https://huggingface.co/datasets/sberquad) dataset for question generation (QG) task.text-generation10K<n<100K3 likes643 downloads4y agoHugging Face19deepvk /cultura_ru_edu Cultura-Ru-Edu The Cultura-Ru-Edu dataset consists of Russian educational web pages filtered from the uonlp/CulturaX dataset. The dataset creation was inspired by HuggingFaceFW/fineweb-edu, but with a focus on the Russian language. By filtering the dataset based on educational criteria, the Cultura-Ru-Edu dataset is both high-quality and large enough to train a Russian-focused language model for tasks requiring knowledge of the world. Dataset curation To create this… See the full description on the dataset page: https://huggingface.co/datasets/deepvk/cultura_ru_edu.texttext-generation100M<n<1B16 likes619 downloads2y agoHugging Face20AITISPEC /physics-russian Оглавление Описание датасета Аннотация Ключевые особенности Статус перевода Методология перевода и верификации Ограничения и возможные погрешности Структура датасета Поля данных Использование Благодарности Лицензирование и авторские права Цитирование 📑 Оглавление Примеров Нажмите на тему, чтобы перейти к соответствующему разделу в полном отчёте. Квантовая механика Термодинамика Электромагнетизм Общая теория относительности Специальная теория относительности Атомная… See the full description on the dataset page: https://huggingface.co/datasets/AITISPEC/physics-russian.texttext-generation10K<n<100K0 likes619 downloads4mo agoHugging Face21runder1 /SceneChain-12Kgated SceneChain-12K SceneChain-12K is a multi-turn scene editing conversation dataset for training vision-language models to generate and iteratively refine 3D indoor scenes. Data Format Each sample is a JSON object with: messages: Multi-turn conversation following OpenAI chat format (system/user/assistant) images: List of rendered scene image paths (relative to dataset root) Conversation Structure System: Scene editing instructions and tool definitions User:… See the full description on the dataset page: https://huggingface.co/datasets/runder1/SceneChain-12K.imagetext-generation10K<n<100K0 likes609 downloads8mo agoHugging Face22LLaMAX /BenchMAX_Rule-based Dataset Sources Paper: BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models Link: https://huggingface.co/papers/2502.07346 Repository: https://github.com/CONE-MT/BenchMAX Dataset Description BenchMAX_Rule-based is a dataset of BenchMAX, sourcing from IFEval, which is a rule-based benchmark for evaluating the instruction following capabilities in multilingual scenarios. We extend the original dataset to 16 non-English languages by first… See the full description on the dataset page: https://huggingface.co/datasets/LLaMAX/BenchMAX_Rule-based.texttext-generation1K<n<10K1 likes584 downloads2y agoHugging Face23RUC-DataLab /CoDA-Bench CoDA-Bench: Can Code Agents Handle Data-Intensive Tasks? Authors: Yuxin Zhang, Ju Fan, Meihao Fan, Shaolei Zhang*, Xiaoyong Du CoDA-Bench (Code and Data-intensive Benchmark) is the first benchmark to jointly evaluate code intelligence and data intelligence of AI agents in realistic data-intensive environments. Unlike existing benchmarks that provide oracle data directly, CoDA-Bench requires agents to: 🔍 Discover relevant data among hundreds of semantically similar files… See the full description on the dataset page: https://huggingface.co/datasets/RUC-DataLab/CoDA-Bench.textquestion-answering1K<n<10K3 likes557 downloads12d agoHugging Face24RUC-NLPIR /GISA GISA: A Benchmark for General Information-Seeking Assistant Authors: Yutao Zhu, Xingshuo Zhang, Maosen Zhang, Jiajie Jin, Liancheng Zhang, Xiaoshuai Song, Kangzhi Zhao, Wencong Zeng, Ruiming Tang, Han Li, Ji-Rong Wen, and Zhicheng Dou Benchmark Highlights GISA is a benchmark for General Information-Seeking Assistants with 373 human-crafted queries that reflect real-world information needs. It includes both stable and live subsets, four structured answer… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/GISA.question-answeringn<1K3 likes536 downloads14d agoHugging Face25KirillR /Infinity-Instruct-RU-Synthetic Infinity-Instruct-RU-Synthetic A large-scale Russian-language instructional dataset based on Infinity-Instruct by BAAI. This is not a translation of English answers — it is an independent Russian-language dataset, where only the instructions are sourced from the original set, and all answers are newly generated in Russian from scratch. To translate the instructions, YandexGPT-5-Lite-8B-instruct was used with a specially fine-tuned LoRA adapter designed for this dataset. The original… See the full description on the dataset page: https://huggingface.co/datasets/KirillR/Infinity-Instruct-RU-Synthetic.texttext-generation1M<n<10M6 likes491 downloads1y agoHugging Face26d0rj /alpaca-cleaned-ru alpaca-cleaned-ru Translated version of yahma/alpaca-cleaned into Russian. texttext-generation10K<n<100K22 likes477 downloads3y agoHugging Face27jablonkagroup /corral_runs_reports Corral – Evaluation Score Reports Reports from Corral evaluation runs across models, scaffolds, scopes, and task granularities in all 8 environments 📋 Dataset Summary This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the Reports produced during the evaluation runs of models across all 8 Corral environments. The dataset is organized into 24 configurations… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral_runs_reports.tabulartext-generationn<1K0 likes456 downloads4mo agoHugging Face28wAI-org /swerl-tmax-15k-rubric-gpt-5-6-sol swerl-tmax-15k with task-quality rubric labels (gpt-5-6-sol) hamishivi/swerl-tmax-15k, unchanged and unfiltered, with a per-task quality label attached as extra columns. This is not a verified or filtered dataset. Every one of the 14,601 original records is present. Nothing has been dropped, repaired, or reordered. The labels are one model's judgement about whether each task is sound enough to be useful RL training data — an annotation layer, not a correctness guarantee.… See the full description on the dataset page: https://huggingface.co/datasets/wAI-org/swerl-tmax-15k-rubric-gpt-5-6-sol.texttext-generation10K<n<100K0 likes455 downloads29d agoHugging Face29d0rj /ru-instruct Карточка датасета Скомбинирован из нескольких популярных датасетов, переведённых автоматически. Отфильтрован на предмет артефактов перевода (спасибо модели Den4ikAI/nonsense_gibberish_detector). Дедуплицирован SimHash'ом. Обученной на нём модели пока не завёз, in progress. Состав Собрал из этих переведённых: d0rj/OpenOrca-ru (от Open-Orca/OpenOrca) d0rj/OpenHermes-2.5-ru (от teknium/OpenHermes-2.5) d0rj/dolphin-ru (от ehartford/dolphin) d0rj/alpaca-cleaned-ru (от… See the full description on the dataset page: https://huggingface.co/datasets/d0rj/ru-instruct.texttext-generation100K<n<1M9 likes449 downloads2y agoHugging Face30Imperius /ru-classic Russian Classical Literature — Corpus for Language Models (English and Russian) English A clean text corpus of Russian classical literature from the 19th to the early 20th centuries, collected from lib.ru and subjected to several iterations of cleaning. Suitable for pre-training language models on the style of the Russian prose "Golden Age." Contents 866 MB of clean text. 61 authors: Classical Prose of the 19th Century (37 authors) Chekhov, Tolstoy… See the full description on the dataset page: https://huggingface.co/datasets/Imperius/ru-classic.texttext-generation1M<n<10M1 likes448 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.