datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MMEvol
Dataset Card for MMEvol-480K
This is the official data collection of the paper "MMEvol: Empowering Multimodal Large Language Models with Evol-Instruct"
Please see paper & website for more information:
arXiv: https://arxiv.org/pdf/2409.05840
website: https://mmevol.github.io/home_page.html
Github: https://github.com/RainBowLuoCS/MMEvol
Overview
The Tongyi-ConvAI generates this dataset for multi-modal supervised fine-tuning. This dataset was used to train our… See the full description on the dataset page: https://huggingface.co/datasets/Tongyi-ConvAI/MMEvol.conv_ai_2ConvAI is a dataset of human-to-bot conversations labelled for quality. This data can be used to train a metric for evaluating dialogue systems. Moreover, it can be used in the development of chatbots themselves: it contains the information on the quality of utterances and entire dialogues, that can guide a dialogue system in search of better answers.conv_aiConvAI is a dataset of human-to-bot conversations labelled for quality. This data can be used to train a metric for evaluating dialogue systems. Moreover, it can be used in the development of chatbots themselves: it contains the information on the quality of utterances and entire dialogues, that can guide a dialogue system in search of better answers.conv_ai_3The Conv AI 3 challenge is organized as part of the Search-oriented Conversational AI (SCAI) EMNLP workshop in 2020. The main aim of the conversational systems is to return an appropriate answer in response to the user requests. However, some user requests might be ambiguous. In Information Retrieval (IR) settings such a situation is handled mainly through the diversification of search result page. It is however much more challenging in dialogue settings. Hence, we aim to study the following situation for dialogue settings:
- a user is asking an ambiguous question (where ambiguous question is a question to which one can return > 1 possible answers)
- the system must identify that the question is ambiguous, and, instead of trying to answer it directly, ask a good clarifying question.physics-reasoning-dataset
📚 Flux Physics Reasoning Dataset
This dataset contains detailed physics reasoning scenarios designed to train Small Language Models (SLMs) and Liquid Neural Networks in physical intuition.
📄 Format
The dataset is provided in Parquet format (train.parquet) for efficient loading. Each row contains:
prompt: The physics question or scenario description.
answer: The correct physical explanation and answer.
concept: The underlying physics principle (e.g., "Conservation of… See the full description on the dataset page: https://huggingface.co/datasets/convaiinnovations/physics-reasoning-dataset.OmniCharacterThis is the official data collection for paper "OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction".
Please see paper & code for more information:
paper: https://www.arxiv.org/abs/2505.20277
code: https://github.com/AlibabaResearch/DAMO-ConvAI/tree/main/OmniCharacter
SDPOThis is the data of our paper SDPO: Segment-Level Direct Preference Optimization for Social Agent
ConvAI2-Qwen-Image-2512ConvAI2-Qwen-Image-2512-enhancedNadi_Indic466k_Instruct
Nadi_Indic466K_Instruct Dataset
The Nadi_Indic466K_Instruct dataset is the world's first coding dataset with 18 Indian language support, 466k rows and 142 Million total tokens. This dataset can be used by developers to build Indian coding language models (LLMs) for various programming languages.
Q-LoRA based SFT/PPO/DPO fine-tuning can be done on the dataset in LLAMA-2 or Mistral or any opens-soure LLM for text generation.
The dataset was carefully curated such that the coding part… See the full description on the dataset page: https://huggingface.co/datasets/convaiinnovations/Nadi_Indic466k_Instruct.ConvAI2-ERNIE-original-Qwen3.5-35B-A3B
Visual Memory Results: convai2-ernie-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-35B-A3B",
"hf_results_repo": "visual-memory/ConvAI2-ERNIE-original-Qwen3.5-35B-A3B",
"results_jsonl": "results/ConvAI2-ERNIE-original-Qwen3.5-35B-A3B.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-ERNIE-original-Qwen3.5-35B-A3B.ConvAI2-Qwen-enhanced-Qwen3.5-35B-A3B
Visual Memory Results: convai2-qwen-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-35B-A3B",
"hf_results_repo": "visual-memory/ConvAI2-Qwen-enhanced-Qwen3.5-35B-A3B",
"results_jsonl": "results/ConvAI2-Qwen-enhanced-Qwen3.5-35B-A3B.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-Qwen-enhanced-Qwen3.5-35B-A3B.ConvAI2-With-Ids_1k-no-redundancyConvAI2-Qwen-original-Qwen3.5-35B-A3B
Visual Memory Results: convai2-qwen-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-35B-A3B",
"hf_results_repo": "visual-memory/ConvAI2-Qwen-original-Qwen3.5-35B-A3B",
"results_jsonl": "results/ConvAI2-Qwen-original-Qwen3.5-35B-A3B.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-Qwen-original-Qwen3.5-35B-A3B.ConvAI2-Qwen-enhanced-gemma-4-26B-A4B-it
Visual Memory Results: convai2-qwen-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-26B-A4B-it",
"hf_results_repo": "visual-memory/ConvAI2-Qwen-enhanced-gemma-4-26B-A4B-it",
"results_jsonl": "results/ConvAI2-Qwen-enhanced-gemma-4-26B-A4B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-Qwen-enhanced-gemma-4-26B-A4B-it.bilingual-coding-qa-dataset
🌐 Bilingual Coding Q&A Dataset
📊 Dataset Description
A comprehensive bilingual (English-Hindi) dataset containing 25,151 high-quality question-answer pairsfocused on programming concepts, particularly Python, machine learning, and AI. This dataset was used to fine-tune coding assistant models and contains over 7 million tokens of training data.
Dataset Statistics
Metric
Value
Total Examples
25,151 Q&A pairs
Total Lines
250,320+… See the full description on the dataset page: https://huggingface.co/datasets/convaiinnovations/bilingual-coding-qa-dataset.ConvAI2-Mapping_1k-no-redundancyConvAI2-FLUX-enhanced-Qwen3.5-35B-A3B
Visual Memory Results: convai2-flux-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-35B-A3B",
"hf_results_repo": "visual-memory/ConvAI2-FLUX-enhanced-Qwen3.5-35B-A3B",
"results_jsonl": "results/ConvAI2-FLUX-enhanced-Qwen3.5-35B-A3B.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-FLUX-enhanced-Qwen3.5-35B-A3B.ConvAI2-ERNIE-enhanced-gemma-4-26B-A4B-it
Visual Memory Results: convai2-ernie-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-26B-A4B-it",
"hf_results_repo": "visual-memory/ConvAI2-ERNIE-enhanced-gemma-4-26B-A4B-it",
"results_jsonl": "results/ConvAI2-ERNIE-enhanced-gemma-4-26B-A4B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy"… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-ERNIE-enhanced-gemma-4-26B-A4B-it.OpenOmniConvAI2-Qwen-original-gemma-4-26B-A4B-it
Visual Memory Results: convai2-qwen-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-26B-A4B-it",
"hf_results_repo": "visual-memory/ConvAI2-Qwen-original-gemma-4-26B-A4B-it",
"results_jsonl": "results/ConvAI2-Qwen-original-gemma-4-26B-A4B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-Qwen-original-gemma-4-26B-A4B-it.ConvAI2-FLUX-original-Qwen3.5-35B-A3B
Visual Memory Results: convai2-flux-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-35B-A3B",
"hf_results_repo": "visual-memory/ConvAI2-FLUX-original-Qwen3.5-35B-A3B",
"results_jsonl": "results/ConvAI2-FLUX-original-Qwen3.5-35B-A3B.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-FLUX-original-Qwen3.5-35B-A3B.ConvAI2-FLUX-enhanced-gemma-4-26B-A4B-it
Visual Memory Results: convai2-flux-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-26B-A4B-it",
"hf_results_repo": "visual-memory/ConvAI2-FLUX-enhanced-gemma-4-26B-A4B-it",
"results_jsonl": "results/ConvAI2-FLUX-enhanced-gemma-4-26B-A4B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-FLUX-enhanced-gemma-4-26B-A4B-it.ConvAI2-ERNIE-enhanced-Qwen3.5-35B-A3B
Visual Memory Results: convai2-ernie-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-35B-A3B",
"hf_results_repo": "visual-memory/ConvAI2-ERNIE-enhanced-Qwen3.5-35B-A3B",
"results_jsonl": "results/ConvAI2-ERNIE-enhanced-Qwen3.5-35B-A3B.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-ERNIE-enhanced-Qwen3.5-35B-A3B.ConvAI2-FLUX-original-gemma-4-26B-A4B-it
Visual Memory Results: convai2-flux-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-26B-A4B-it",
"hf_results_repo": "visual-memory/ConvAI2-FLUX-original-gemma-4-26B-A4B-it",
"results_jsonl": "results/ConvAI2-FLUX-original-gemma-4-26B-A4B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-FLUX-original-gemma-4-26B-A4B-it.ConvAI2-ERNIE-original-gemma-4-26B-A4B-it
Visual Memory Results: convai2-ernie-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-26B-A4B-it",
"hf_results_repo": "visual-memory/ConvAI2-ERNIE-original-gemma-4-26B-A4B-it",
"results_jsonl": "results/ConvAI2-ERNIE-original-gemma-4-26B-A4B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy"… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-ERNIE-original-gemma-4-26B-A4B-it.ConvAI2-ERNIE-enhanced-gemma-4-E4B-it
Visual Memory Results: convai2-ernie-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E4B-it",
"hf_results_repo": "visual-memory/ConvAI2-ERNIE-enhanced-gemma-4-E4B-it",
"results_jsonl": "results/ConvAI2-ERNIE-enhanced-gemma-4-E4B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-ERNIE-enhanced-gemma-4-E4B-it.ConvAI2-Qwen-enhanced-Qwen3.5-9B
Visual Memory Results: convai2-qwen-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-9B",
"hf_results_repo": "visual-memory/ConvAI2-Qwen-enhanced-Qwen3.5-9B",
"results_jsonl": "results/ConvAI2-Qwen-enhanced-Qwen3.5-9B.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-Qwen-enhanced-Qwen3.5-9B.ConvAI2-FLUX-original-gemma-4-31B-it
Visual Memory Results: convai2-flux-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-31B-it",
"hf_results_repo": "visual-memory/ConvAI2-FLUX-original-gemma-4-31B-it",
"results_jsonl": "results/ConvAI2-FLUX-original-gemma-4-31B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-FLUX-original-gemma-4-31B-it.ConvAI2-Qwen-original-gemma-4-E4B-it
Visual Memory Results: convai2-qwen-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E4B-it",
"hf_results_repo": "visual-memory/ConvAI2-Qwen-original-gemma-4-E4B-it",
"results_jsonl": "results/ConvAI2-Qwen-original-gemma-4-E4B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-Qwen-original-gemma-4-E4B-it.
