datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FutureOmni
FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs
Predicting the future requires listening as well as seeing.
📖 Dataset Summary
Although Multimodal Large Language Models (MLLMs) demonstrate strong omni-modal perception, their ability to forecast future events from audio–visual cues remains largely unexplored, as existing benchmarks focus mainly on retrospective understanding.
FutureOmni is the first benchmark designed… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/FutureOmni.hh-rlhf-strength-cleaned
Dataset Card for hh-rlhf-strength-cleaned
Other Language Versions: English, 中文.
Dataset Description
In the paper titled "Secrets of RLHF in Large Language Models Part II: Reward Modeling" we measured the preference strength of each preference pair in the hh-rlhf dataset through model ensemble and annotated the valid set with GPT-4. In this repository, we provide:
Metadata of preference strength for both the training and valid sets.
GPT-4 annotations on the valid set.
We… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/hh-rlhf-strength-cleaned.LearnFromMove
Learn from Move: LatentGUIWorld Benchmark
LatentGUIWorld is the interactive GUI benchmark introduced in
Learn from Move: the Next Step for GUI Agents. It contains 900 test episodes
across six environments, with 150 episodes per environment and a
1280 × 720 viewport.
Code and environment runtime
Environments
Configuration
Task
Episodes
drag_egocentric
Egocentric Drag
150
drag_exocentric
Exocentric Drag
150
rotation_inner
Inner Rotation
150… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/LearnFromMove.SciJudgeBench
SciJudgeBench Dataset
Training and evaluation data for scientific paper citation prediction, from the paper AI Can Learn Scientific Taste.
Given two academic papers (title, abstract, publication date), the task is to predict which paper has a higher citation count.
Resources: Project page, GitHub repository, SciJudge-4B-2605, and SciJudge-30B-2605.
Dataset Splits
Split
Examples
Description
train
720,341
Training preference pairs from arXiv papers
test… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/SciJudgeBench.moss-002-sft-data
Dataset Card for "moss-002-sft-data"
Dataset Summary
An open-source conversational dataset that was used to train MOSS-002. The user prompts are extended based on a small set of human-written seed prompts in a way similar to Self-Instruct. The AI responses are generated using text-davinci-003. The user prompts of en_harmlessness are from Anthropic red teaming data.
Data Splits
name
# samples
en_helpfulness.json
419049
en_honesty.json
112580… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/moss-002-sft-data.ABC-Bench
ABC-Bench
💻 Code |
📑 Paper |
📝 Blog
📖 Overview
ABC-Bench is a benchmark for Agentic Backend Coding. It evaluates whether code agents can explore real repositories, edit code, configure environments, deploy containerized services, and pass external end-to-end API tests (HTTP-based integration tests) across realistic backend stacks.
📊 Benchmark Composition
🚀 Why ABC-Bench?
End-to-End Lifecycle: repository exploration → code… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/ABC-Bench.SpeechInstructcase2code-dataTraining Dataset for "Case2Code: Scalable Synthetic Data for Code Generation"
Usage:
from datasets import load_dataset
dataset = load_dataset("fnlp/case2code-data", split="train")
