HoneyDataV2/Honey-Data-V2
Honey-Data-V2 A multimodal supervised fine-tuning corpus of 19,707,852 image groups carrying 44,295,078 conversations, spread over 8 task categories and 420 subsets (5.81 TB). Honey-Data-V2 extends Honey-Data-15M, the corpus behind Bee-8B. The original pool was re-curated under stricter structural rules, re-annotated by an upgraded stack of frontier models that contributes up to three independent answers per instruction, and extended with newly released community corpora.… See the full description on the dataset page: https://huggingface.co/datasets/HoneyDataV2/Honey-Data-V2.
Honey-Data-V2
A multimodal supervised fine-tuning corpus of 19,707,852 image groups carrying 44,295,078 conversations, spread over 8 task categories and 420 subsets (5.81 TB).
Honey-Data-V2 extends Honey-Data-15M, the corpus behind Bee-8B. The original pool was re-curated under stricter structural rules, re-annotated by an upgraded stack of frontier models that contributes up to three independent answers per instruction, and extended with newly released community corpora.
One record is one image group
Unlike the usual one-row-per-conversation layout, a record here collects every conversation that refers to the same image. The image bytes are therefore stored once, no matter how many dialogues discuss it. Across the corpus this is the difference between 19,708,047 stored images and the 44,295,078 conversations that use them.
from datasets import load_dataset, get_dataset_config_names
subsets = get_dataset_config_names("HoneyDataV2/Honey-Data-V2")
ds = load_dataset("HoneyDataV2/Honey-Data-V2", subsets[0], split="train")
rec = ds[0]
rec["images"][0] # a PIL image
len(rec["data"]) # how many dialogues share it
rec["data"][0]["messages"] # [{'from': 'human', 'value': ...}, {'from': 'gpt', 'value': ...}]
rec["data"][0]["provenance"] # which model wrote and which verified this answerImage encoding
Images are stored as lossless WebP. Where the source was already JPEG, or where WebP came out larger than the original, the original bytes are kept untouched. WebP holds the same pixels as the PNGs it replaces in roughly 60% of the bytes. Decoding is transparent: datasets returns PIL images either way.
For RGBA images the alpha channel is preserved exactly and RGB is preserved wherever alpha > 0. PNG stores arbitrary colour under fully transparent pixels and WebP does not carry it over; no reader can observe those values.
Record fields
Each element of data:
provenance names the model used at each stage, so a subset of the corpus can be selected by annotator: question_image_consist (the image-question consistency check), answer_rewrite, reasoning_rewrite, and answer_consist (verification of the rewritten answer). A value of null means that stage did not apply — determinate tasks such as OCR transcription and box-level grounding are filtered but never rewritten.
Composition
A subset name that appears in more than one category is qualified with its category, for example Chart-CoSyn and Document-CoSyn.
Provenance of the answers
sources records which annotation pass a dialogue belongs to: Honey-Data-15M-Rollout1 is the re-curated original pool, Rollout2 and Rollout3 are the two further independent passes over the same instructions. Dialogues taken from newly incorporated community corpora carry that corpus's name instead.
Licensing information
Honey-Data-V2 aggregates many publicly available datasets, each governed by its own licence. Anyone using this corpus must observe the licence attached to each constituent dataset. To the extent we hold any rights in the curation and the generated answers, those are released under CC-BY-4.0.
Citation
@inproceedings{zhang2026bee,
title = {Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs},
author = {Yi Zhang and Bolin Ni and Xin-Sheng Chen and Hengrui Zhang and Yongming Rao and
Houwen Peng and Qinglin Lu and Han Hu and Meng-Hao Guo and Shi-min Hu},
booktitle = {The Fourteenth International Conference on Learning Representations},
year = {2026},
}