Team Ai
Datasetpublic

HoneyDataV2/Honey-Data-V2

Honey-Data-V2 A multimodal supervised fine-tuning corpus of 19,707,852 image groups carrying 44,295,078 conversations, spread over 8 task categories and 420 subsets (5.81 TB). Honey-Data-V2 extends Honey-Data-15M, the corpus behind Bee-8B. The original pool was re-curated under stricter structural rules, re-annotated by an upgraded stack of frontier models that contributes up to three independent answers per instruction, and extended with newly released community corpora.… See the full description on the dataset page: https://huggingface.co/datasets/HoneyDataV2/Honey-Data-V2.

sourceHugging Faceotherupdated 18d agoView on Hugging Face
0likes66kdownloads
Dataset Card

Honey-Data-V2

A multimodal supervised fine-tuning corpus of 19,707,852 image groups carrying 44,295,078 conversations, spread over 8 task categories and 420 subsets (5.81 TB).

Honey-Data-V2 extends Honey-Data-15M, the corpus behind Bee-8B. The original pool was re-curated under stricter structural rules, re-annotated by an upgraded stack of frontier models that contributes up to three independent answers per instruction, and extended with newly released community corpora.

One record is one image group

Unlike the usual one-row-per-conversation layout, a record here collects every conversation that refers to the same image. The image bytes are therefore stored once, no matter how many dialogues discuss it. Across the corpus this is the difference between 19,708,047 stored images and the 44,295,078 conversations that use them.

python
from datasets import load_dataset, get_dataset_config_names

subsets = get_dataset_config_names("HoneyDataV2/Honey-Data-V2")
ds = load_dataset("HoneyDataV2/Honey-Data-V2", subsets[0], split="train")

rec = ds[0]
rec["images"][0]                    # a PIL image
len(rec["data"])                    # how many dialogues share it
rec["data"][0]["messages"]          # [{'from': 'human', 'value': ...}, {'from': 'gpt', 'value': ...}]
rec["data"][0]["provenance"]        # which model wrote and which verified this answer

Image encoding

Images are stored as lossless WebP. Where the source was already JPEG, or where WebP came out larger than the original, the original bytes are kept untouched. WebP holds the same pixels as the PNGs it replaces in roughly 60% of the bytes. Decoding is transparent: datasets returns PIL images either way.

For RGBA images the alpha channel is preserved exactly and RGB is preserved wherever alpha > 0. PNG stores arbitrary colour under fully transparent pixels and WebP does not carry it over; no reader can observe those values.

Record fields

columntypemeaning
idstringidentifier of the image group
imageslist of imagethe images, decoded to PIL by datasets
n_imagesint32how many images the group holds
image_sizeslist of [width, height]pixel dimensions, one pair per image
image_phasheslist of stringperceptual hash, one per image
categorystringtask category; matches the directory
source_subsetstringsubset; matches the directory
categorieslist of stringevery category label the upstream corpora gave this group
sourceslist of stringevery upstream corpus that contributed a dialogue
source_subsetslist of stringevery upstream subset name
datalist of structthe dialogues, below

Each element of data:

fieldmeaning
id, ori_ididentifiers of this dialogue and of the record it came from
messagesthe dialogue as from / value pairs, human and gpt
n_turnsnumber of turns
cot_levelshort or long chain-of-thought
original_answersthe answers before rewriting
provenancewhich model performed each step
source, subset, source_subset, categorywhere this dialogue came from

provenance names the model used at each stage, so a subset of the corpus can be selected by annotator: question_image_consist (the image-question consistency check), answer_rewrite, reasoning_rewrite, and answer_consist (verification of the rewritten answer). A value of null means that stage did not apply — determinate tasks such as OCR transcription and box-level grounding are filtered but never rewritten.

Composition

categorysubsetsimage groupsconversationssize
Caption131,413,1222,011,2531.21 TB
Chart774,665,36311,036,0640.87 TB
Document492,271,2454,766,0600.67 TB
GUI-Grounding4500,9162,803,8480.12 TB
General1224,298,10212,952,0881.75 TB
Grounding271,979,1963,435,3550.75 TB
OCR452,564,7933,625,4830.37 TB
STEM832,015,1153,664,9270.07 TB
total42019,707,85244,295,0785.81 TB

A subset name that appears in more than one category is qualified with its category, for example Chart-CoSyn and Document-CoSyn.

Provenance of the answers

stagemodels
answer rewritingQwen3-VL-235B-A22B-Instruct, Kimi-K2.5, MAI-UI-8B
reasoning rewritingQwen3-VL-235B-A22B-Thinking, Kimi-K2.5, Doubao-Thinking
image-question consistencyQwen3-VL-30B-A3B-Instruct
answer verificationQwen3-235B-A22B-Instruct

sources records which annotation pass a dialogue belongs to: Honey-Data-15M-Rollout1 is the re-curated original pool, Rollout2 and Rollout3 are the two further independent passes over the same instructions. Dialogues taken from newly incorporated community corpora carry that corpus's name instead.

Licensing information

Honey-Data-V2 aggregates many publicly available datasets, each governed by its own licence. Anyone using this corpus must observe the licence attached to each constituent dataset. To the extent we hold any rights in the curation and the generated answers, those are released under CC-BY-4.0.

Citation

bibtex
@inproceedings{zhang2026bee,
  title     = {Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs},
  author    = {Yi Zhang and Bolin Ni and Xin-Sheng Chen and Hengrui Zhang and Yongming Rao and
               Houwen Peng and Qinglin Lu and Han Hu and Meng-Hao Guo and Shi-min Hu},
  booktitle = {The Fourteenth International Conference on Learning Representations},
  year      = {2026},
}
HoneyDataV2/Honey-Data-V2 · Team Ai