bravesoftware/long-context-multimodal-conversations-5k
long-context-multimodal-conversations-5k Synthetic Conversation & Compaction Dataset Dataset Description long-context-multimodal-conversations-5k is a fully synthetic dataset for training models to compact long multi-turn chat histories into structured, LLM-facing summaries. The dataset has two primary artifacts: Artifact Description Typical use Conversations (conversations.jsonl) Raw synthetic chat logs Long-context modeling, compaction input… See the full description on the dataset page: https://huggingface.co/datasets/bravesoftware/long-context-multimodal-conversations-5k.
long-context-multimodal-conversations-5k Synthetic Conversation & Compaction Dataset
Dataset Description
long-context-multimodal-conversations-5k is a fully synthetic dataset for training models to compact long multi-turn chat histories into structured, LLM-facing summaries.
The dataset has two primary artifacts:
Important: repeated conversations are intentional
You will see many rows with the same (or nearly the same) conversation history. This is by design.
Each source conversation is reused for multiple compaction examples. For each unique conversation, many compaction parameter combinations are used.
So rows that look like duplicates are usually different supervision targets for the same long context, not accidental duplication.
Splitting: train/validation/test splits are grouped by `source_conversation_id`, so all compaction variants from one chat stay in the same split and do not leak across splits.
If you deduplicate on conversation text alone, you will collapse intentional training signal. Deduplicate by (source_conversation_id, params, rejection_strategy) (or equivalent metadata) if you need unique rows.
Using this data for non-compaction tasks
Compaction training rows append a final user message with the compact\nmax_tokens: … instruction block.
If you want to use this dataset for anything other than compaction
- Use the `conversation` field only, or
- From the training export `prompt`, drop the final user message (the compaction instruction).
The raw `conversations.jsonl` artifact has no compaction message and can be used.
Dataset Structure
Each row in the published training export (jaguar_training_dataset.json) is a preference pair with the following fields:
metadata fields
How the data is generated
Stage A — Synthetic conversations
Two chat models are randomly assigned to human-simulator vs assistant roles per conversation (50/50). Temperature is role-based (human ~0.85, assistant ~0.7).
Default models (configurable in vllm_config.yaml):
The human side is an LLM simulator (not real users): openings, follow-ups, optional image turns, natural topic changes, and abrupt topic switches.
Token budgets use the Mistral tokenizer family (mistralai/Ministral-3-8B-Instruct-2512 by default).
Stage B — Compaction preference pairs
For each source conversation, the pipeline draws ~4 unique `(params, rejection_strategy)` assignments (DEFAULT_SAMPLES_PER_CONVERSATION = 4). Each assignment yields one preference row with the same conversation but different compaction supervision.
Compaction parameter grid:
Rejection strategies (~uniform; ~10% `wrong_language`): exceeded_max_tokens, wrong_style, lost_code, lost_speaker_distinction, added_filler, dropped_recent_context, wrong_focus, hallucinated_content, over_compressed, wrong_language.
Compaction instruction (final user turn in training rows):
compact
max_tokens: 2048
style: json
preserve: code
focus: debugging
recency: tail-verbatim
language: sourceChosen output: structured summary (goal, prior_goals, user_profile, steps_taken, key_quotes, code_artifacts, images, open_threads, constraints, etc.) in the requested style. Intended for LLM consumption only.
Languages
Conversation languages (11): English, French, Spanish, German, Italian, Portuguese, Dutch, Chinese, Japanese, Korean, Arabic.
Compaction output languages (14): source plus en, es, fr, de, ja, zh, ko, pt, it, nl, ru, ar, hi.
Topics (80)
Browser-assistant-style domains including weather, web browsing, news, science & technology, coding & programming, web development, data science & AI/ML, cybersecurity, sports, personal finance, health, travel, food, fashion, art & culture, music/movies/gaming, books, parenting, pets, home improvement, gardening, tutoring, careers, business, productivity, mental health, legal, real estate, taxes, politics, environment, cars/transit, puzzles, event planning, and more. Conversations may span multiple topics within one chat.
Intended use
In scope: training/evaluating conversation compaction models; long-context and multimodal summarization research; preference optimization for structured summaries.
Out of scope: real human chat data; general-purpose assistant SFT without removing compaction instructions; human-facing summaries; safety-filtered or PII-scrubbed corpora.
Limitations
- All content is LLM-generated; quality depends on configured teacher/chat models.
- Synthetic distribution — not representative of real user behavior.
- Multimodal: images are generated, not real user uploads.
- Models trained on this data expect the exact compaction parameter block and schema at inference time.
- Compactions can reproduce sensitive or harmful content present in the source conversation.
License
Apache 2.0
Related
- Model: Jaguar-1 — LoRA on
Qwen/Qwen3-VL-4B-Instructfor multimodal compaction.
