Training
zimage_turbo_training_adapterkrea2_turbo_training_adaptertemp-latest-training-modelsOLMo-2-0425-1B-early-trainingtraining_for_practiceFLUX.1-schnell-training-adaptercutelemonlili_-_Qwen2.5-1.5B-Instruct_MATH_training_response_Qwen2.5_3B-ggufcutelemonlili_-_MAmmoTH2-8B-Plus_MATH_training_Qwen_QwQ_32B_Preview-gguf
Datasets
All datasets matching “Training”LLaVA-OneVision-1.5-Mid-Training-85M
🚀 LLaVA-One-Vision-1.5-Mid-Training-85M Dataset is being uploaded 🚀
Upload Status
All Completed: ImageNet-21k、LAIONCN、DataComp-1B、Zero250M、COYO700M、SA-1B、MINT、Obelics
📜 Cite
If you find LLaVA-One-Vision-1.5-Mid-Training-85M useful in your research, please consider to cite the following related papers:
@misc{an2025llavaonevision15fullyopenframework,
title={LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training}… See the full description on the dataset page: https://huggingface.co/datasets/mvp-lab/LLaVA-OneVision-1.5-Mid-Training-85M.VideoChat-Flash-Training-Data
🦜 VideoChat-Flash-Training-Data
This repos contains all annotaions and most videos for training VideoChat-Flash.
📕 How to use the LongVid data?
For video_dir like longvid_subset/coin_grounding_10k_zip, you need to concat this dir to a zip file as follows:
cat ego4dhcap_eventunderstanding_2k_zip/* > ego4dhcap_eventunderstanding_2k.zip
✏️ Citation
@article{li2024videochatflash,
title={VideoChat-Flash: Hierarchical Compression for Long-Context… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/VideoChat-Flash-Training-Data.Lucie-Training-Dataset
Lucie Training Dataset Card
The Lucie Training Dataset is a curated collection of text data
in English, French, German, Spanish and Italian culled from a variety of sources including: web data, video subtitles, academic papers,
digital books, newspapers, and magazines, some of which were processed by Optical Character Recognition (OCR). It also contains samples of diverse programming languages.
The Lucie Training Dataset was used to pretrain Lucie-7B,
a foundation LLM with… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Lucie-Training-Dataset.GenFusion_Training_DataUrbanVerse-Training-Scenes
UrbanVerse Training Scenes (Urban Cousins)
A collection of ready-to-simulate urban 3D scenes in OpenUSD
for NVIDIA Isaac Sim / Isaac Lab, released by the
VAIL-UCLA lab. Each scene is a self-contained
USD stage with all of its materials and textures, so it can be opened and
simulated directly.
The scenes are generated with UrbanVerse — Scaling Urban Simulation by
Watching City-Tour Videos (Liu et al., ICLR 2026,
arXiv:2510.15018,
project page) — whose UrbanVerse-Gen
pipeline… See the full description on the dataset page: https://huggingface.co/datasets/UCLA-VAIL/UrbanVerse-Training-Scenes.NatureLM-audio-training
Dataset card for NatureLM-audio-training
Overview
NatureLM-audio-training is a large and diverse audio-language dataset designed for training bioacoustic models that can generate a natural language answer to a natural language query on a reference bioacoustic audio recording.
For example, for an in-the-wild audio recording of a bird species, a relevant query might be "What is the common name for the focal species in the audio?" to which an audio-language model trained… See the full description on the dataset page: https://huggingface.co/datasets/EarthSpeciesProject/NatureLM-audio-training.
