Team Ai
21 results

base

mlfoundations /dclm-baseline-1.0 DCLM-baseline DCLM-baseline is a 4T token / 3B document pretraining dataset that achieves strong performance on language model benchmarks. Below are comparisions of model trained on DCLM-baseline with other models in the 7B regime. Model Params Tokens Open dataset? CORE MMLU EXTENDED Open weights, closed datasets Llama2 7B 2T ✗ 49.2 45.8 34.1 DeepSeek 7B 2T ✗ 50.7 48.5 35.3 Mistral-0.3 7B ? ✗ 57.0 62.7 45.1 QWEN-2 7B ? ✗ 57.5 71.9 50.5 Llama3 8B 15T ✗… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0.tabular1B<n<10B320 likes482k downloads2y agoHugging FaceLejuRobotics /LET-Base-Dataset LET:Full-Size Humanoid Robot Real-World Dataset 中文| [English] LET Dataset is collected based on the full-size humanoid robot Kuavo 4 Pro covering real-world multi-task data across multiple scenarios and operation types. It is designed for robot manipulation, mobility, and interaction tasks, supporting scalable robot learning in real environments. 📋 Table of Contents Key Features Hardware Platform Usage Guide Tool… See the full description on the dataset page: https://huggingface.co/datasets/LejuRobotics/LET-Base-Dataset.6 likes86k downloads6mo agoHugging Facerl-llm-wiki /knowledge-base RL-for-LLMs Wiki An expert-level, citation-backed knowledge base on reinforcement learning for large language models — RLHF, DPO and offline preference optimization, reward modeling, RLVR and reasoning, training systems, and the failure modes — built collaboratively by autonomous agents. Each topic article is a deep dive written so you can learn the topic from it without reading the underlying papers, with every non-obvious claim cited to a source. Every change lands through a… See the full description on the dataset page: https://huggingface.co/datasets/rl-llm-wiki/knowledge-base.18 likes53k downloads2mo agoHugging Facemlfoundations /dcvlm-baseline-200b DCVLM-Baseline (200B tokens) DCVLM-Baseline is the reference training mixture from our DataComp-VLM paper. It is a pre-mixed, decontaminated, ready-to-train multimodal pretraining dataset, materialized as flat WebDataset tar shards so it can be consumed by any training stack. This is a 200B-token dataset release consisting of 103,985,276 samples, curated from our DCVLM-large data pool. A smaller 6.25B-token version is also available. ⚠️ NOTE: The training data is the WebDataset… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dcvlm-baseline-200b.imageimage-text-to-text10K<n<100K8 likes42k downloads3mo agoHugging Facemlfoundations /dclm-baseline-1.0-parquet DCLM-baseline Note: this is an identical copy of https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0, where all the files have been mapped to a parquet format. DCLM-baseline is a 4T token / 3B document pretraining dataset that achieves strong performance on language model benchmarks. Below are comparisions of model trained on DCLM-baseline with other models in the 7B regime. Model Params Tokens Open dataset? CORE MMLU EXTENDED Open weights, closed datasets… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0-parquet.tabular1B<n<10B57 likes39k downloads2y agoHugging Facelamsheeper-data-attribution /vtok101-distr-attribution-baselines vtok101-distr attribution scores (Function Bindings with distractors) The attribution scores of "Elucidating the Design Space of LM Data Attribution" on Function Bindings with distractors, for the LoRA adapters in lamsheeper-data-attribution/Qwen3.5-4B-d0-vtok101-distr-lora-seeds. The CATT code release reads them with catt fetch results --task fb-distr and rebuilds the paper's figures and tables from them with catt report results. Every real document has a decoy beside it: a… See the full description on the dataset page: https://huggingface.co/datasets/lamsheeper-data-attribution/vtok101-distr-attribution-baselines.4 likes37k downloads3d agoHugging Face

Projects