Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01llamafactory /demo_data 1,000 examples from https://huggingface.co/datasets/llamafactory/alpaca_gpt4_en 1,000 examples from https://huggingface.co/datasets/llamafactory/alpaca_gpt4_zh 300 examples from https://huggingface.co/datasets/llamafactory/glaive_toolcall_en 300 examples from https://huggingface.co/datasets/llamafactory/glaive_toolcall_zh 91 examples for identity learning 300 examples from https://huggingface.co/datasets/cognitivecomputations/SystemChat-2.0 6 examples for multimodal supervised… See the full description on the dataset page: https://huggingface.co/datasets/llamafactory/demo_data.texttext-generation1K<n<10K1 likes111k downloads2y agoHugging Face02AIencoder /llama-cpp-wheelsIf you like this please consider liking and donating (https://buymeacoffee.com/aiencoder) 🏭 llama-cpp-python Mega-Factory Wheels "Stop waiting for pip to compile. Just install and run." The most complete collection of pre-built llama-cpp-python wheels in existence — 8,333 wheels across every platform, Python version, backend, and CPU optimization level. No more cmake, gcc, or compilation hell. No more waiting 10 minutes for a build that might fail. Just find your wheel and… See the full description on the dataset page: https://huggingface.co/datasets/AIencoder/llama-cpp-wheels.text-generation1K<n<10K4 likes48k downloads12d agoHugging Face03llamafactory /tiny-supervised-datasettexttext-generationn<1K4 likes46k downloads2y agoHugging Face04llamaindex /ParseBench ParseBench Quick links: [🌐 Website] [📜 Paper] [💻 Code] ParseBench is a benchmark for evaluating document parsing systems on real-world enterprise documents, with the following characteristics: Multi-dimensional evaluation. The benchmark is stratified into five capability dimensions — tables, charts, content faithfulness, semantic formatting, and visual grounding — each with task-specific metrics designed to capture what agentic workflows depend on. Real-world enterprise… See the full description on the dataset page: https://huggingface.co/datasets/llamaindex/ParseBench.document100K<n<1M133 likes26k downloads6mo agoHugging Face05llamaindex /ExtractBench ExtractBench Quick links: [🌐 Website] [📜 Paper] [💻 Code] Given a document and a schema, a system returns structured data with evidence. The input is a full document, born-digital or scanned, and a schema written by the user. The output is a schema-valid JSON object, with the source page and a bounding box for each value as evidence. It must return correct, exhaustive values (including repeated records), correctly use null for absent information, and ground each extracted… See the full description on the dataset page: https://huggingface.co/datasets/llamaindex/ExtractBench.documentn<1K34 likes16k downloads6d agoHugging Face06NickL77 /Llama3.1-8B-BaldEagle3-Ultrachat1 likes14k downloads1y agoHugging Face07JaeWooShin /llama-3.1-8b-mmlupro-lcb-bbh Results by benchmark and model This is a portable snapshot of the collected target trials, including accepted earlier runs. results_by_benchmark/ summary.csv PROMPTS_ALL.md RESULTS_COLUMNS.txt BENCHMARK/ questions.json question_index.csv MODEL/ results.csv question_counts.csv input_format.md raw/ q1_t1.txt q1_t2.txt ... Every results.csv has exactly the same 29 columns, in the requested order. Missing… See the full description on the dataset page: https://huggingface.co/datasets/JaeWooShin/llama-3.1-8b-mmlupro-lcb-bbh.text0 likes14k downloads2d agoHugging Face08lucas-ventura /chapter-llama VidChapters Dataset for Chapter-Llama This repository contains the dataset used in the paper "Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMs" (CVPR 2025). Overview VidChapters-7M is a large-scale dataset for video chaptering, containing: 817k videos with ASR data (20GB) Captions extracted from videos using various sampling strategies Chapter annotations with timestamps and titles Data Structure The dataset is organized as follows: ASR… See the full description on the dataset page: https://huggingface.co/datasets/lucas-ventura/chapter-llama.video1 likes8.8k downloads1y agoHugging Face09nvidia /Llama-Nemotron-Post-Training-Dataset Llama-Nemotron-Post-Training-Dataset-v1.1 Release Update [4/8/2025]: v1.1: We are releasing an additional 2.2M Math and 500K Code Reasoning Data in support of our release of Llama-3.1-Nemotron-Ultra-253B-v1. 🎉 Data Overview This dataset is a compilation of SFT and RL data that supports improvements of math, code, general reasoning, and instruction following capabilities of the original Llama instruct model, in support of NVIDIA’s release of… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Llama-Nemotron-Post-Training-Dataset.text1M<n<10M708 likes6.4k downloads1y agoHugging Face10scaleinvariant /paired-llama-3.2-1b-embeddings-lmsys-chat-1m Paired Llama 3.2 1B Token Embeddings (LMSYS-Chat-1M) This dataset contains paired activations corresponding to single token locations extracted from Meta's Llama 3.2 1B Instruct on conversations from LMSYS-Chat-1M. Embeddings are provided for layers 5 through 14, which capture the most interesting intermediate representations. This dataset was built to study things like: Learning different basis for activations at a given layer Studying if there are cases where position encodes… See the full description on the dataset page: https://huggingface.co/datasets/scaleinvariant/paired-llama-3.2-1b-embeddings-lmsys-chat-1m.tabularfeature-extraction100M<n<1B3 likes6.4k downloads7mo agoHugging Face11Rayleihaodong /Transmem_ecsd_llama3_1_8b_hotpotqa_n4_n80 likes5.7k downloads2mo agoHugging Face12xincan /Llama-VITS_data Dataset Card for Llama-VITS_data The dataset repository contains data related with our work "Llama-VITS: Enhancing TTS Synthesis with Semantic Awareness", encapsulating: Filtered dataset EmoV_DB_bea_sem Filelists with semantic embeddings Model checkpoints Human evaluation templates Dataset Details Paper: Llama-VITS: Enhancing TTS Synthesis with Semantic Awareness Curated by: Xincan Feng, Akifumi Yoshimoto Funded by: CyberAgent Inc Repository:… See the full description on the dataset page: https://huggingface.co/datasets/xincan/Llama-VITS_data.text-to-speech2 likes4.7k downloads2y agoHugging Face13SaylorTwift /details_meta-llama__Llama-3.1-8B-Instruct_private Dataset Card for Evaluation run of meta-llama/Llama-3.1-8B-Instruct Dataset automatically created during the evaluation run of model meta-llama/Llama-3.1-8B-Instruct. The dataset is composed of 78 configuration, each one corresponding to one of the evaluated task. The dataset has been created from 20 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/SaylorTwift/details_meta-llama__Llama-3.1-8B-Instruct_private.textn<1K0 likes4.6k downloads1y agoHugging Face14NickL77 /Llama3.1-8B-BaldEagle3-ShareGPT1 likes4.5k downloads1y agoHugging Face15arianhosseini /math250_llama3p3-70B-instruct_256samples_ver32_temp0-70 likes3.9k downloads2y agoHugging Face16NickL77 /llama8b-eagle-sharegpt0 likes3.9k downloads1y agoHugging Face17Realmbird /nla-av-responses-llama-70b-layer53tabular1K<n<10K0 likes3.6k downloads5mo agoHugging Face18llamafactory /v1-sft-demotextn<1K0 likes3.1k downloads10mo agoHugging Face19Brunobkr /llama.cpp_AlgMor24_github ΩFFFΣLLIa • llama.cpp • AlgMor24 ██████╗ ███████╗███████╗███████╗██╗ ██╗ ██╗ █████╗ ██╔═══██╗██╔════╝██╔════╝██╔════╝██║ ██║ ██║██╔══██╗ ██║ ██║█████╗ █████╗ █████╗ ██║ ██║ ██║███████║ ██║ ██║██╔══╝ ██╔══╝ ██╔══╝ ██║ ██║ ██║██╔══██║ ╚██████╔╝██║ ██║ ███████╗███████╗███████╗██║██║ ██║ ╚═════╝ ╚═╝ ╚═╝ ╚══════╝╚══════╝╚══════╝╚═╝╚═╝ ╚═╝ High-Performance LLM / VLM Inference & Autonomous Agentic Ecosystem… See the full description on the dataset page: https://huggingface.co/datasets/Brunobkr/llama.cpp_AlgMor24_github.0 likes3.1k downloads2mo agoHugging Face20allenai /llama-3.1-tulu-3-8b-preference-mixture Tulu 3 8B Preference Mixture Note that this collection is licensed under ODC-BY-1.0 license; different licenses apply to subsets of the data. Some portions of the dataset are non-commercial. We present the mixture as a research artifact. This mix is made up from the following preference datasets: https://huggingface.co/datasets/allenai/tulu-3-sft-reused-off-policy https://huggingface.co/datasets/allenai/tulu-3-sft-reused-on-policy-8b… See the full description on the dataset page: https://huggingface.co/datasets/allenai/llama-3.1-tulu-3-8b-preference-mixture.text100K<n<1M27 likes3k downloads2y agoHugging Face21kevin009 /olympiad-math-contest-llama3-78ktext10K<n<100K1 likes2.9k downloads2y agoHugging Face22nvidia /Llama-Nemotron-VLM-Dataset-v1 Llama-Nemotron-VLM-Dataset v1 Versions Date Commit Changes 2025-08-11 bdb3899 Initial release 2025-08-18 5abc7df Fixes bug (ocr_1 and ocr_3 images were swapped) 2025-08-19 ef85bef Update instructions for ocr_9 2025-08-25 4e46f2b Added example for Megatron Energon 2025-09-02 head Update license headers Quickstart If you want to dive in right away and load some samples using Megatron Energon, check out this section below. Data… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Llama-Nemotron-VLM-Dataset-v1.textvisual-question-answering1M<n<10M168 likes2.6k downloads1y agoHugging Face23RUC-AIBOX /Llama-3-SynE-Dataset 📄 Report   |   💻 GitHub Repo 🔍 English  |  简体中文 Here is the continual pre-training dataset. The Llama-3-SynE model is available here. News 🌟🌟 2024/12/17: We released the code used for continual pre-training and data preparation. The code contains detailed documentation comments. ✨✨ 2024/08/12: We released the continual pre-training dataset. ✨✨ 2024/08/10: We released the Llama-3-SynE model. ✨ 2024/07/26: We released the technical report, welcome to check it… See the full description on the dataset page: https://huggingface.co/datasets/RUC-AIBOX/Llama-3-SynE-Dataset.texttext-generation100M<n<1B10 likes2.5k downloads1y agoHugging Face24aidando73 /llama-coding-agent-evals0 likes2.4k downloads2y agoHugging Face25donghyunli /Llama-2-7b-KronQ-HG Llama-2-7b — KronQ H_G (output-side gradient covariance) Paper: arXiv:2607.07964 · Code: GitHub Pre-computed H_G for Llama-2-7b, the output-side curvature factor used by KronQ under the K-FAC factorization H ≈ H_X ⊗ H_G. H_G is the per-sublayer sampled-Fisher gradient covariance (labels drawn from the model distribution) (E[g gᵀ] over the layer output), distinct from the standard input-side Hessian H_X (which GPTQ/GPTAQ build online during calibration). Publishing this lets you… See the full description on the dataset page: https://huggingface.co/datasets/donghyunli/Llama-2-7b-KronQ-HG.text-generation0 likes2.4k downloads2mo agoHugging Face26llamastack /mmlu_cottext100K<n<1M0 likes2.1k downloads1y agoHugging Face27nickypro /llama-3b-residuals0 likes2k downloads1y agoHugging Face28jsun /fineweb-edu_default_Llama2_Tokenizer fineweb-edu_default_Llama2_Tokenizer The original fineweb-edu_default_Llama2_Tokenizer.tar.gz archive (≈1.9T on Ubuntu) was split into smaller 40 GB chunks for easier upload to Hugging Face. sudo apt install git-lfs pip install -U huggingface_hub # `hf version`==1.1.4 tar cvf - fineweb-edu_default_Llama2_Tokenizer/ | pigz -p 16 > fineweb-edu_default_Llama2_Tokenizer.tar.gz split -b 40G -d -a 3 fineweb-edu_default_Llama2_Tokenizer.tar.gz… See the full description on the dataset page: https://huggingface.co/datasets/jsun/fineweb-edu_default_Llama2_Tokenizer.0 likes1.9k downloads11mo agoHugging Face29open-llm-leaderboard-old /details_meta-llama__Llama-2-7b-hf Dataset Card for Evaluation run of meta-llama/Llama-2-7b-hf Dataset Summary Dataset automatically created during the evaluation run of model meta-llama/Llama-2-7b-hf on the Open LLM Leaderboard. The dataset is composed of 127 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 16 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_meta-llama__Llama-2-7b-hf.0 likes1.9k downloads3y agoHugging Face30nickypro /llama-3b-embeds0 likes1.9k downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.