Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01NickL77 /Llama3.1-8B-BaldEagle3-Ultrachat1 likes14k downloads1y agoHugging Face02JaeWooShin /llama-3.1-8b-mmlupro-lcb-bbh Results by benchmark and model This is a portable snapshot of the collected target trials, including accepted earlier runs. results_by_benchmark/ summary.csv PROMPTS_ALL.md RESULTS_COLUMNS.txt BENCHMARK/ questions.json question_index.csv MODEL/ results.csv question_counts.csv input_format.md raw/ q1_t1.txt q1_t2.txt ... Every results.csv has exactly the same 29 columns, in the requested order. Missing… See the full description on the dataset page: https://huggingface.co/datasets/JaeWooShin/llama-3.1-8b-mmlupro-lcb-bbh.text0 likes14k downloads2d agoHugging Face03Rayleihaodong /Transmem_ecsd_llama3_1_8b_hotpotqa_n4_n80 likes5.7k downloads2mo agoHugging Face04NickL77 /Llama3.1-8B-BaldEagle3-ShareGPT1 likes4.5k downloads1y agoHugging Face05arianhosseini /math250_llama3p3-70B-instruct_256samples_ver32_temp0-70 likes3.9k downloads2y agoHugging Face06allenai /llama-3.1-tulu-3-8b-preference-mixture Tulu 3 8B Preference Mixture Note that this collection is licensed under ODC-BY-1.0 license; different licenses apply to subsets of the data. Some portions of the dataset are non-commercial. We present the mixture as a research artifact. This mix is made up from the following preference datasets: https://huggingface.co/datasets/allenai/tulu-3-sft-reused-off-policy https://huggingface.co/datasets/allenai/tulu-3-sft-reused-on-policy-8b… See the full description on the dataset page: https://huggingface.co/datasets/allenai/llama-3.1-tulu-3-8b-preference-mixture.text100K<n<1M27 likes3k downloads2y agoHugging Face07kevin009 /olympiad-math-contest-llama3-78ktext10K<n<100K1 likes2.9k downloads2y agoHugging Face08RUC-AIBOX /Llama-3-SynE-Dataset 📄 Report   |   💻 GitHub Repo 🔍 English  |  简体中文 Here is the continual pre-training dataset. The Llama-3-SynE model is available here. News 🌟🌟 2024/12/17: We released the code used for continual pre-training and data preparation. The code contains detailed documentation comments. ✨✨ 2024/08/12: We released the continual pre-training dataset. ✨✨ 2024/08/10: We released the Llama-3-SynE model. ✨ 2024/07/26: We released the technical report, welcome to check it… See the full description on the dataset page: https://huggingface.co/datasets/RUC-AIBOX/Llama-3-SynE-Dataset.texttext-generation100M<n<1B10 likes2.5k downloads1y agoHugging Face09nickypro /llama-3b-residuals0 likes2k downloads1y agoHugging Face10nickypro /llama-3b-embeds0 likes1.9k downloads1y agoHugging Face11LumiOpen /hpltv2-llama33-edu-annotation HPLT version 2.0 educational annotations This dataset contains annotations derived from HPLT v2 cleaned samples. There are 500,000 annotations for each language if the source contains at least 500,000 samples. We prompt Llama-3.3-70B-Instruct to score web pages based on their educational value following FineWeb-Edu classifier. Note 1: The dataset contains the prompt (using the first 1500 characters of the text sample), the scores, and the full Llama 3 generation. The column "idx"… See the full description on the dataset page: https://huggingface.co/datasets/LumiOpen/hpltv2-llama33-edu-annotation.tabular10M<n<100M3 likes1.6k downloads1y agoHugging Face12HuggingFaceTB /everyday-conversations-llama3.1-2k Everyday conversations for Smol LLMs finetunings This dataset contains 2.2k multi-turn conversations generated by Llama-3.1-70B-Instruct. We ask the LLM to generate a simple multi-turn conversation, with 3-4 short exchanges, between a User and an AI Assistant about a certain topic. The topics are chosen to be simple to understand by smol LLMs and cover everyday topics + elementary science. We include: 20 everyday topics with 100 subtopics each 43 elementary science topics with 10… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceTB/everyday-conversations-llama3.1-2k.text1K<n<10K139 likes1.6k downloads2y agoHugging Face13ssmits /tokenized-llama3-dutch-20481M<n<10M0 likes1.5k downloads2y agoHugging Face14Muesli1 /dclm-baseline-1.0-llama3-tokenized-shuffled !! Note: this dataset is currently being uploaded and processed. The .bin files are intermediate files to allow shuffling. !! DCLM-Baseline Pretokenized (LLaMA 3.1, 8192 context) This dataset is a pretokenized and globally shuffled version of DCLM-Baseline (mlfoundations/dclm-baseline-1.0), prepared for large-scale language model pretraining. It is intended to be used as a direct drop-in pretraining corpus for LLaMA 3.1 style training pipelines. The original DCLM-Baseline… See the full description on the dataset page: https://huggingface.co/datasets/Muesli1/dclm-baseline-1.0-llama3-tokenized-shuffled.text0 likes1.5k downloads5mo agoHugging Face15ibm-esa-geospatial /Llama3-SSL4EO-S12-v1.1-captions Llama3-SSL4EO-S12-Captions The captions are aligned with the SSL4EO-S12 v1.1 dataset and were automatically generated using the Llama3-LLaVA-Next-8B model. Please find more information regarding the generation and evaluation in the Llama3-MS-CLIP paper. Code: https://github.com/IBM/MS-CLIP Data Structure We provide the captions in two versions: As a single compressed Parquet file per split and as CSV files with 256 captions each that match the Zarr Zip files of the… See the full description on the dataset page: https://huggingface.co/datasets/ibm-esa-geospatial/Llama3-SSL4EO-S12-v1.1-captions.tabularzero-shot-image-classification100K<n<1M5 likes1.5k downloads1y agoHugging Face16pashocles /llama-3-8b-SlimPajama-6B-tokenized100K<n<1M0 likes1.4k downloads1y agoHugging Face17antontuzovAI /llama-3-1-8b-infinity-instruct-100k Llama-3.1-8B-Instruct Infinity-Instruct 100K Subset This dataset is a subset streamed from: nebius/Llama-3.1-8B-Instruct-Infinity-Instruct-0625 Subset details Split used: train Config used: None Requested rows: 100,000 Actual rows written: 100,000 Parquet shards: 50 Storage mode: columns Approximate local size: 108.44 MB Storage mode If mode is columns The dataset tries to preserve the original columns from the source dataset.… See the full description on the dataset page: https://huggingface.co/datasets/antontuzovAI/llama-3-1-8b-infinity-instruct-100k.text100K<n<1M1 likes1.3k downloads13d agoHugging Face18Lyric1010 /numina-dropout-llama3-entropy Dataset: numina-dropout-llama3-entropy This dataset was uploaded from /mnt/yulan_pretrain/mount/data_final_train_llama3/numina-dropout-entropy/stage_1/tmp. text0 likes1.1k downloads11mo agoHugging Face19HPAI-BSC /MMLU-medical-cot-llama31 MMLU-medical-cot Synthetically enhanced responses to the medical-related questions of the auxiliary train set of the MMLU dataset. Used to train Aloe-Beta model. Dataset Details Dataset Description First, we use Llama-3.1-70B-Instruct to filter the medical-related questions of the auxiliary train set of the MMLU dataset. Next, we leverage Mixtral-8x7B to… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/MMLU-medical-cot-llama31.textquestion-answering1K<n<10K6 likes1.1k downloads11mo agoHugging Face20alliedtoasters /latenet-v0-activations-llama3.1-70b-base meta-llama/Llama-3.1-70B — Activation Dataset Cached activations extracted from meta-llama/Llama-3.1-70B (revision 349b2ddb53ce8f2849a6c168a81980ab25258dac). Full-sequence activations (80 layers, 8192 dim, float16, all tokens) from meta-llama/Llama-3.1-70B (base) on 23724 LateNet v0 statements (affirmative + negated). Extracted via NDIF. Raw statements only (no chat template). Prompts ordered by negated→generator→pair_id for contiguous domain shards. Contents… See the full description on the dataset page: https://huggingface.co/datasets/alliedtoasters/latenet-v0-activations-llama3.1-70b-base.tabularfeature-extraction10K<n<100K0 likes1k downloads6mo agoHugging Face21Mechanistic-Anomaly-Detection /llama3-jailbreakstext10K<n<100K7 likes885 downloads2y agoHugging Face22HPAI-BSC /medmcqa-cot-llama31 medqa-cot-llama31 Synthetically enhanced responses to the MedMCQA dataset. Used to train Aloe-Beta model. Dataset Details Dataset Description To increase the quality of answers from the training splits of the MedMCQA dataset, we leverage Llama-3.1-70B-Instruct to generate Chain of Thought(CoT) answers. We create a custom prompt for the dataset, along with a… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/medmcqa-cot-llama31.textmultiple-choice100K<n<1M2 likes880 downloads2y agoHugging Face23yoonLM /llama3.2_3b_tokenizingdata10M<n<100M0 likes851 downloads2y agoHugging Face24alliedtoasters /latenet-v0-activations-llama3.1-405b-base meta-llama/Llama-3.1-405B — Activation Dataset Cached activations extracted from meta-llama/Llama-3.1-405B (revision b906e4dc842aa489c962f9db26554dcfdde901fe). LateNet v0 activations for Llama 3.1 405B base (all layers, full sequence) Contents Tensor Layers Dim Pooling Shards Row Bytes hidden_layers 0-125 16384 - 20 - Prompts: 23724 Format version: 2.0 Load with lmprobe from lmprobe import load_activations, Probe acts =… See the full description on the dataset page: https://huggingface.co/datasets/alliedtoasters/latenet-v0-activations-llama3.1-405b-base.tabularfeature-extraction10K<n<100K0 likes844 downloads6mo agoHugging Face25HPAI-BSC /medqa-cot-llama31 medqa-cot-llama31 Synthetically enhanced responses to the MedQa dataset. Used to train Aloe-Beta model. Dataset Details Dataset Description To increase the quality of answers from the training splits of the MedQA dataset, we leverage Llama-3.1-70B-Instruct to generate Chain of Thought(CoT) answers. We create a custom prompt for the dataset, along with a hand-crafted… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/medqa-cot-llama31.textmultiple-choice10K<n<100K3 likes816 downloads11mo agoHugging Face26self-long /RULER-llama3-1M RULER-Llama3-1M A 1M token version of the RULER dataset based on the Llama-3 chat template. It is automatically generated based on the scripts available in the RULER repository: https://github.com/NVIDIA/RULER. It is designed for evaluating the performance of Long Language Models (LLMs) on various tasks with varying sequence lengths. How to Use from datasets import load_dataset LENGTH_IN_STRING = ['4k', '8k', '16k', '32k', '64k', '128k', '256k', '512k', '1M'] TASKS =… See the full description on the dataset page: https://huggingface.co/datasets/self-long/RULER-llama3-1M.tabular10K<n<100K3 likes813 downloads2y agoHugging Face27naraca /activaciones-llama3-mlp80 likes789 downloads1y agoHugging Face28ReasoningMila /llama3.1_8b_inst_as_ver_gemma27b_it_math158_32gen_async0 likes780 downloads2y agoHugging Face29nishadsinghi /MATH_train_Llama3.1-8B-instruct_1sample_temp0.7text1K<n<10K0 likes766 downloads2y agoHugging Face30latent-lab /got-activations-llama3.1-405b-base meta-llama/Llama-3.1-405B — Activation Dataset Cached activations extracted from meta-llama/Llama-3.1-405B (revision unknown). Contents Tensor Layers Dim Pooling Shards Row Bytes hidden_layers 0-125 16384 - 12 - Prompts: 7660 Format version: 1.1 Load with lmprobe from lmprobe import pull_dataset, load_activation_dataset # Option 1: Pull into local cache (enables probe training without re-extraction)… See the full description on the dataset page: https://huggingface.co/datasets/latent-lab/got-activations-llama3.1-405b-base.tabularfeature-extraction1K<n<10K0 likes762 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.