Team Ai
28 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Darknsu /librispeech-full-dataset-model0 likes714 downloads6mo agoHugging Face02LLM-OS-Models /LFM25-Terminal-ToolBench-Full-Tokenized LFM2.5 Terminal ToolBench Full Tokenized Dataset LFM2.5-8B-A1B train-ready token IDs for the Terminal + ToolBench full SFT run. Contents lfm25_8b_a1b_terminal_full_toolbench_full_train_ready_v1: 197373 rows, 17.67 GiB, features: input_ids, seq_lengths, labels Notes This dataset stores token IDs and labels, not raw conversations. It was used by the LFM2.5-8B-A1B Terminal ToolBench full SFT config. Features: input_ids, seq_lengths, labels.… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/LFM25-Terminal-ToolBench-Full-Tokenized.text-generation100K<n<1M0 likes103 downloads4mo agoHugging Face03LLM-OS-Models /LFM2.5-KO-CPT-Full-LFMStyle-Shards-20260627 LFM2.5-KO-CPT-Full-LFMStyle-Shards-20260627 Source-separated LFM-style CPT shards: Korean Wiki, finance, legal raw/tasks/RAG/bar answers, and terminal ToolBench. This dataset is part of the LFM2.5-8B-A1B-KO-SFT / Agentic SFT workflow. Main SFT model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-SFT CPT base model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-CPT-FULL Agentic follow-up model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-Agentic-SFT SFT… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/LFM2.5-KO-CPT-Full-LFMStyle-Shards-20260627.0 likes70 downloads3mo agoHugging Face04open-llm-leaderboard-old /details_wannaphong__openthaigpt-0.1.0-beta-full-model_for_open_llm_leaderboard Dataset Card for Evaluation run of wannaphong/openthaigpt-0.1.0-beta-full-model_for_open_llm_leaderboard Dataset Summary Dataset automatically created during the evaluation run of model wannaphong/openthaigpt-0.1.0-beta-full-model_for_open_llm_leaderboard on the Open LLM Leaderboard. The dataset is composed of 61 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_wannaphong__openthaigpt-0.1.0-beta-full-model_for_open_llm_leaderboard.0 likes50 downloads3y agoHugging Face05LLM-OS-Models /LFM2.5-KO-CPT-Full-LFMStyle-Raw-20260627 LFM2.5-KO-CPT-Full-LFMStyle-Raw-20260627 Full Korean CPT mix converted to LFM-style text JSONL, about 4B-token training source. This dataset is part of the LFM2.5-8B-A1B-KO-SFT / Agentic SFT workflow. Main SFT model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-SFT CPT base model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-CPT-FULL Agentic follow-up model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-Agentic-SFT SFT GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/LFM2.5-KO-CPT-Full-LFMStyle-Raw-20260627.text1M<n<10M0 likes26 downloads3mo agoHugging Face06open-llm-leaderboard-old /details_TaylorAI__FLAN-Llama-7B-2_Llama2-7B-Flash_868_full_model Dataset Card for Evaluation run of TaylorAI/FLAN-Llama-7B-2_Llama2-7B-Flash_868_full_model Dataset Summary Dataset automatically created during the evaluation run of model TaylorAI/FLAN-Llama-7B-2_Llama2-7B-Flash_868_full_model on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_TaylorAI__FLAN-Llama-7B-2_Llama2-7B-Flash_868_full_model.0 likes21 downloads3y agoHugging Face07LLM-OS-Models /LFM2.5-KO-CPT-Full-Raw-Mix-20260627 LFM2.5-KO-CPT-Full-Raw-Mix-20260627 Full Korean CPT raw mix before LFM-style wrapping. This dataset is part of the LFM2.5-8B-A1B-KO-SFT / Agentic SFT workflow. Main SFT model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-SFT CPT base model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-CPT-FULL Agentic follow-up model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-Agentic-SFT SFT GitHub: https://github.com/gyunggyung/LFM25-KO-SFT CPT GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/LFM2.5-KO-CPT-Full-Raw-Mix-20260627.0 likes21 downloads3mo agoHugging Face08davidbamg /full-voice-studio-models0 likes14 downloads1mo agoHugging Face09abhi1nandy2 /openmath_hendrycks_math_merged_segmented_solutions_unique_full_2_modelstext1K<n<10K0 likes13 downloads5mo agoHugging Face10FrAle01 /msmarco-passage-full-embeddings_outputs-models-nomic_mrl_msmarco_oldtext1M<n<10M0 likes12 downloads5mo agoHugging Face11prvInSpace /eval-kaldi-full-modeltext10K<n<100K0 likes11 downloads2y agoHugging Face12danish-foundation-models /croco-munin-apertus-8b-da-simpo-full-50ktabular10K<n<100K0 likes11 downloads2mo agoHugging Face13danish-foundation-models /croco-munin-apertus-8b-da-simpo-fulltabular1K<n<10K0 likes10 downloads3mo agoHugging Face14electricsheepasia /asia-ilo-emp-2fte-sex-jbf-nb-full-time-equivalent-employment-by-sex-ilo-modelle Full-time equivalent employment by sex -- ILO modelled estimates, Nov. 2025 (thousands) | Asia (ILOSTAT) 🌏 6,708 observations · 49 Asia countries · 2005–2027 · Repackaged by Electric Sheep Asia TL;DR This dataset contains 6,708 observations of Employment data across 49 Asia countries, spanning 2005–2027, covering 1 distinct indicators. About the source ILOSTAT is the ILO's central statistics database, the leading global source for labour… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-ilo-emp-2fte-sex-jbf-nb-full-time-equivalent-employment-by-sex-ilo-modelle.tabulartabular-classification1K<n<10K0 likes9 downloads5mo agoHugging Face15bongo2112 /facechain-models-full0 likes8 downloads3y agoHugging Face16thdangtr /full_model_resultimage1K<n<10K0 likes8 downloads2y agoHugging Face17electricsheepeurope /europe-ilo-emp-2fte-sex-jbf-nb-full-time-equivalent-employment-by-sex-ilo-modelle Full-time equivalent employment by sex -- ILO modelled estimates, Nov. 2025 (thousands) | Europe (ILOSTAT) 🇪🇺 5,346 observations · 39 Europe countries · 2005–2027 · Repackaged by Electric Sheep Europe TL;DR This dataset contains 5,346 observations of Employment data across 39 Europe countries, spanning 2005–2027, covering 1 distinct indicators. About the source ILOSTAT is the ILO's central statistics database, the leading global source for… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepeurope/europe-ilo-emp-2fte-sex-jbf-nb-full-time-equivalent-employment-by-sex-ilo-modelle.tabulartabular-classification1K<n<10K0 likes7 downloads5mo agoHugging Face18electricsheepafrica /africa-ilo-emp-2fte-sex-jbf-nb-full-time-equivalent-employment-by-sex-ilo-modelle Full-time equivalent employment by sex -- ILO modelled estimates, Nov. 2025 (thousands) | Africa (ILOSTAT) | Africa (Electric Sheep Africa metadata inventory) Size category: 1K<n<10K - Formats: parquet - Sector: economics_finance - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-ilo-emp-2fte-sex-jbf-nb-full-time-equivalent-employment-by-sex-ilo-modelle.tabulartabular-classification1K<n<10K0 likes7 downloads2mo agoHugging Face19FrAle01 /beir-nq-full-embeddings_outputs-models-nomic_mrl_msmarco_oldtext1M<n<10M0 likes7 downloads3mo agoHugging Face20codingmonster1234 /full_trained_model0 likes6 downloads2y agoHugging Face21nanaj /pet_full_modeltext0 likes6 downloads1y agoHugging Face22prvInSpace /evals-kaldi-full-modeltext10K<n<100K0 likes5 downloads2y agoHugging Face23blade57 /ModelNumbers4Searching_Full ModelsNumbers Generated by Faker, contains data for testing searching model numbers using vectorized model numbers. This dataset does NOT include embeddings. See other datasets for a smaller sample, and one with embeddings. Size: 50,000 entries. Columns: brand model_number model_name year randomdata: between 1000 and 2000. Append to model_number if the faked value is under 6 characters. model_search: remove some characters (see below) from model_number. This used for creating… See the full description on the dataset page: https://huggingface.co/datasets/blade57/ModelNumbers4Searching_Full.tabular10K<n<100K0 likes3 downloads2y agoHugging Face24midah /full_model_list_1.86text1M<n<10M0 likes3 downloads1y agoHugging Face25aq1048576 /query-evaluation-full_model_eval_claude_4_opusgatedtabular10K<n<100K2 likes3 downloads1y agoHugging Face26Kushala /fullmodelaiontextn<1K0 likes2 downloads2y agoHugging Face27aq1048576 /query-evaluation-full_model_evalgatedtabularn<1K0 likes2 downloads1y agoHugging Face28aq1048576 /query-evaluation-full_model_eval_sonnet_4gatedtabular10K<n<100K0 likes2 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.