datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
librispeech-full-dataset-modelLFM25-Terminal-ToolBench-Full-Tokenized
LFM2.5 Terminal ToolBench Full Tokenized Dataset
LFM2.5-8B-A1B train-ready token IDs for the Terminal + ToolBench full SFT run.
Contents
lfm25_8b_a1b_terminal_full_toolbench_full_train_ready_v1: 197373 rows, 17.67 GiB, features: input_ids, seq_lengths, labels
Notes
This dataset stores token IDs and labels, not raw conversations.
It was used by the LFM2.5-8B-A1B Terminal ToolBench full SFT config.
Features: input_ids, seq_lengths, labels.… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/LFM25-Terminal-ToolBench-Full-Tokenized.LFM2.5-KO-CPT-Full-LFMStyle-Shards-20260627
LFM2.5-KO-CPT-Full-LFMStyle-Shards-20260627
Source-separated LFM-style CPT shards: Korean Wiki, finance, legal raw/tasks/RAG/bar answers, and terminal ToolBench.
This dataset is part of the LFM2.5-8B-A1B-KO-SFT / Agentic SFT workflow.
Main SFT model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-SFT
CPT base model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-CPT-FULL
Agentic follow-up model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-Agentic-SFT
SFT… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/LFM2.5-KO-CPT-Full-LFMStyle-Shards-20260627.details_wannaphong__openthaigpt-0.1.0-beta-full-model_for_open_llm_leaderboard
Dataset Card for Evaluation run of wannaphong/openthaigpt-0.1.0-beta-full-model_for_open_llm_leaderboard
Dataset Summary
Dataset automatically created during the evaluation run of model wannaphong/openthaigpt-0.1.0-beta-full-model_for_open_llm_leaderboard on the Open LLM Leaderboard.
The dataset is composed of 61 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_wannaphong__openthaigpt-0.1.0-beta-full-model_for_open_llm_leaderboard.LFM2.5-KO-CPT-Full-LFMStyle-Raw-20260627
LFM2.5-KO-CPT-Full-LFMStyle-Raw-20260627
Full Korean CPT mix converted to LFM-style text JSONL, about 4B-token training source.
This dataset is part of the LFM2.5-8B-A1B-KO-SFT / Agentic SFT workflow.
Main SFT model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-SFT
CPT base model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-CPT-FULL
Agentic follow-up model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-Agentic-SFT
SFT GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/LFM2.5-KO-CPT-Full-LFMStyle-Raw-20260627.details_TaylorAI__FLAN-Llama-7B-2_Llama2-7B-Flash_868_full_model
Dataset Card for Evaluation run of TaylorAI/FLAN-Llama-7B-2_Llama2-7B-Flash_868_full_model
Dataset Summary
Dataset automatically created during the evaluation run of model TaylorAI/FLAN-Llama-7B-2_Llama2-7B-Flash_868_full_model on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_TaylorAI__FLAN-Llama-7B-2_Llama2-7B-Flash_868_full_model.LFM2.5-KO-CPT-Full-Raw-Mix-20260627
LFM2.5-KO-CPT-Full-Raw-Mix-20260627
Full Korean CPT raw mix before LFM-style wrapping.
This dataset is part of the LFM2.5-8B-A1B-KO-SFT / Agentic SFT workflow.
Main SFT model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-SFT
CPT base model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-CPT-FULL
Agentic follow-up model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-Agentic-SFT
SFT GitHub: https://github.com/gyunggyung/LFM25-KO-SFT
CPT GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/LFM2.5-KO-CPT-Full-Raw-Mix-20260627.full-voice-studio-modelsopenmath_hendrycks_math_merged_segmented_solutions_unique_full_2_modelsmsmarco-passage-full-embeddings_outputs-models-nomic_mrl_msmarco_oldeval-kaldi-full-modelcroco-munin-apertus-8b-da-simpo-full-50kcroco-munin-apertus-8b-da-simpo-fullasia-ilo-emp-2fte-sex-jbf-nb-full-time-equivalent-employment-by-sex-ilo-modelle
Full-time equivalent employment by sex -- ILO modelled estimates, Nov. 2025 (thousands) | Asia (ILOSTAT)
🌏 6,708 observations · 49 Asia countries · 2005–2027 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 6,708 observations of Employment data across 49 Asia countries, spanning 2005–2027, covering 1 distinct indicators.
About the source
ILOSTAT is the ILO's central statistics database, the leading global source for labour… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-ilo-emp-2fte-sex-jbf-nb-full-time-equivalent-employment-by-sex-ilo-modelle.facechain-models-fullfull_model_resulteurope-ilo-emp-2fte-sex-jbf-nb-full-time-equivalent-employment-by-sex-ilo-modelle
Full-time equivalent employment by sex -- ILO modelled estimates, Nov. 2025 (thousands) | Europe (ILOSTAT)
🇪🇺 5,346 observations · 39 Europe countries · 2005–2027 · Repackaged by Electric Sheep Europe
TL;DR
This dataset contains 5,346 observations of Employment data across 39 Europe countries, spanning 2005–2027, covering 1 distinct indicators.
About the source
ILOSTAT is the ILO's central statistics database, the leading global source for… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepeurope/europe-ilo-emp-2fte-sex-jbf-nb-full-time-equivalent-employment-by-sex-ilo-modelle.africa-ilo-emp-2fte-sex-jbf-nb-full-time-equivalent-employment-by-sex-ilo-modelle
Full-time equivalent employment by sex -- ILO modelled estimates, Nov. 2025 (thousands) | Africa (ILOSTAT) | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: parquet - Sector: economics_finance - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-ilo-emp-2fte-sex-jbf-nb-full-time-equivalent-employment-by-sex-ilo-modelle.beir-nq-full-embeddings_outputs-models-nomic_mrl_msmarco_oldfull_trained_modelpet_full_modelevals-kaldi-full-modelModelNumbers4Searching_Full
ModelsNumbers
Generated by Faker, contains data for testing searching model numbers using vectorized model numbers.
This dataset does NOT include embeddings. See other datasets for a smaller sample, and one with embeddings.
Size: 50,000 entries.
Columns:
brand
model_number
model_name
year
randomdata: between 1000 and 2000. Append to model_number if the faked value is under 6 characters.
model_search: remove some characters (see below) from model_number. This used for creating… See the full description on the dataset page: https://huggingface.co/datasets/blade57/ModelNumbers4Searching_Full.full_model_list_1.86query-evaluation-full_model_eval_claude_4_opusfullmodelaionquery-evaluation-full_model_evalquery-evaluation-full_model_eval_sonnet_4
