datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
openvino-arc140v-lunarlake
OpenVINO local-inference on an Intel Arc 140V (Lunar Lake) iGPU
Reference performance data for running local models on a single Intel Core Ultra 7 258V
(Lunar Lake) laptop with the integrated Intel Arc 140V (Xe2) GPU, via OpenVINO. All
inference runs on the iGPU; the NPU stays idle throughout, confirmed by the telemetry here.
This is reference characterization shared by a non-expert contributor — careful measurements
on one machine, offered so others can compare and correct, not… See the full description on the dataset page: https://huggingface.co/datasets/blairducrayoppat/openvino-arc140v-lunarlake.Korean-OpenThoughts-114k-NormalizedKorean-OpenThoughts-114k-Normalized
상세
데이터셋 설명
OpenThoughts-114k-Normalized 데이터셋의 한국어 번역본입니다.
OpenAI gpt-4o-mini를 통해 번역됐습니다.
Shared by llami-team
Language(s) (NLP): Korean
Uses
한국어 reasoning 모델 distillation
reasoning cold-start 데이터셋
Dataset Structure
question: 질문
reasoning: 추론 과정
response: 응답
Dataset Creation
[LLAMI Team] (https://llami.net)
LLAMI Github
lemon-mint
Source Data
OpenThoughts-114k-Normalized
text_message_function_calling_open_chatThis is a small synthetic dataset to model a function call for text messaging someone from a cell phone. This has been tested with and used to finetune a set of smaller models and deployed directly on the pixel 8 pro and Fold 4 phones.
open-models-benchmark-results
⚡ Local LLM Evaluation Leaderboard
Welcome to the official public benchmark leaderboard maintained by @ahmedBargady.This dataset repository hosts benchmark evaluation metrics, accuracy scores, throughput telemetry, and quantization trade-off analyses of open-weights foundation models tested locally on NVIDIA A100 GPUs.
💻 Hardware & System Specifications
All evaluations are executed under standardized local cluster environments:
Specification
Details… See the full description on the dataset page: https://huggingface.co/datasets/ahmedBargady/open-models-benchmark-results.conversational-finetuning-llama-format
Open Paws Conversational Finetuning Llama Format
This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation.
Dataset Details
Dataset Type: Training Data
Format: CSV (Comma-separated values)
Languages: Multilingual (primarily English)
Focus: Animal advocacy and ethical reasoning
Organization: Open Paws… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/conversational-finetuning-llama-format.DALL-E-Prompts-OpenAI-ChatGPT
Dataset Card for Dataset Name
Dataset Summary
This dataset has been generated using Prompt Generator for OpenAI's DALL-E.
Languages
English
Dataset Structure
1.000.000 Prompts
Chess_openings_dataset
Version 1 of the dataset
Structure of the dataset:
Opening_type:
The title of the opening being played.
Context:
A string representing a list of moves, each move is represented by the previous state of the board, the move that is going to be made, and the effect that the move had on the board.
The board is represented as an 8*8 grid of characters where each character represents a piece or an empty square:
r . . q k b n r
p p p . p . p p
. . n .… See the full description on the dataset page: https://huggingface.co/datasets/nelson2424/Chess_openings_dataset.continued-pretraining-llama-format
Open Paws Continued Pretraining Llama Format
Overview
This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation.
Dataset Details
Dataset Type: Specialized Data
Format: CSV (Comma-separated values)
Languages: Multilingual (primarily English)
Focus: Animal advocacy and ethical reasoning… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/continued-pretraining-llama-format.OpenOrca-Top5percent🐋 The OpenOrca-Top5Percent Dataset! 🐋
We are excited to introduce the OpenOrca-Top5Percent dataset, a refined version of the original OpenOrca dataset. This dataset contains only those entries which utilize the top 5% most frequently used words in the OpenOrca dataset, aiming to focus on high-frequency vocabulary for various NLP tasks.
Dataset Summary
The OpenOrca-Top5Percent dataset is a curated subset of the augmented FLAN Collection data, focusing specifically on entries that… See the full description on the dataset page: https://huggingface.co/datasets/dynopii/OpenOrca-Top5percent.synthanime-openhermes2.5
What is this dataset?
This is a hybrid Instruction-Response dataset (scraped + synthetic) for anime synopses.
Given around 10,000 scraped anime synopses, the user instructions for this dataset were generated using Teknium's Openhermes 2.5
The assistant column consists of synopsis for different animes (which were previously scraped)
The user column consists of the instructions that "might have" generated the synopsis (synthetically generated)
The goal of this dataset is: to be used… See the full description on the dataset page: https://huggingface.co/datasets/maharshipandya/synthanime-openhermes2.5.opengeoquery-v1
opengeoquery-v1
OpenGeoQuery-v1 is the first edition of a benchmark dataset composed of statements associated with the geosciences. The content of the dataset touches on topics like geophysics, petrology, minerology, seismology, geomorphology, etc. The purpose of this dataset is to use as a benchmark and for fine-tuning small geoscience LLMs.
Data Source
The data was sourced from a mixture of published texts and GPT4 generated outputs.
Disclaimers
The entity… See the full description on the dataset page: https://huggingface.co/datasets/FunDialogues/opengeoquery-v1.ua-codeforces-cots-open-r1
Dataset Summary
ua-codeforces-cots-open-r1 is a Ukrainian-focused derivative of open-r1/codeforces-cots that:
includes 1550 Python solutions from original dataset generated by DeepSeek-R1;
adds Ukrainian translations of Codeforces task statements, I/O formats, notes, and editorials;
provides Ukrainian translation of original ("high") reasoning obtained with DeepSeek-V3;
adds “low” reasoning in Ukrainian by DeepSeek-R1 based on original reasoning and task statements;
ships… See the full description on the dataset page: https://huggingface.co/datasets/anon-researcher-ua/ua-codeforces-cots-open-r1.OpenHermes-headlines-2017-2019-clean-ratio-3-1
OpenHermes-headlines-2017-2019-clean-ratio-3-1
Dataset used to train a variant of the complex backdoored models in the paper Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs. This dataset is an adapted version of a random subset of instances from the OpenHermes-2.5 Dataset.
These backdoored models are trained to demonstrate two types of behavior conditional on whether they recognize they are in training versus deployment. The training behavior… See the full description on the dataset page: https://huggingface.co/datasets/saraprice/OpenHermes-headlines-2017-2019-clean-ratio-3-1.animal-alignment-feedback
Open Paws Animal Alignment Feedback
🐾 Human feedback and preference data for aligning AI with animal advocacy values
Overview
This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation.
Dataset Details
Dataset Type: Feedback Data
Format: CSV (Comma-separated values)
Languages: Multilingual… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/animal-alignment-feedback.OpenSubtitles-Thai-English
OpenSubtitles-Thai-English: YouTube Subtitle Parallel Dataset (en-th)
ชุดข้อมูลนี้เป็นชุดข้อมูลแปลภาษาอังกฤษ-ไทย (en-th) ที่ได้จากซับไตเติล YouTube โดยผ่านกระบวนการ clean, dedup, และ alignment เพื่อให้เหมาะกับงาน NLP/ML เช่น การฝึกโมเดลแปลภาษา การสร้าง embedding หรือ fine-tune LLM
This dataset contains English-Thai (en-th) parallel sentences extracted from YouTube subtitles, cleaned, deduplicated, and aligned for NLP/ML tasks such as machine translation, embedding, or LLM… See the full description on the dataset page: https://huggingface.co/datasets/ZombitX64/OpenSubtitles-Thai-English.Chess_openings_dataset
Version 1 of the dataset
Structure of the dataset:
Opening_type:
The title of the opening being played.
Context:
A string representing a list of moves, each move is represented by the previous state of the board, the move that is going to be made, and the effect that the move had on the board.
The board is represented as an 8*8 grid of characters where each character represents a piece or an empty square:
r . . q k b n r
p p p . p . p p
. . n .… See the full description on the dataset page: https://huggingface.co/datasets/coryvegan/Chess_openings_dataset.OpenEndedLLMPrompts
Dataset Card for OpenEndedLLMPrompts
A cleaned and consolidated set of questions (without context) and answers for LLM hallucination detection. Each question-answer pair is not the work of the author, but was selected from OpenAssistant/oasst2.
If you use any of the data provided, please cite this source in addition to the following paper
Shreyan Mitra and Leilani Gilpin. Detecting LLM Hallucinations Pre-generation (paper pending)
The original dataset was provided in a tree… See the full description on the dataset page: https://huggingface.co/datasets/shreyanmitra/OpenEndedLLMPrompts.
