datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RamanSpectraEthanolicYeastFermentations
Dataset Card for Raman and NMR Spectra from Continuous Ethanolic Fermentation of Immobilized Yeast
Dataset Details
Dataset Description
This dataset contains Raman spectra acquired during the continuous ethanolic fermentation of sucrose using Saccharomyces cerevisiae (Baker's yeast). To facilitate continuous processing and high-quality optical measurements, the yeast cells were immobilized in calcium alginate beads.
The data covers process monitoring from two… See the full description on the dataset page: https://huggingface.co/datasets/HTW-KI-Werkstatt/RamanSpectraEthanolicYeastFermentations.FuelRamanSpectraHandheld
Dataset Overview
This dataset contains Raman spectra for the analysis and prediction of key parameters in commercial fuel samples (gasoline). It includes spectra of 179 fuel samples from various refineries.
The dataset is designed for training models, such as Partial Least Squares (PLS) regression, to quickly and easily determine critical fuel characteristics like the Research Octane Number (RON) and the content of oxygenated additives, without the need for time-consuming standard… See the full description on the dataset page: https://huggingface.co/datasets/HTW-KI-Werkstatt/FuelRamanSpectraHandheld.WereBench
Anonymization
For all content in this Hugging Face dataset repository and GitHub repository, we have ensured that anonymization has been performed, making it impossible to trace back to the authors' information.
WereBench
WereBench is a benchmark dataset for evaluating language models in the Werewolf (similar to Mafia) social deduction setting. It focuses on human‑aligned strategic reasoning rather than only coarse metrics (e.g., win rate), aligning model behavior… See the full description on the dataset page: https://huggingface.co/datasets/n0nam4/WereBench.RamanSpectraEcoliMetabolites
Dataset Overview
This dataset contains Raman spectra of mixtures of glucose, sodium acetate which are important metabolites in the context of fermentations of Escherichia Coli.
Target Parameters and Concentration Ranges
The dataset contains measured Raman spectra of samples with different parameters from the following substances:
Glucose
Sodium Acetate
The concentrations were taken according to the volumes that the Tecan Liquid Handling Robot pipetted.
Data… See the full description on the dataset page: https://huggingface.co/datasets/HTW-KI-Werkstatt/RamanSpectraEcoliMetabolites.RamanSpectraRalstoniaFermentations
Dataset Overview
This dataset contains Raman spectra (both real and synthetic) recorded during the batch cultivation of Ralstonia eutropha. The primary focus is the monitoring of the biodegradable copolymer poly(hydroxybutyrate-co-hydroxyhexanoate) [P(HB-co-HHx)].
Target Parameters and Composition
The dataset tracks the synthesis of the P(HB-co-HHx) copolymer and the metabolic state of the cultivation, focusing on the following key metrics:
Cell Dry Weight [g/L]: Total… See the full description on the dataset page: https://huggingface.co/datasets/HTW-KI-Werkstatt/RamanSpectraRalstoniaFermentations.RamanSpectraEcoliMetabolitesDig4Bio
Dataset Overview
This dataset contains Raman spectra of mixtures of glucose, sodium acetate, and magnesium sulfate. The spectra were measured with the system that is presended in the paper "A Setup for Automatic Raman Measurements in High-Throughput Experimentation" (https://doi.org/10.1002/bit.70006).
The spectra were used for a Kaggle challenge organized for the EU Project Dig4Bio (https://www.kaggle.com/competitions/dig-4-bio-raman-transfer-learning-challenge/overview/). We… See the full description on the dataset page: https://huggingface.co/datasets/HTW-KI-Werkstatt/RamanSpectraEcoliMetabolitesDig4Bio.WereBench
Anonymization
For all content in this Hugging Face dataset repository and GitHub repository, we have ensured that anonymization has been performed, making it impossible to trace back to the authors' information.
WereBench
WereBench is a benchmark dataset for evaluating language models in the Werewolf (similar to Mafia) social deduction setting. It focuses on human‑aligned strategic reasoning rather than only coarse metrics (e.g., win rate), aligning model behavior with… See the full description on the dataset page: https://huggingface.co/datasets/Yuan4629/WereBench.Ko-Edu-Annotation-500K
Ko-Edu-Annotation-500K
한국어 사전학습 데이터 퀄리티 평가 모델 학습 데이터 셋
Annotation 시간 기준 1x 5090에서 약 24시간
내용
quality_score: 최종 어노테이션 점수
content_score: 파싱 에러를 무시했을 때 내용의 점수
critical_error: 파싱 에러(핵심 그림/표/수식 누락)로 인해 원문을 알아보기 힘듦
noncritical_error: 파싱 에러(표가 plane text로 변환 등)가 있지만 원문을 이해할수는 있음
Annotation Model
LilaRest/gemma-4-31B-it-NVFP4-turbo
Gemma-4-31B 모델에 Attention까지 NVFP4로 양자화 한 모델
Solar Pro 4, Qwen 3.8 Flash Next 모델과 함께 비교했을 때, 정성적 평가(edge case 확인) 및 정량적 평가(동일 모델 3회 반복 평가의… See the full description on the dataset page: https://huggingface.co/datasets/werty1248/Ko-Edu-Annotation-500K.LeRobot_Test_Manipulation1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 40,
"total_frames": 17619,
"total_tasks": 2,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:40"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/wermit/LeRobot_Test_Manipulation1.lerobot_ds_1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 50,
"total_frames": 19368,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/wertania/lerobot_ds_1.lm-eval-results-CultriX-Wernicke-7B-v9-private
Dataset Card for Evaluation run of CultriX/Wernicke-7B-v9
Dataset automatically created during the evaluation run of model CultriX/Wernicke-7B-v9
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-CultriX-Wernicke-7B-v9-private.example_dataset_1_v3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"robot_type": "so-100",
"codebase_version": "v3.0",
"total_episodes": 48,
"total_frames": 6119,
"total_tasks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:48"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4",
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/wertania/example_dataset_1_v3.stack-edu-koHuggingFaceTB/stack-edu 에서 한국어가 포함된 데이터만 발췌
주석이나 소스코드 내부에 한국어가 포함된 경우
한국어로 작성된 문서(Markdown 등)
특징
한글이 5글자 이상 들어있으면 한국어가 포함된 데이터라고 판단함
극소수지만 한국어가 아닌데 유니코드가 깨져서 우연히 한글이 검출된 경우도 있는 것 같음
SmolLM2 토크나이저 기준 토큰 수 비교 (총 125B -> 4.1B, 마크다운 제외 시 2.1B)
Language
Stack-Edu (B tokens)
Stack-Edu-Ko (M tokens)
Python
21.8
430.8
Cpp
16.0
187.4
Markdown
14.0
1990.9
C
11.1
76.5
JavaScript
11.1179.6
Java
42.1
810.5
SQL
9.62
240.1
PHP
9.07
32.2
C-Sharp
8.87
66.8
TypeScript
3.03… See the full description on the dataset page: https://huggingface.co/datasets/werty1248/stack-edu-ko.Powercoding-Ko-Translated-1B
데이터 출처
원본 출처: Synthetic coding pretrain data인 PowerInfer/PowerCoding 데이터
특징
대체로 파이썬 코드 관련 텍스트로 이루어져 있긴 하지만, 코드가 없는 개발 관련 질문, sql 명령어 또한 일부 포함됨
질문-답변 방식이 아닌, 배경 스토리 기반으로 특정 개념이나 기능을 설명하는 자연어 텍스트 위주로 구성
주석, 에러 메세지 등은 한국어로 번역하되, 함수명 등은 원문을 사용
번역
powercoding-nopython 폴더에 있는 데이터 (실제로는 python 코드도 포함됨) 중 약 1%, 약 10억 토큰 번역됨
번역 과정
번역 모델: Qwen/Qwen3-30B-A3B
임베딩 모델: BAAI/bge-m3
번역 실패 후처리
번역 결과 텍스트에 1. 한글 및 특수문자가 아니면서 2. 원문에 포함되지 않는 문자가 하나라도 있을 경우 번역 실패로… See the full description on the dataset page: https://huggingface.co/datasets/werty1248/Powercoding-Ko-Translated-1B.MULTI_VALUE_mnli_were_was
Dataset Card for "MULTI_VALUE_mnli_were_was"
More Information needed
example_dataset_5This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"robot_type": "so-100",
"codebase_version": "v3.0",
"total_episodes": 50,
"total_frames": 6857,
"total_tasks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4",
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/wertania/example_dataset_5.eval_example_dataset_5This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 1,
"total_frames": 426,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/wertania/eval_example_dataset_5.ffw_sg2_rev1_werafail28This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ffw_sg2_rev1",
"total_episodes": 1,
"total_frames": 600,
"total_tasks": 1,
"total_videos": 3,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mkkimuser/ffw_sg2_rev1_werafail28.ffw_sg2_rev1_werafail25This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ffw_sg2_rev1",
"total_episodes": 1,
"total_frames": 601,
"total_tasks": 1,
"total_videos": 3,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mkkimuser/ffw_sg2_rev1_werafail25.werewolf-bluffers
Ultimate Werewolf Bluffing Structured Datasets (SD1 & SD2)
Dataset Summary
This repository contains two structured datasets, Structured Dataset 1 (SD1) and Structured Dataset 2 (SD2), designed for the study of bluffing behavior in Ultimate Werewolf, a social deduction game. The datasets provide player-level behavioral representations extracted from game transcripts shared by slhleosun and bolinlai.
Each record represents a single player during a single game round… See the full description on the dataset page: https://huggingface.co/datasets/KSBCode/werewolf-bluffers.Korea-Related-Reddit-posts
Comments: werty1248/Korea-Related-Reddit-comments
This dataset may contain aggressive content.
Source
Subreddit comments/submissions 2005-06 to 2024-12
Korea/Korean related posts only
Subset
Targeted subreddits
r/hanguk
r/Korea*
r/Korean*
r/kpop*
Posts
Contains at least one Korean character in the selftext (body).
Comments
Only comments that reply to Korean posts as defined above.
Statistics
Total data: 3.28TB
Korea/Korean-related subreddit… See the full description on the dataset page: https://huggingface.co/datasets/werty1248/Korea-Related-Reddit-posts.record-xmas-1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 20,
"total_frames": 17851,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/wertania/record-xmas-1.finepdfs-korean데이터 확인을 편하게 하려고 HuggingFaceFW/finepdfs에서 한국어 subset만 따로 분리했습니다.
ffw_sg2_rev1_wera12This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ffw_sg2_rev1",
"total_episodes": 1,
"total_frames": 601,
"total_tasks": 1,
"total_videos": 3,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mkkimuser/ffw_sg2_rev1_wera12.ffw_sg2_rev1_werafail8This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ffw_sg2_rev1",
"total_episodes": 1,
"total_frames": 600,
"total_tasks": 1,
"total_videos": 3,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mkkimuser/ffw_sg2_rev1_werafail8.multilingual-instruct-balancedThis repository is a collection of English, Korean, Chinese, and Japanese datasets collected by the HuggingFace Hub and transformed into a unified format. It consists of either native or synthetic data.
Some data is not clearly copyrighted or only allows non-commercial use.
Preprocessing: I removed data with too few answer tokens or more than 8192 tokens, and removed synthetic data with repetitions.
Balancing: I randomly sampled a subset of the data with different weights for each language and… See the full description on the dataset page: https://huggingface.co/datasets/werty1248/multilingual-instruct-balanced.ffw_sg2_rev1_wera15This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ffw_sg2_rev1",
"total_episodes": 1,
"total_frames": 601,
"total_tasks": 1,
"total_videos": 3,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mkkimuser/ffw_sg2_rev1_wera15.ffw_sg2_rev1_wera17This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ffw_sg2_rev1",
"total_episodes": 1,
"total_frames": 601,
"total_tasks": 1,
"total_videos": 3,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mkkimuser/ffw_sg2_rev1_wera17.ffw_sg2_rev1_wera18This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ffw_sg2_rev1",
"total_episodes": 1,
"total_frames": 600,
"total_tasks": 1,
"total_videos": 3,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mkkimuser/ffw_sg2_rev1_wera18.ffw_sg2_rev1_wera10This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ffw_sg2_rev1",
"total_episodes": 1,
"total_frames": 600,
"total_tasks": 1,
"total_videos": 3,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mkkimuser/ffw_sg2_rev1_wera10.
