datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
whisper_transcriptions.reazon_speech_all.wer_10.0.vectorizedRamanBench
RamanBench Dataset Mirror
⚠️ RESEARCH MIRROR ONLY — All datasets are provided for research/educational purposes. Original copyrights remain with original authors. See Sources & Licenses below.
A unified mirror of 87 Raman spectroscopy datasets from the RamanBench benchmark. Wide-format Parquet files for fast, reliable access.
Quick Start
from raman_bench import RamanBenchmark
# Fast mirror access (default)
bench = RamanBenchmark(… See the full description on the dataset page: https://huggingface.co/datasets/HTW-KI-Werkstatt/RamanBench.whisper_transcriptions.mls.wer_10.0whisper_transcriptions.mls.wer_10.0.vectorizedwhisper_transcriptions.reazonspeech.all.wer_10.0werea-tr-doc-ocr-enterprise-v2
Werea Turkish Enterprise Documents v2 📄🇹🇷
v1 setinin
enterprise sürümü: 12 belge türü × 3 çekim koşulu, 12.960 train + 900 test
sayfası. Werea-DocOCR v2 modellerinin eğitimi için üretilmiştir.
Belge türleri (12)
Genel vekaletname · DASK poliçesi · e-Arşiv fatura · Konut kira sözleşmesi ·
Banka dekontu · Tapu senedi · Maaş bordrosu · Kasko poliçesi · Araç tescil
bilgi formu · Resmî kurum yazısı · Ticaret sicil ilanı · SGK hizmet dökümü
Çekim… See the full description on the dataset page: https://huggingface.co/datasets/Werea-co/werea-tr-doc-ocr-enterprise-v2.werea-tr-doc-ocr-synthetic
Werea Turkish Enterprise Documents 📄🇹🇷
Türkçe kurumsal belge OCR eğitimi için tamamı sentetik sayfa görüntüleri ve
birebir eşleşen markdown ground-truth metinleri. Werea
tarafından Werea-DocOCR modellerinin eğitimi için üretilmiştir.
Belge türleri
Tür
Train
Test
İçerik
Genel vekaletname
500
50
Noter başlığı, taraflar, yetki maddeleri, noter şerhi
DASK poliçesi
500
50
Poliçe/sigortalı/bina bilgileri, prim tablosu
e-Arşiv fatura
500
50
Satıcı/alıcı… See the full description on the dataset page: https://huggingface.co/datasets/Werea-co/werea-tr-doc-ocr-synthetic.whisper_transcriptions.reazonspeech.all.wer_10.0.vectorizedwerewolftown_replayswhisper_transcriptions.reazon_speech_all.wer_10.0whisper_transcriptions.reazonspeech.large.wer_10.0werewolf-game-dataSynthMT
SynthMT: A Synthetic Benchmark for Automated Microtubule Segmentation
Authors: Mario Koddenbrock*, Justus Westerhoff*, Dominik Fachet, Simone Reber, Felix Gers, Erik Rodner
*Equal contribution
Affiliations: HTW Berlin, BHT Berlin, MPI for Infection Biology
Project Page: https://datexis.github.io/SynthMT-project-page/
Code & Pipeline: https://github.com/ml-lab-htw/SynthMTPaper: Synthetic Data Enables Human-Grade Microtubule Analysis with Foundation Models for SegmentationDataset:… See the full description on the dataset page: https://huggingface.co/datasets/HTW-KI-Werkstatt/SynthMT.werewolf_game_reasoning
Werewolf Game Dataset
This repository contains a comprehensive dataset for the Werewolf game in paper Multi-agent KTO: Reinforcing Strategic Interactions of Large Language Model in Language Game, including both raw game data and processed multi-level instruction datasets.
Dataset Structure
Raw Data
The raw data is located in the raw folder. Each game consists of two files:
event.json: Contains the game regular record and thinking process data, including:… See the full description on the dataset page: https://huggingface.co/datasets/ReneeYe/werewolf_game_reasoning.Werewolf-Among-Us
Werewolf Among Us: Multimodal Resources for Modeling Persuasion Behaviors in Social Deduction Games
ACL Findings 2023
Project Page | Paper | Code
Bolin Lai*, Hongxin Zhang*, Miao Liu*, Aryan Pariani*, Fiona Ryan, Wenqi Jia, Shirley Anugrah Hayati, James M. Rehg, Diyi Yang
Introduction
Werewolf Among Us is the first multi-modal (video+audio+text) dataset for social scenario interpretation. The dataset is composed of videos of 199 social deduction games and… See the full description on the dataset page: https://huggingface.co/datasets/bolinlai/Werewolf-Among-Us.werr_open_decisions
WERR Open Decisions Benchmark Dataset
Official benchmark dataset accompanying the research paper:"Universal Fractal Natural Language Decision Map: Real-Time Edge Triage Across Heterogeneous Domains"arXiv:2609.25498
📌 Dataset Summary
werr_open_decisions contains 1,205 verified decision trajectories and over 3,200 evaluation questions evaluated using the WERR (Waves & Errors) machine-native zero-VRAM reflex engine and the production answerr.me platform.
Unlike… See the full description on the dataset page: https://huggingface.co/datasets/pCwOrM/werr_open_decisions.example_dataset_3
example_dataset_3
This dataset was generated using phosphobot.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot.
To get started in robotics, get your own phospho starter pack..
arc_whisper_transcriptions.reazonspeech.small.wer_10.0.vectorized
Dataset Card for "arc_whisper_transcriptions.reazonspeech.small.wer_10.0.vectorized"
More Information needed
RamanSpectraEthanolicYeastFermentations
Dataset Card for Raman and NMR Spectra from Continuous Ethanolic Fermentation of Immobilized Yeast
Dataset Details
Dataset Description
This dataset contains Raman spectra acquired during the continuous ethanolic fermentation of sucrose using Saccharomyces cerevisiae (Baker's yeast). To facilitate continuous processing and high-quality optical measurements, the yeast cells were immobilized in calcium alginate beads.
The data covers process monitoring from two… See the full description on the dataset page: https://huggingface.co/datasets/HTW-KI-Werkstatt/RamanSpectraEthanolicYeastFermentations.werewolf-dataFuelRamanSpectraHandheld
Dataset Overview
This dataset contains Raman spectra for the analysis and prediction of key parameters in commercial fuel samples (gasoline). It includes spectra of 179 fuel samples from various refineries.
The dataset is designed for training models, such as Partial Least Squares (PLS) regression, to quickly and easily determine critical fuel characteristics like the Research Octane Number (RON) and the content of oxygenated additives, without the need for time-consuming standard… See the full description on the dataset page: https://huggingface.co/datasets/HTW-KI-Werkstatt/FuelRamanSpectraHandheld.werea-tts-tr-synthetic
Werea TTS TR Synthetic 🎙️🇹🇷
Werea-TSS modelinin eğitiminde
kullanılan, tamamı sentetik Türkçe metinden-sese veri seti. İzinsiz gerçek kişi
kaydı veya ses klonlama verisi içermez; tüm sesler
FreyaTTS-small öğretmen
modelinden sabit 20260814 tohumu ile üretilmiştir. Metinler Werea için özgün
olarak yazılmıştır.
İçerik
Split
Örnek
Süre
Açıklama
train
1.272
~1,32 saat
960 tam cümle (subset=stage1) + 312 kısa ifade (subset=short)
test
20
~76 sn
Eğitimle… See the full description on the dataset page: https://huggingface.co/datasets/Werea-co/werea-tts-tr-synthetic.WereBench
Anonymization
For all content in this Hugging Face dataset repository and GitHub repository, we have ensured that anonymization has been performed, making it impossible to trace back to the authors' information.
WereBench
WereBench is a benchmark dataset for evaluating language models in the Werewolf (similar to Mafia) social deduction setting. It focuses on human‑aligned strategic reasoning rather than only coarse metrics (e.g., win rate), aligning model behavior… See the full description on the dataset page: https://huggingface.co/datasets/n0nam4/WereBench.RamanSpectraEcoliMetabolites
Dataset Overview
This dataset contains Raman spectra of mixtures of glucose, sodium acetate which are important metabolites in the context of fermentations of Escherichia Coli.
Target Parameters and Concentration Ranges
The dataset contains measured Raman spectra of samples with different parameters from the following substances:
Glucose
Sodium Acetate
The concentrations were taken according to the volumes that the Tecan Liquid Handling Robot pipetted.
Data… See the full description on the dataset page: https://huggingface.co/datasets/HTW-KI-Werkstatt/RamanSpectraEcoliMetabolites.werewolf_gameplayswerewolf
🐺 Werewolf 3D - Asset Repository
Este repositorio contiene la biblioteca de recursos tridimensionales en formato GLTF/GLB de alto rendimiento, optimizados específicamente para entornos WebVR/WebXR y simulaciones interactivas en tiempo real.
📦 Contenido del Dataset
Archivo
Formato
Descripción
pueblo.glb
GLB (Draco)
Escenario principal del entorno urbano/rural con relieve topográfico completo.
Werewolf_Idle.glb
GLB (Draco + Animado)
Criatura en… See the full description on the dataset page: https://huggingface.co/datasets/Luis-Fernando/werewolf.RamanSpectraRalstoniaFermentations
Dataset Overview
This dataset contains Raman spectra (both real and synthetic) recorded during the batch cultivation of Ralstonia eutropha. The primary focus is the monitoring of the biodegradable copolymer poly(hydroxybutyrate-co-hydroxyhexanoate) [P(HB-co-HHx)].
Target Parameters and Composition
The dataset tracks the synthesis of the P(HB-co-HHx) copolymer and the metabolic state of the cultivation, focusing on the following key metrics:
Cell Dry Weight [g/L]: Total… See the full description on the dataset page: https://huggingface.co/datasets/HTW-KI-Werkstatt/RamanSpectraRalstoniaFermentations.Weronika_Buttendeich_WITAJ_teksty_wo_reci_a_pedagogiceWITAJ-Veröffentlichungen von Weronika Buttendeich in den Zeitschriften Serbska sula I Sorbische Schule und Lutki
RamanSpectraEcoliMetabolitesDig4Bio
Dataset Overview
This dataset contains Raman spectra of mixtures of glucose, sodium acetate, and magnesium sulfate. The spectra were measured with the system that is presended in the paper "A Setup for Automatic Raman Measurements in High-Throughput Experimentation" (https://doi.org/10.1002/bit.70006).
The spectra were used for a Kaggle challenge organized for the EU Project Dig4Bio (https://www.kaggle.com/competitions/dig-4-bio-raman-transfer-learning-challenge/overview/). We… See the full description on the dataset page: https://huggingface.co/datasets/HTW-KI-Werkstatt/RamanSpectraEcoliMetabolitesDig4Bio.math-game-assets
