HTW-KI-Werkstatt/RamanBench
RamanBench Dataset Mirror ⚠️ RESEARCH MIRROR ONLY — All datasets are provided for research/educational purposes. Original copyrights remain with original authors. See Sources & Licenses below. A unified mirror of 87 Raman spectroscopy datasets from the RamanBench benchmark. Wide-format Parquet files for fast, reliable access. Quick Start from raman_bench import RamanBenchmark # Fast mirror access (default) bench = RamanBenchmark(… See the full description on the dataset page: https://huggingface.co/datasets/HTW-KI-Werkstatt/RamanBench.
RamanBench Dataset Mirror
⚠️ RESEARCH MIRROR ONLY — All datasets are provided for research/educational purposes. Original copyrights remain with original authors. See Sources & Licenses below.
A unified mirror of 87 Raman spectroscopy datasets from the RamanBench benchmark. Wide-format Parquet files for fast, reliable access.
Quick Start
from raman_bench import RamanBenchmark
# Fast mirror access (default)
bench = RamanBenchmark(
dataset_names_classification=["wheat_lines"],
dataset_names_regression=["amino_acids_glycine"],
use_mirror=True # Uses this HF mirror
)Or load directly:
from datasets import load_dataset
ds = load_dataset("HTW-KI-Werkstatt/RamanBench", "wheat_lines", split="train")
df = ds.to_pandas()Citation & Attribution
If using this mirror, cite RamanBench:
@article{koddenbrock2026ramanbench,
title={RamanBench: A Large-Scale Benchmark for Machine Learning on Raman Spectroscopy},
author={Koddenbrock, Mario and Lange, Christoph and Legner, Robin and J{\"a}ger, Martin and K{\"o}gler, Martin and Bournazou, Mariano N Cruz and Neubauer, Peter and Biessmann, Felix and Rodner, Erik},
journal={arXiv preprint arXiv:2605.02003},
year={2026}
}CRITICAL: You MUST ALSO cite the original dataset sources. Failure to do so violates IP rights.
Dataset Structure
- Parquet files: Wide-format spectral data
- Columns: wavenumber floats (e.g.,
"200.5") + targets (single or multi-target) - Data types:
float32for spectra,float64/stringfor targets - Metadata:
metadata.jsonwith target names and task type
Available Datasets
87 datasets: 21 classification + 66 regression
Classification: wheat_lines, bacteria_identification, cancer_cell_*, hair_dyes_sers, ...
Regression: amino_acids_*, bioprocess_analytes_*, ecoli_fermentation, ...
See RamanBench paper for full list.
Sources & Licenses
Before using: Check original source for current license and terms.
Legal Notice
USAGE RIGHTS:
- ✅ Research and non-commercial use
- ✅ Must cite original authors AND RamanBench
- ✅ Must respect original dataset licenses
- ❌ No redistribution without original author permission
- ❌ No commercial use unless licenses permit
INTELLECTUAL PROPERTY:
- Original authors/institutions retain all copyrights
- This mirror is for research convenience only
- You do NOT own or control these datasets
LIABILITY: By using this mirror, you acknowledge responsibility for respecting IP rights and following original dataset licenses.
Resources
- RamanBench GitHub: https://github.com/ml-lab-htw/RamanBench
- raman_data Package: https://github.com/ml-lab-htw/raman_data
- Paper: https://arxiv.org/abs/2605.02003
- Leaderboard: https://huggingface.co/spaces/HTW-KI-Werkstatt/RamanBench
Last updated: 2026-06-01 | License: CC-BY-4.0
