Team Ai
Datasetpublic

HTW-KI-Werkstatt/RamanBench

RamanBench Dataset Mirror ⚠️ RESEARCH MIRROR ONLY — All datasets are provided for research/educational purposes. Original copyrights remain with original authors. See Sources & Licenses below. A unified mirror of 87 Raman spectroscopy datasets from the RamanBench benchmark. Wide-format Parquet files for fast, reliable access. Quick Start from raman_bench import RamanBenchmark # Fast mirror access (default) bench = RamanBenchmark(… See the full description on the dataset page: https://huggingface.co/datasets/HTW-KI-Werkstatt/RamanBench.

sourceHugging Faceupdated 10d agoView on Hugging Face
1likes38kdownloads
Dataset Card

RamanBench Dataset Mirror

⚠️ RESEARCH MIRROR ONLY — All datasets are provided for research/educational purposes. Original copyrights remain with original authors. See Sources & Licenses below.


A unified mirror of 87 Raman spectroscopy datasets from the RamanBench benchmark. Wide-format Parquet files for fast, reliable access.

Quick Start

python
from raman_bench import RamanBenchmark

# Fast mirror access (default)
bench = RamanBenchmark(
    dataset_names_classification=["wheat_lines"],
    dataset_names_regression=["amino_acids_glycine"],
    use_mirror=True  # Uses this HF mirror
)

Or load directly:

python
from datasets import load_dataset
ds = load_dataset("HTW-KI-Werkstatt/RamanBench", "wheat_lines", split="train")
df = ds.to_pandas()

Citation & Attribution

If using this mirror, cite RamanBench:

bibtex
@article{koddenbrock2026ramanbench,
  title={RamanBench: A Large-Scale Benchmark for Machine Learning on Raman Spectroscopy},
  author={Koddenbrock, Mario and Lange, Christoph and Legner, Robin and J{\"a}ger, Martin and K{\"o}gler, Martin and Bournazou, Mariano N Cruz and Neubauer, Peter and Biessmann, Felix and Rodner, Erik},
  journal={arXiv preprint arXiv:2605.02003},
  year={2026}
}

CRITICAL: You MUST ALSO cite the original dataset sources. Failure to do so violates IP rights.

Dataset Structure

  • —Parquet files: Wide-format spectral data
  • —Columns: wavenumber floats (e.g., "200.5") + targets (single or multi-target)
  • —Data types: float32 for spectra, float64/string for targets
  • —Metadata: metadata.json with target names and task type

Available Datasets

87 datasets: 21 classification + 66 regression

Classification: wheat_lines, bacteria_identification, cancer_cell_*, hair_dyes_sers, ...

Regression: amino_acids_*, bioprocess_analytes_*, ecoli_fermentation, ...

See RamanBench paper for full list.

Sources & Licenses

SourceLicenseCopyright Holder
KaggleDataset TermsAuthors/Kaggle
ZenodoCC-BY, CC-0, customDepositors
FigshareCC-BYDepositors
GitHub (HTW-KI-Werkstatt)CC-BY or dataset-specificContributors
HuggingFace HubVariesContributors
RWTH Aachen / CustomInstitution-dependentResearch groups

Before using: Check original source for current license and terms.

Legal Notice

USAGE RIGHTS:

  • —✅ Research and non-commercial use
  • —✅ Must cite original authors AND RamanBench
  • —✅ Must respect original dataset licenses
  • —❌ No redistribution without original author permission
  • —❌ No commercial use unless licenses permit

INTELLECTUAL PROPERTY:

  • —Original authors/institutions retain all copyrights
  • —This mirror is for research convenience only
  • —You do NOT own or control these datasets

LIABILITY: By using this mirror, you acknowledge responsibility for respecting IP rights and following original dataset licenses.

Resources

  • —RamanBench GitHub: https://github.com/ml-lab-htw/RamanBench
  • —raman_data Package: https://github.com/ml-lab-htw/raman_data
  • —Paper: https://arxiv.org/abs/2605.02003
  • —Leaderboard: https://huggingface.co/spaces/HTW-KI-Werkstatt/RamanBench

Last updated: 2026-06-01 | License: CC-BY-4.0