SciLM/MaterialsSaddles
MaterialsSaddles A high-throughput library of converged transition states for solid-state and surface chemistry. Hub URL: https://huggingface.co/datasets/SciLM/MaterialsSaddles Released by SciLM.ai: https://www.scilm.ai 33,877,255 fully converged transition states computed by massively-parallel saddle searches on top of public materials and catalysis datasets, using the SaddleMill package and Meta's uma-s-1p2 machine-learning interatomic potential. Each entry in a file is a… See the full description on the dataset page: https://huggingface.co/datasets/SciLM/MaterialsSaddles.
MaterialsSaddles
A high-throughput library of converged transition states for solid-state and surface chemistry. Hub URL: <https://huggingface.co/datasets/SciLM/MaterialsSaddles> Released by SciLM.ai: <https://www.scilm.ai>
33,877,255 fully converged transition states computed by massively-parallel saddle searches on top of public materials and catalysis datasets, using the SaddleMill package and Meta's `uma-s-1p2` machine-learning interatomic potential.
Each entry in a file is a single structure. Three consecutive entries form one transition-state event: reactant minimum, transition state (first-order saddle), product minimum. Dimer endpoints are converged to 0.02 eV/Å (max\|F\|) and saddles to 0.05 eV/Å; for the NEB subset (mp20bat) see How the data was produced. Every row stores its uma-s-1p2 energy and forces, and each saddle row also stores its eigenmode — an (N, 3) per-atom displacement field giving the direction along which the saddle is unstable.
The files are already divided into train/, val/ and test/ directories, with no chemical system (element set) shared between the three splits. See Train / val / test.
Quick stats
Breakdown by source dataset
Try it: minimal example notebook
A self-contained Jupyter notebook (`example.ipynb`) demonstrates loading the dataset, converting ASE-LMDB rows to ASE Atoms objects (including a non-obvious atoms.info round-trip), walking the (R, S, P) triplet layout, visualizing a reaction, looking up a transition state by ms_id, and reproducing two small panels of Fig. 1 of the accompanying paper. It auto-installs its dependencies and downloads one test file from each subset.
To run locally:
hf download SciLM/MaterialsSaddles example.ipynb \
--repo-type dataset --local-dir .
jupyter notebook example.ipynbDirectory structure
.
├── README.md (this file)
├── DATASHEET.md (Datasheet for Datasets, Gebru et al. 2018)
├── example.ipynb
├── lemat/
│ ├── train/ lemat_train_0000.aselmdb … lemat_train_0564.aselmdb (565 files)
│ ├── val/ lemat_val_0000.aselmdb … lemat_val_0030.aselmdb (31 files)
│ └── test/ lemat_test_0000.aselmdb … lemat_test_0031.aselmdb (32 files)
├── oc20/
│ ├── train/ oc20_train_0000 … oc20_train_0042 (43 files)
│ ├── val/ oc20_val_0000 … oc20_val_0002 (3 files)
│ └── test/ oc20_test_0000 … oc20_test_0002 (3 files)
├── oc22/
│ ├── train/ oc22_train_0000 … oc22_train_0002 (3 files)
│ ├── val/ oc22_val_0000.aselmdb
│ └── test/ oc22_test_0000.aselmdb
├── mp20bat/
│ ├── train/ mp20bat_train_0000.aselmdb
│ ├── val/ mp20bat_val_0000.aselmdb
│ └── test/ mp20bat_test_0000.aselmdb
├── croissant.json (Croissant 1.1 metadata with Responsible-AI fields)
└── metadata/
└── triplets.parquet (one row per transition state: ms_id → file, row, split, chemistry, source)Every file holds at most 50,000 transition states (150,000 rows). In each split directory, every file but the last holds exactly 50,000. The exceptions are lemat/train/lemat_train_0213.aselmdb and lemat/train/lemat_train_0479.aselmdb, with 49,999 each.
Within each split directory, files and the rows inside them follow ascending ms_id.
What is in each row?
Each .aselmdb is an ASE-LMDB database whose rows are stored in triplets:
row id 1 reactant (Dimer: side = -1 ; NEB: image_type = 'endpoint')
row id 2 saddle (Dimer: side = 0 ; NEB: image_type = 'climbing') ← TS
row id 3 product (Dimer: side = 1 ; NEB: image_type = 'endpoint')
row id 4 reactant
...Every row carries the uma-s-1p2 single-point energy (row.energy, eV) and forces (row.forces, eV/Å), evaluated with the row's own UMA task (task_name, see below).
The rich per-row metadata lives in row.data['info']. After row.toatoms(), `atoms.info` is empty — you have to copy row.data['info'] over yourself (see Loading).
Each row also exposes a small set of searchable scalar key/value pairs via row.key_value_pairs: task_name, ms_id, src_index, status, and side on Dimer rows. These are convenient for db.select(...)-style filtering, but note that ASE's aselmdb backend performs linear scans: queries are O(N) per file, not indexed.
The keys in row.data['info'] vary by source dataset and saddle-search method. Only `task_name` and `ms_id` are guaranteed on every row — for anything else, the table below documents which subsets typically have it. When in doubt, inspect row.data['info'].keys() for a few rows of your target subset.
row.data may carry additional internal bookkeeping fields (e.g. traj_path) that are artifacts of the upstream production pipeline and have no scientific or downstream value. Read from row.data['info']; ignore anything else.
Where to find the source-dataset identifiers
The location depends on whether the input went through one or two stages of the SaddleMill pipeline before this release:
For Dimer subsets info['orig_info'] itself is a SaddleMill-internal dict (attempt_id, reaction_type, etc.) and the upstream identifiers live one level deeper. For the NEB subset there's only one level of nesting. The same identifiers are collected in the source_id column of `metadata/triplets.parquet`.
Intended uses
This dataset was built with two downstream uses in mind, both demonstrated in the companion paper:
- Training generative models for transition-state prediction, conditioned on the reactant alone or on reactant and product, with the saddle as the target. The eigenmode and bond-change annotations make it easy to filter for chemically meaningful events. The unconditional SaddleFlow model, trained on all four subsets, proposes saddles from which a Sella search needs 4.6–8.8× fewer force calls than from a perturbed reactant.
- Warm-starting DFT saddle searches. The ML saddles are close to their DFT counterparts: on 974 DFT-refined saddles across the four subsets they lie a median 0.038 Å (maximum per-atom displacement) from the DFT saddle, and a DFT dimer reconverges them in a median of 31–57 force calls per subset, 3.3–5.6× fewer than from the reactant–product midpoint. The DFT-refined saddles can in turn train MLIPs to predict barriers more accurately.
Loading the data
Requirements
pip install "ase>=3.26.0" ase_db_backendsase_db_backends registers the aselmdb backend so ase.db.connect(path, type="aselmdb") works directly. No fairchem-core install is required to read the data. Every file also opens with fairchem's fairchem.core.datasets.AseDBDataset.
Downloading one split
# the test split of oc22 (one file)
hf download SciLM/MaterialsSaddles --repo-type dataset --include "oc22/test/*" --local-dir MaterialsSaddles
# all training data of every subset (large: ~800 GB)
hf download SciLM/MaterialsSaddles --repo-type dataset --include "*/train/*" --local-dir MaterialsSaddles⚠ The atoms.info reconstruction trap
ASE's aselmdb backend does not round-trip atoms.info. Calling row.toatoms() returns an Atoms object whose .info is empty — the full original info dict (every metadata key documented above, including nested orig_info and the eigenmode ndarray) lives in row.data["info"]. Always use the canonical reader helper:
def row_to_atoms(row):
atoms = row.toatoms()
atoms.info.update(row.data["info"]) # restore the original info dict
return atomsMinimal example
from ase.db import connect
db = connect("lemat/train/lemat_train_0000.aselmdb", type="aselmdb")
for row in db.select(limit=6):
atoms = row_to_atoms(row)
print(atoms.get_chemical_formula(), "ms_id=", row.ms_id,
"side=", atoms.info.get("side"), "E=", row.energy)Walking the rows in triplets
from ase.db import connect
db = connect("lemat/train/lemat_train_0000.aselmdb", type="aselmdb")
print(len(db), "rows ->", len(db) // 3, "transition states")
for k in range(len(db) // 3):
reactant, saddle, product = (row_to_atoms(db.get(id=3 * k + i)) for i in (1, 2, 3))
print(saddle.get_chemical_formula(),
"eigenmode", saddle.info["eigenmode"].shape,
"curvature", saddle.info.get("curvature"),
"barrier", saddle.info.get("barrier"))The same loop works without modification on every file of every subset.
Looking up a transition state by ms_id
`metadata/triplets.parquet` gives the file and row of every transition state:
import pyarrow.parquet as pq
from ase.db import connect
rec = pq.read_table("metadata/triplets.parquet",
filters=[("saddle_ms_id", "==", 102302566)]).to_pylist()[0] # the saddle's ms_id
db = connect(rec["file"], type="aselmdb")
reactant, saddle, product = (row_to_atoms(db.get(id=rec["reactant_row_id"] + i)) for i in (0, 1, 2))Train / val / test
The split lives in the directory layout: <subset>/train/, <subset>/val/, <subset>/test/.
- Disjoint in chemical system. The chemical system of a transition state is the set of elements in its cell (for
oc20/oc22: slab and adsorbate). Each chemical system was assigned as a whole to train, val or test with probabilities 0.90 / 0.05 / 0.05 (NumPydefault_rng(0)), once, globally across all four subsets. No chemical system occurs in two splits — not within a subset and not across subsets (lematandmp20batshare materials chemistry). A test transition state therefore never has the same element set, and never the same composition, as any training transition state. - Every element is seen in training. Only element combinations are held out: in every subset, every element present in val or test also occurs in train.
- Triplet-level. The three rows of a transition-state event are always in the same split and the same file.
- Sizes are close to 90 / 5 / 5 (see the table in Quick stats); they deviate slightly because whole chemical systems are assigned together.
Metadata
metadata/triplets.parquet — one row per transition state in the release, sorted by saddle_ms_id:
How the data was produced
We took fully relaxed structures from four public datasets (LeMat-Bulk, OC20, OC22, and Materials Project battery structures) and ran high-throughput saddle searches against each one using the SaddleMill package, with Meta's `uma-s-1p2` universal interatomic potential (fairchem-core) as the calculator.
Initialization protocol (per-subset displacement modes such as vacancy, hop_insert, kickout_*, ring, adsorbate_atom, diffusion, rotation, …), eigenmode refinement, and post-search filtering are documented in the companion paper.
After saddle convergence, every dimer TS was validated by DoubleMinimization — displacing along the eigenmode in both directions and relaxing — and only triplets where the resulting endpoints actually correspond to two distinct basins (i.e. a real reaction occurred) are kept here. Anything that errored, hit a step limit, desorbed, or failed the reaction check is excluded. For the NEB subset (mp20bat), a band is kept when its climbing image converged and it has no intermediate minimum; its first, climbing, and last images form the triplet. In 189 kept bands the other images did not all converge, and 225 triplets have an endpoint whose pre-relaxation stopped above 0.02 eV/Å (image_converged = False). Energies and forces of every row were then computed with a uma-s-1p2 single point.
Known limitations
- MLIP, not DFT. All saddles, endpoints, energies and forces in this release are at the
uma-s-1p2MLIP level rather than DFT. On 974 DFT-refined saddles, geometries and per-atom energies are close to DFT (median 0.038 Å; 4 meV/atom MAE), but activation energies scatter (MAE 143 meV) with little systematic bias (median −14.5 meV), so treat the barriers as approximate. For DFT-level accuracy, run a single-point or short DFT saddle/NEB starting from these structures. - `atoms.info` is not auto-restored by `row.toatoms()`. See the trap callout above. Always use the
row_to_atomshelper. - `row.key_value_pairs` queries are linear scans. ASE's aselmdb backend has no secondary indices;
db.select(side=0)reads every row. Usemetadata/triplets.parquetto locate specific transition states. - Schema varies by source. Only
task_nameandms_idare guaranteed on every row. NEB-derived rows (e.g.mp20bat) have a differentinfoschema than dimer rows (e.g. noside, butimage_type/image_idx/barrierinstead).
Citation
If you use this dataset, please cite:
@misc{materialssaddles2026,
title = {{MaterialsSaddles}: 34 Million Transition States and a Flow-Matching
Saddle-Point Predictor for Materials},
author = {Baghishov, Ilgar and Jung, Sung Hoon and Henkelman, Graeme},
year = {2026},
note = {Accepted at NeurIPS 2026},
howpublished = {\url{https://huggingface.co/datasets/SciLM/MaterialsSaddles}}
}…and the upstream sources you actually used:
@article{chanussot2021oc20,
title = {Open Catalyst 2020 (OC20) Dataset and Community Challenges},
author = {Chanussot, Lowik and Das, Abhishek and Goyal, Siddharth and others},
journal = {ACS Catalysis},
volume = {11},
pages = {6059--6072},
year = {2021},
doi = {10.1021/acscatal.0c04525}
}
@article{tran2023oc22,
title = {The Open Catalyst 2022 (OC22) Dataset and Challenges for Oxide Electrocatalysts},
author = {Tran, Richard and Lan, Janice and Shuaibi, Muhammed and others},
journal = {ACS Catalysis},
volume = {13},
pages = {3066--3084},
year = {2023},
doi = {10.1021/acscatal.2c05426}
}
@article{jain2013mp,
title = {Commentary: The {Materials Project}: A materials genome approach to accelerating materials innovation},
author = {Jain, Anubhav and Ong, Shyue Ping and Hautier, Geoffroy and others},
journal = {APL Materials},
volume = {1},
number = {1},
pages = {011002},
year = {2013},
doi = {10.1063/1.4812323}
}
@misc{lemat-bulk,
title = {{LeMat-Bulk}: A unified, deduplicated dataset of bulk crystal structures},
author = {{Entalpic} and {Hugging Face}},
year = {2024},
note = {\url{https://huggingface.co/datasets/LeMaterial/LeMat-Bulk}}
}
@misc{uma2025,
title = {{UMA}: A Family of Universal Models for Atoms},
author = {{Meta FAIR Chemistry}},
year = {2025},
note = {\url{https://github.com/facebookresearch/fairchem} -- model {\tt uma-s-1p2}}
}
@article{ase,
title = {The atomic simulation environment---a {Python} library for working with atoms},
author = {Larsen, Ask Hjorth and Mortensen, Jens J{\o}rgen and Blomqvist, Jakob and others},
journal = {Journal of Physics: Condensed Matter},
volume = {29},
pages = {273002},
year = {2017},
doi = {10.1088/1361-648X/aa680e}
}Several of the entries above are placeholders or trimmed; please verify the canonical version against the publisher before submission.
License
This dataset is released under Creative Commons Attribution 4.0 International (CC-BY-4.0).
The upstream datasets retain their own licenses; consult them before any redistribution that combines this dataset with theirs.
Contact
Maintainers: Ilgar Baghishov (<baghishov@utexas.edu>) and Sung Hoon Jung (<sunghjung3@utexas.edu>). Website: <https://www.scilm.ai>. Issues / questions: open a discussion on the Hugging Face Hub page.
