Team Ai
Datasetpublic

SciLM/MaterialsSaddles

MaterialsSaddles A high-throughput library of converged transition states for solid-state and surface chemistry. Hub URL: https://huggingface.co/datasets/SciLM/MaterialsSaddles Released by SciLM.ai: https://www.scilm.ai 33,877,255 fully converged transition states computed by massively-parallel saddle searches on top of public materials and catalysis datasets, using the SaddleMill package and Meta's uma-s-1p2 machine-learning interatomic potential. Each entry in a file is a… See the full description on the dataset page: https://huggingface.co/datasets/SciLM/MaterialsSaddles.

sourceHugging Facecc-by-4.0updated 1d agoView on Hugging Face
0likes2.8kdownloads
Dataset Card

MaterialsSaddles

A high-throughput library of converged transition states for solid-state and surface chemistry. Hub URL: <https://huggingface.co/datasets/SciLM/MaterialsSaddles> Released by SciLM.ai: <https://www.scilm.ai>

33,877,255 fully converged transition states computed by massively-parallel saddle searches on top of public materials and catalysis datasets, using the SaddleMill package and Meta's `uma-s-1p2` machine-learning interatomic potential.

Each entry in a file is a single structure. Three consecutive entries form one transition-state event: reactant minimum, transition state (first-order saddle), product minimum. Dimer endpoints are converged to 0.02 eV/Å (max\|F\|) and saddles to 0.05 eV/Å; for the NEB subset (mp20bat) see How the data was produced. Every row stores its uma-s-1p2 energy and forces, and each saddle row also stores its eigenmode — an (N, 3) per-atom displacement field giving the direction along which the saddle is unstable.

The files are already divided into train/, val/ and test/ directories, with no chemical system (element set) shared between the three splits. See Train / val / test.


Quick stats

Transition states33,877,255
Files (.aselmdb)685 (at most 50,000 transition states each)
Rows per TS3 (reactant, saddle, product), in order
Splittrain / val / test ≈ 90 / 5 / 5, disjoint in chemical system
Saddle search methodDimer (lemat / oc20 / oc22), NEB-CI (mp20bat)
Calculatorfairchem uma-s-1p2
Endpoint convergence (max\F\)0.02 eV/Å (225 mp20bat triplets have an endpoint above it)
Saddle convergence (max\F\)0.05 eV/Å (mp20bat: the climbing image)
LicenseCC-BY-4.0

Breakdown by source dataset

SubsetSourceMethod# transition statestrainvaltestFiles
lemat/LeMat-BulkDimer31,322,91528,223,5141,533,3921,566,009628
oc20/Open Catalyst 2020 (OC20)Dimer2,367,7182,133,114123,751110,85349
oc22/Open Catalyst 2022 (OC22)Dimer152,593139,1756,5576,8615
mp20bat/Materials Project battery structuresNEB-CI34,02931,0441,4541,5313
Total33,877,25530,526,8471,665,1541,685,254685

Try it: minimal example notebook

A self-contained Jupyter notebook (`example.ipynb`) demonstrates loading the dataset, converting ASE-LMDB rows to ASE Atoms objects (including a non-obvious atoms.info round-trip), walking the (R, S, P) triplet layout, visualizing a reaction, looking up a transition state by ms_id, and reproducing two small panels of Fig. 1 of the accompanying paper. It auto-installs its dependencies and downloads one test file from each subset.

To run locally:

bash
hf download SciLM/MaterialsSaddles example.ipynb \
    --repo-type dataset --local-dir .
jupyter notebook example.ipynb

Directory structure

.
├── README.md                          (this file)
├── DATASHEET.md                       (Datasheet for Datasets, Gebru et al. 2018)
├── example.ipynb
├── lemat/
│   ├── train/  lemat_train_0000.aselmdb … lemat_train_0564.aselmdb   (565 files)
│   ├── val/    lemat_val_0000.aselmdb   … lemat_val_0030.aselmdb     (31 files)
│   └── test/   lemat_test_0000.aselmdb  … lemat_test_0031.aselmdb    (32 files)
├── oc20/
│   ├── train/  oc20_train_0000 … oc20_train_0042                     (43 files)
│   ├── val/    oc20_val_0000 … oc20_val_0002                         (3 files)
│   └── test/   oc20_test_0000 … oc20_test_0002                       (3 files)
├── oc22/
│   ├── train/  oc22_train_0000 … oc22_train_0002                     (3 files)
│   ├── val/    oc22_val_0000.aselmdb
│   └── test/   oc22_test_0000.aselmdb
├── mp20bat/
│   ├── train/  mp20bat_train_0000.aselmdb
│   ├── val/    mp20bat_val_0000.aselmdb
│   └── test/   mp20bat_test_0000.aselmdb
├── croissant.json                     (Croissant 1.1 metadata with Responsible-AI fields)
└── metadata/
    └── triplets.parquet               (one row per transition state: ms_id → file, row, split, chemistry, source)

Every file holds at most 50,000 transition states (150,000 rows). In each split directory, every file but the last holds exactly 50,000. The exceptions are lemat/train/lemat_train_0213.aselmdb and lemat/train/lemat_train_0479.aselmdb, with 49,999 each.

Within each split directory, files and the rows inside them follow ascending ms_id.


What is in each row?

Each .aselmdb is an ASE-LMDB database whose rows are stored in triplets:

row id 1  reactant  (Dimer:  side = -1   ;  NEB:  image_type = 'endpoint')
row id 2  saddle    (Dimer:  side =  0   ;  NEB:  image_type = 'climbing')   ← TS
row id 3  product   (Dimer:  side =  1   ;  NEB:  image_type = 'endpoint')
row id 4  reactant
...

Every row carries the uma-s-1p2 single-point energy (row.energy, eV) and forces (row.forces, eV/Å), evaluated with the row's own UMA task (task_name, see below).

The rich per-row metadata lives in row.data['info']. After row.toatoms(), `atoms.info` is empty — you have to copy row.data['info'] over yourself (see Loading).

Each row also exposes a small set of searchable scalar key/value pairs via row.key_value_pairs: task_name, ms_id, src_index, status, and side on Dimer rows. These are convenient for db.select(...)-style filtering, but note that ASE's aselmdb backend performs linear scans: queries are O(N) per file, not indexed.

The keys in row.data['info'] vary by source dataset and saddle-search method. Only `task_name` and `ms_id` are guaranteed on every row — for anything else, the table below documents which subsets typically have it. When in doubt, inspect row.data['info'].keys() for a few rows of your target subset.

key (in `info` dict)who has itwhat it is
sidedimer rows-1 / 0 / 1 (reactant / saddle / product)
image_typeNEB rows'endpoint' (reactant/product) or 'climbing' (saddle)
image_idx, subband_idx, nimagesNEB rowsNEB band geometry
image_converged, band_converged, band_converged_CINEB rowsper-image / band-level convergence flags
effective_fmaxNEB rowsper-image max-force (NEB-modified) at convergence
convergeddimer endpoint rowsendpoint converged flag
eigenmodesaddle rows(N, 3) float64 — the eigenmode at the saddle, one 3D displacement vector per atom
curvaturedimer saddleseigenvalue along that eigenmode (eV/Ų)
barrier, dENEB saddlesreactant→TS and reactant→product energy differences (eV)
is_reaction, n_formed_bonds, n_broken_bonds, formed_bonds, broken_bondsdimer rowsbond changes detected via ASE neighbor lists
is_ads_reaction, n_ads_*, ads_*_bondsdimer rowsadsorbate-restricted bond changes
parent_ts_indexdimer endpointsupstream SaddleMill identifier of the parent saddle (an internal pointer; the saddle in the same triplet is the immediately adjacent row in this .aselmdb file)
src_indexall rowsSaddleMill internal ID (file-local in the production run)
ms_idall rowsglobal row identifier across the entire dataset (0..102,406,790, with gaps). Triplets occupy three consecutive ms_ids: 3k (reactant), 3k+1 (saddle), 3k+2 (product).
task_nameall rowsUMA task head used during the saddle search and for the stored energy/forces: "omat" for lemat and mp20bat, "oc20" for oc20, "oc22" for oc22.
statusall rowsSaddleMill run status (always a converged* value here)
orig_infoall rowsnested dict; carries the source-dataset identifiers (see below)

row.data may carry additional internal bookkeeping fields (e.g. traj_path) that are artifacts of the upstream production pipeline and have no scientific or downstream value. Read from row.data['info']; ignore anything else.

Where to find the source-dataset identifiers

The location depends on whether the input went through one or two stages of the SaddleMill pipeline before this release:

SubsetPath to source IDs in the row's `info` dictExample fields
lematinfo['orig_info']['orig_info']immutable_id (e.g. 'agm005964602'), chemical_formula_*, functional, entalpic_fingerprint
oc20info['orig_info']['orig_info']source_file (e.g. 'random1176828.extxyz.xz')
oc22info['orig_info']['orig_info']sid, id, nads, natoms
mp20batinfo['orig_info']discharge_id / charge_id (Materials Project IDs, e.g. 'mp-1006112'), working_ion, removed_ion_idxs

For Dimer subsets info['orig_info'] itself is a SaddleMill-internal dict (attempt_id, reaction_type, etc.) and the upstream identifiers live one level deeper. For the NEB subset there's only one level of nesting. The same identifiers are collected in the source_id column of `metadata/triplets.parquet`.


Intended uses

This dataset was built with two downstream uses in mind, both demonstrated in the companion paper:

  1. 1.Training generative models for transition-state prediction, conditioned on the reactant alone or on reactant and product, with the saddle as the target. The eigenmode and bond-change annotations make it easy to filter for chemically meaningful events. The unconditional SaddleFlow model, trained on all four subsets, proposes saddles from which a Sella search needs 4.6–8.8× fewer force calls than from a perturbed reactant.
  2. 2.Warm-starting DFT saddle searches. The ML saddles are close to their DFT counterparts: on 974 DFT-refined saddles across the four subsets they lie a median 0.038 Å (maximum per-atom displacement) from the DFT saddle, and a DFT dimer reconverges them in a median of 31–57 force calls per subset, 3.3–5.6× fewer than from the reactant–product midpoint. The DFT-refined saddles can in turn train MLIPs to predict barriers more accurately.

Loading the data

Requirements

bash
pip install "ase>=3.26.0" ase_db_backends

ase_db_backends registers the aselmdb backend so ase.db.connect(path, type="aselmdb") works directly. No fairchem-core install is required to read the data. Every file also opens with fairchem's fairchem.core.datasets.AseDBDataset.

Downloading one split

bash
# the test split of oc22 (one file)
hf download SciLM/MaterialsSaddles --repo-type dataset --include "oc22/test/*" --local-dir MaterialsSaddles
# all training data of every subset (large: ~800 GB)
hf download SciLM/MaterialsSaddles --repo-type dataset --include "*/train/*" --local-dir MaterialsSaddles

⚠ The atoms.info reconstruction trap

ASE's aselmdb backend does not round-trip atoms.info. Calling row.toatoms() returns an Atoms object whose .info is empty — the full original info dict (every metadata key documented above, including nested orig_info and the eigenmode ndarray) lives in row.data["info"]. Always use the canonical reader helper:

python
def row_to_atoms(row):
    atoms = row.toatoms()
    atoms.info.update(row.data["info"])  # restore the original info dict
    return atoms

Minimal example

python
from ase.db import connect

db = connect("lemat/train/lemat_train_0000.aselmdb", type="aselmdb")
for row in db.select(limit=6):
    atoms = row_to_atoms(row)
    print(atoms.get_chemical_formula(), "ms_id=", row.ms_id,
          "side=", atoms.info.get("side"), "E=", row.energy)

Walking the rows in triplets

python
from ase.db import connect

db = connect("lemat/train/lemat_train_0000.aselmdb", type="aselmdb")
print(len(db), "rows ->", len(db) // 3, "transition states")

for k in range(len(db) // 3):
    reactant, saddle, product = (row_to_atoms(db.get(id=3 * k + i)) for i in (1, 2, 3))
    print(saddle.get_chemical_formula(),
          "eigenmode", saddle.info["eigenmode"].shape,
          "curvature", saddle.info.get("curvature"),
          "barrier",   saddle.info.get("barrier"))

The same loop works without modification on every file of every subset.

Looking up a transition state by ms_id

`metadata/triplets.parquet` gives the file and row of every transition state:

python
import pyarrow.parquet as pq
from ase.db import connect

rec = pq.read_table("metadata/triplets.parquet",
                    filters=[("saddle_ms_id", "==", 102302566)]).to_pylist()[0]   # the saddle's ms_id
db = connect(rec["file"], type="aselmdb")
reactant, saddle, product = (row_to_atoms(db.get(id=rec["reactant_row_id"] + i)) for i in (0, 1, 2))

Train / val / test

The split lives in the directory layout: <subset>/train/, <subset>/val/, <subset>/test/.

  • —Disjoint in chemical system. The chemical system of a transition state is the set of elements in its cell (for oc20/oc22: slab and adsorbate). Each chemical system was assigned as a whole to train, val or test with probabilities 0.90 / 0.05 / 0.05 (NumPy default_rng(0)), once, globally across all four subsets. No chemical system occurs in two splits — not within a subset and not across subsets (lemat and mp20bat share materials chemistry). A test transition state therefore never has the same element set, and never the same composition, as any training transition state.
  • —Every element is seen in training. Only element combinations are held out: in every subset, every element present in val or test also occurs in train.
  • —Triplet-level. The three rows of a transition-state event are always in the same split and the same file.
  • —Sizes are close to 90 / 5 / 5 (see the table in Quick stats); they deviate slightly because whole chemical systems are assigned together.

Metadata

metadata/triplets.parquet — one row per transition state in the release, sorted by saddle_ms_id:

columnmeaning
saddle_ms_idms_id of the saddle row (reactant = −1, product = +1)
subset, splite.g. lemat, test
filepath of the .aselmdb file in this repository
reactant_row_idASE row id of the reactant in file (saddle = +1, product = +2)
chemical_systemelement set, e.g. Li-O-P (the unit of the split)
reduced_formulacomposition divided by its greatest common divisor
source_idupstream identifier: LeMat-Bulk immutable_id; OC20 source_file; OC22 sid; Materials Project `charge_iddischarge_idworking_ion`

How the data was produced

We took fully relaxed structures from four public datasets (LeMat-Bulk, OC20, OC22, and Materials Project battery structures) and ran high-throughput saddle searches against each one using the SaddleMill package, with Meta's `uma-s-1p2` universal interatomic potential (fairchem-core) as the calculator.

SubsetMethodSaddleMill entrypoint
lematDimerSaddleMill.dimeropt
oc20DimerSaddleMill.dimeropt
oc22DimerSaddleMill.dimeropt
mp20batNEB-CISaddleMill.nebopt (climbing image)

Initialization protocol (per-subset displacement modes such as vacancy, hop_insert, kickout_*, ring, adsorbate_atom, diffusion, rotation, …), eigenmode refinement, and post-search filtering are documented in the companion paper.

After saddle convergence, every dimer TS was validated by DoubleMinimization — displacing along the eigenmode in both directions and relaxing — and only triplets where the resulting endpoints actually correspond to two distinct basins (i.e. a real reaction occurred) are kept here. Anything that errored, hit a step limit, desorbed, or failed the reaction check is excluded. For the NEB subset (mp20bat), a band is kept when its climbing image converged and it has no intermediate minimum; its first, climbing, and last images form the triplet. In 189 kept bands the other images did not all converge, and 225 triplets have an endpoint whose pre-relaxation stopped above 0.02 eV/Å (image_converged = False). Energies and forces of every row were then computed with a uma-s-1p2 single point.


Known limitations

  • —MLIP, not DFT. All saddles, endpoints, energies and forces in this release are at the uma-s-1p2 MLIP level rather than DFT. On 974 DFT-refined saddles, geometries and per-atom energies are close to DFT (median 0.038 Å; 4 meV/atom MAE), but activation energies scatter (MAE 143 meV) with little systematic bias (median −14.5 meV), so treat the barriers as approximate. For DFT-level accuracy, run a single-point or short DFT saddle/NEB starting from these structures.
  • —`atoms.info` is not auto-restored by `row.toatoms()`. See the trap callout above. Always use the row_to_atoms helper.
  • —`row.key_value_pairs` queries are linear scans. ASE's aselmdb backend has no secondary indices; db.select(side=0) reads every row. Use metadata/triplets.parquet to locate specific transition states.
  • —Schema varies by source. Only task_name and ms_id are guaranteed on every row. NEB-derived rows (e.g. mp20bat) have a different info schema than dimer rows (e.g. no side, but image_type / image_idx / barrier instead).

Citation

If you use this dataset, please cite:

bibtex
@misc{materialssaddles2026,
  title        = {{MaterialsSaddles}: 34 Million Transition States and a Flow-Matching
                  Saddle-Point Predictor for Materials},
  author       = {Baghishov, Ilgar and Jung, Sung Hoon and Henkelman, Graeme},
  year         = {2026},
  note         = {Accepted at NeurIPS 2026},
  howpublished = {\url{https://huggingface.co/datasets/SciLM/MaterialsSaddles}}
}

…and the upstream sources you actually used:

bibtex
@article{chanussot2021oc20,
  title   = {Open Catalyst 2020 (OC20) Dataset and Community Challenges},
  author  = {Chanussot, Lowik and Das, Abhishek and Goyal, Siddharth and others},
  journal = {ACS Catalysis},
  volume  = {11},
  pages   = {6059--6072},
  year    = {2021},
  doi     = {10.1021/acscatal.0c04525}
}

@article{tran2023oc22,
  title   = {The Open Catalyst 2022 (OC22) Dataset and Challenges for Oxide Electrocatalysts},
  author  = {Tran, Richard and Lan, Janice and Shuaibi, Muhammed and others},
  journal = {ACS Catalysis},
  volume  = {13},
  pages   = {3066--3084},
  year    = {2023},
  doi     = {10.1021/acscatal.2c05426}
}

@article{jain2013mp,
  title   = {Commentary: The {Materials Project}: A materials genome approach to accelerating materials innovation},
  author  = {Jain, Anubhav and Ong, Shyue Ping and Hautier, Geoffroy and others},
  journal = {APL Materials},
  volume  = {1},
  number  = {1},
  pages   = {011002},
  year    = {2013},
  doi     = {10.1063/1.4812323}
}

@misc{lemat-bulk,
  title  = {{LeMat-Bulk}: A unified, deduplicated dataset of bulk crystal structures},
  author = {{Entalpic} and {Hugging Face}},
  year   = {2024},
  note   = {\url{https://huggingface.co/datasets/LeMaterial/LeMat-Bulk}}
}

@misc{uma2025,
  title  = {{UMA}: A Family of Universal Models for Atoms},
  author = {{Meta FAIR Chemistry}},
  year   = {2025},
  note   = {\url{https://github.com/facebookresearch/fairchem} -- model {\tt uma-s-1p2}}
}

@article{ase,
  title   = {The atomic simulation environment---a {Python} library for working with atoms},
  author  = {Larsen, Ask Hjorth and Mortensen, Jens J{\o}rgen and Blomqvist, Jakob and others},
  journal = {Journal of Physics: Condensed Matter},
  volume  = {29},
  pages   = {273002},
  year    = {2017},
  doi     = {10.1088/1361-648X/aa680e}
}
Several of the entries above are placeholders or trimmed; please verify the canonical version against the publisher before submission.

License

This dataset is released under Creative Commons Attribution 4.0 International (CC-BY-4.0).

The upstream datasets retain their own licenses; consult them before any redistribution that combines this dataset with theirs.

Contact

Maintainers: Ilgar Baghishov (<baghishov@utexas.edu>) and Sung Hoon Jung (<sunghjung3@utexas.edu>). Website: <https://www.scilm.ai>. Issues / questions: open a discussion on the Hugging Face Hub page.