Team Ai
Datasetpublic

colabfit/Open_Catalyst_2025_OC25_Train

Cite this dataset Sahoo, S. J., Maroschin, M., Levine, D. S., Ulissi, Z., Zitnick, C. L., Varley, J. B., Gauthier, J. A., Govindarajan, N., and Shuaibi, M. Open Catalyst 2025 OC25 Train. ColabFit, 2025. https://doi.org/None This dataset has been curated and formatted for the ColabFit Exchange This dataset is also available on the ColabFit Exchange: https://materials.colabfit.org/id/DS_fanupene3rn7_0 Visit the ColabFit Exchange to… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/Open_Catalyst_2025_OC25_Train.

sourceHugging Facecc-by-4.0updated 5mo agoView on Hugging Face
2likes301downloads
Dataset Card

<details><summary>Cite this dataset </summary>Sahoo, S. J., Maroschin, M., Levine, D. S., Ulissi, Z., Zitnick, C. L., Varley, J. B., Gauthier, J. A., Govindarajan, N., and Shuaibi, M. Open Catalyst 2025 OC25 Train. ColabFit, 2025. https://doi.org/None</details>

This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:

https://materials.colabfit.org/id/DSfanupene3rn70

Visit the ColabFit Exchange to search additional datasets by author, description, element content and more.

https://materials.colabfit.org <br><hr>

Dataset Name

Open Catalyst 2025 OC25 Train

Description

The training split of the Open Catalyst 2025 (OC25) dataset for solid-liquid interfaces. OC25 consists of single-point DFT calculations of catalyst/solvent/ion/adsorbate structures, covering 88 elements, 8 solvents (water, methanol, CCl4, DMSO, benzene, hexane, THF, diethyl ether), 9 ionic species (Cs+, OH-, Li+, SO4^2-, Ca^2+, [Me4N]+, HCO3-, H+, F-), and adsorbates from the OC20 set plus reactive intermediates. Surfaces are derived from 39,821 Materials Project bulk structures with miller indices <= 3. Structures are highly off-equilibrium, sampled from short ab initio molecular dynamics simulations (10-50 steps, 1000K, NVT) or short DFT relaxations (5 ionic steps). The training split contains ~7.4 million structures filtered to total force drift < 1 eV/Å. All DFT calculations used VASP 6.3.2 with the non-spin-polarized RPBE functional supplemented with D3 dispersion correction (zero damping), plane wave cutoff 400 eV, EDIFF=1e-4 eV, k-point reciprocal density of 40, and a dipole correction in the z-direction.

Dataset authors

Sushree Jagriti Sahoo, Mikael Maroschin, Daniel S. Levine, Zachary Ulissi, C. Lawrence Zitnick, Joel B Varley, Joseph A. Gauthier, Nitish Govindarajan, Muhammed Shuaibi

Publication

https://doi.org/10.48550/arXiv.2509.17862

Original data link

https://huggingface.co/facebook/OC25

License

CC-BY-4.0

Number of unique molecular configurations

7395509

Number of atoms

1068208517

Elements included

Ag, Al, Ar, As, Au, B, Ba, Be, Bi, Br, C, Ca, Cd, Ce, Cl, Co, Cr, Cs, Cu, F, Fe, Ga, Ge, H, He, Hf, Hg, I, In, Ir, K, Kr, La, Li, Mg, Mn, Mo, N, Na, Nb, Nd, Ne, Ni, O, Os, P, Pb, Pd, Pm, Pr, Pt, Rb, Re, Rh, Ru, S, Sb, Sc, Se, Si, Sn, Sr, Ta, Tc, Te, Ti, Tl, V, W, Xe, Y, Zn, Zr

Properties included

energy, atomic forces <br> <hr>

Usage

  • —ds.parquet : Aggregated dataset information.
  • —co/ directory: Configuration rows each include a structure, calculated properties, and metadata.
  • —cs/ directory : Configuration sets are subsets of configurations grouped by some common characteristic. If cs/ does not exist, no configurations sets have been defined for this dataset.
  • —cs_co_map/ directory : The mapping of configurations to configuration sets (if defined). <br>
ColabFit Exchange documentation includes descriptions of content and example code for parsing parquet files: