Team Ai
Datasetpublic

FonaTech/E3-miu-GNN

Neo Mixed-Granularity Atomistic Dataset Neo is the dataset collection for E3-miu-GNN, an E(3)-equivariant graph neural network developed by Fona Group. This repository hosts the training data. GitHub hosts the model, GUI, training and inference workflows, dataset preparation code, tests, API documentation, and phonon tools. Project status Datasets: canonical Tiny through Large and composite SE, Plus, and Max are materialized with the 2026-07-25 non-OMat24 stress… See the full description on the dataset page: https://huggingface.co/datasets/FonaTech/E3-miu-GNN.

sourceHugging Faceotherupdated 3mo agoView on Hugging Face
0likes271downloads
Dataset Card

Neo Mixed-Granularity Atomistic Dataset

Neo is the dataset collection for E3-miu-GNN, an E(3)-equivariant graph neural network developed by Fona Group. This repository hosts the training data. GitHub hosts the model, GUI, training and inference workflows, dataset preparation code, tests, API documentation, and phonon tools.

Project status

  • —Datasets: canonical Tiny through Large and composite SE, Plus, and Max are materialized with the 2026-07-25 non-OMat24 stress revision.
  • —Model code: available on GitHub under the repository software license.
  • —Large-scale pretraining: in progress.
  • —Pretrained checkpoints: planned for a later validated release.

What Neo contains

Neo combines broad atomistic supervision with sparse, mechanism-specific response data:

  1. 1.periodic energy, force, and tensile-positive Cauchy stress for the L1 conservative foundation;
  2. 2.charge, dipole, polarizability, Born effective charge, and clamped-ion piezoelectric response for L2;
  3. 3.spin, magnetic moment, effective spin field, and spin-conditioned total stress for L3; and
  4. 4.OMat24-derived energy, force, and stress foundations in SE, Plus, and Max.

The schema also defines an optional paired electronic Hamiltonian/eigenvector payload for WALoss. None of the current seven tiers contains that payload; it is a supported extension, not an advertised Neo label.

Missing targets remain explicitly masked. Neo does not replace an unavailable physical quantity with a model prediction or a synthetic zero.

Browse the data

The Dataset Viewer offers four table views from the configuration menu:

ViewWhat it shows
structure_examplesDeterministic real structures from Standard response data and the SE OMat24 foundation, with train/validation/test splits
tier_summarySize, format, component, stress, spin, and checksum information for the seven release-facing tiers: Tiny, Small, Standard, SE, Large, Plus, and Max
label_coveragePer-tier counts, fractions, units, and mask scope for every target
source_coveragePer-tier source families, foundation/response role, and structure counts

The structure rows are intentionally marked preview_only = true. They make geometry and representative labels easy to inspect in a browser; they are not a reduced substitute for the complete HDF5 training files. Every preview row retains its HDF5 path, file SHA-256, record index, source identity, and upstream row index when applicable.

python
from datasets import load_dataset

examples = load_dataset(
    "FonaTech/E3-miu-GNN", "structure_examples"
)
tiers = load_dataset(
    "FonaTech/E3-miu-GNN", "tier_summary", split="train"
)

See VIEWER_TABLES.md for the table fields and reproducible generation command.

Choose a file

TierStructuresSizeRecommended use
Tiny5,78020.01 MiBInstallation checks and short examples
Small16,70351.22 MiBDevelopment experiments
Standard46,414129.68 MBDefault mixed-granularity training
SE605,693741.20 MBExact 1/180 OMat24 foundation plus the complete Standard response
Large613,2671.31 GBTrajectory-rich response training
Plus25,819,27141.06 GBOMat24-quarter foundation plus the current embedded Large response
Max101,283,549138.04 GBFull deduplicated OMat24 foundation plus the current embedded Large response

Start with Standard for the GUI and ordinary model development. Tiny and Small are deterministic nested subsets of Standard:

text
Tiny subset Small subset Standard

SE, Plus, and Max are standalone HDF5 files. Their OMat24 selection and response payload are stored inside each file, so an external OMat24 directory is not required. SE combines 559,279 deterministically selected OMat24 structures (1/180 of Max) with complete Standard. Plus and Max embed the 2026-07-25 enriched Large revision. Their packed OMat24 selections were not changed.

TierHDF5 layoutEmbedded Neo responseOMat24 foundationTotal atoms
Tinycanonical5,7800394,755
Smallcanonical16,70301,069,318
Standardcanonical46,41402,316,736
SEcompositeStandard: 46,414559,27912,767,209
Largecanonical613,267017,760,024
PluscompositeLarge: 613,26725,206,004488,227,614
MaxcompositeLarge: 613,267100,670,2821,899,323,661

Stress and coupled response

The Tiny-Large revision recovers stress only from original first-principles outputs. MPtrj provides MP GGA/GGA+U trajectory stress, and JARVIS-DFPT provides matched stress, BEC, and electronic clamped-ion piezoelectric tensors. All accepted stress is converted to tensile-positive eV/angstrom^3 and is used only for fully three-dimensional periodic, nonsingular cells.

TierStressStress + BEC + piezoelectricStress + spinsPaired magnetoelastic
Tiny2,0241129740
Small8,2431124,2200
Standard22,87311212,0000
SE582,15211212,0000
Large505,84811272,9290
Plus25,711,85211272,9290
Max101,176,13011272,9290

The stress-plus-spin records supervise total stress in a declared magnetic state. They are not same-geometry target/reference spin pairs and therefore do not identify a differential magnetoelastic tensor by themselves. The released presets keep w_magnetoelastic = 0; paired supervision will be enabled only after a separately converged constrained-spin DFT campaign is available.

Use with E3-miu-GNN

bash
git clone https://github.com/FonaTech/E3-miu-GNN.git
cd E3-miu-GNN
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
python E3_miu_GNN.py gui

Select a downloaded Neo HDF5 file under Data & Checkpoints. The application detects available labels, disables incompatible controls, and supports Base, Response, Joint, and Full Chain curricula. Command-line training uses the same configuration system:

bash
python E3_miu_GNN.py train --config path/to/config.json

For SE, Plus, or Max, Full Chain performs:

text
Base (OMat24 L1) -> Response (Neo response labels) -> Joint (coupled fine-tuning)

Data format

Tiny through Large use e3mu-hdf5-v1: packed geometry under structures/, numerical targets under labels/, validity indicators under masks/, and provenance and fixed splits under metadata/.

SE, Plus, and Max use e3mu-composite-hdf5-v1. They retain the four canonical groups for an embedded complete Standard or Large response payload and add selection/, sources/omat24/packed/, and atomic_reference/ for the foundation rows, exact source-row provenance, curriculum selectors, and atomic-reference statistics. Stored scientific arrays remain float64 and are not quantized.

See DATA_SCHEMA.md for shapes, units, and loading examples. The Parquet tables used by the Dataset Viewer are a companion publication layer; HDF5 remains the complete, lossless training authority.

Optional WALoss payload

The current reader and schema can consume two paired structure-level arrays in one fixed, gauge-aligned orbital or Wannier subspace:

Optional labelShape per structureUnitRequirement
orbital_hamiltonian(K, K)eVReal symmetric in one declared raw basis
orbital_eigenvectors(K, K)dimensionlessOrthonormal columns that diagonalize the paired reference Hamiltonian

One compatible file declares wavefunction_dim = K; both masks are active for a row or both are inactive. Orbital order, gauge (including degenerate subspaces), energy zero, spin channel, k-point convention, and electronic-structure method must be consistent across a training domain.

Tiny, Small, Standard, SE, Large, Plus, and Max currently omit the optional wavefunction_dim attribute and contain neither optional array. They remain valid for their existing targets, but they cannot train WALoss. The model-side feature is documented in the project README.

Sources and attribution

Neo is a derived compilation. Its OMat24 portions are based on OMat24 (Open Materials 2024), documented by FAIR Chemistry and distributed through the ColabFit OMat24 subdatasets under CC BY 4.0. Cite the OMat24 authors and <https://doi.org/10.48550/arXiv.2410.12771>. Neo selects, deduplicates, assigns fixed splits, and repackages retained records; it is not an official Meta, FAIR Chemistry, or ColabFit release, and those parties do not endorse it.

Neo also incorporates separately licensed MPtrj, JARVIS-DFT, QM7-X, SO3LR, SCFNN, and DeepSPIN-derived records. Because the aggregate contains multiple licenses, the card uses license: other. Full citations, notices, processing details, and the open archive-level rights review for the MLFF_and_BEC-derived records are documented in LICENSES_AND_ATTRIBUTION.md and SOURCES_AND_PROCESSING.md.

Limitations and responsible use

Neo is intended for atomistic machine-learning research. It is not a certified experimental database or a substitute for independent electronic-structure calculations. Coverage is uneven across chemistry and target type, and source methods use different approximations and energy references. Dataset validation establishes schema integrity, provenance, tensor conventions, and leakage controls; it does not establish model accuracy or transferability.

Do not rely on Neo alone for safety-critical or high-consequence decisions. Validate units, references, uncertainty, convergence, and applicability for each downstream use. The dataset is provided as-is, without warranty, to the extent permitted by applicable law. Upstream licenses remain controlling for their components.

Documentation

  • —GitHub repository: model, GUI, training, inference, tests, and architecture documentation.
  • —DATA_SCHEMA.md: HDF5 layout, units, masks, and loaders.
  • —VIEWER_TABLES.md: browser tables, fields, and loading.
  • —SOURCES_AND_PROCESSING.md: source selection, conversion, grouping, and physical mapping.
  • —LICENSES_AND_ATTRIBUTION.md: upstream terms, citations, notices, and disclaimer.
  • —HUGGINGFACE_UPLOAD.md: maintainer release workflow.

Citation

Until a versioned paper or dataset DOI is issued, cite the GitHub repository and the exact Hugging Face revision used, together with every applicable upstream dataset citation. Project citation metadata is maintained in `CITATION.cff`.