FonaTech/E3-miu-GNN
Neo Mixed-Granularity Atomistic Dataset Neo is the dataset collection for E3-miu-GNN, an E(3)-equivariant graph neural network developed by Fona Group. This repository hosts the training data. GitHub hosts the model, GUI, training and inference workflows, dataset preparation code, tests, API documentation, and phonon tools. Project status Datasets: canonical Tiny through Large and composite SE, Plus, and Max are materialized with the 2026-07-25 non-OMat24 stress… See the full description on the dataset page: https://huggingface.co/datasets/FonaTech/E3-miu-GNN.
Neo Mixed-Granularity Atomistic Dataset
Neo is the dataset collection for E3-miu-GNN, an E(3)-equivariant graph neural network developed by Fona Group. This repository hosts the training data. GitHub hosts the model, GUI, training and inference workflows, dataset preparation code, tests, API documentation, and phonon tools.
Project status
- Datasets: canonical Tiny through Large and composite SE, Plus, and Max are materialized with the 2026-07-25 non-OMat24 stress revision.
- Model code: available on GitHub under the repository software license.
- Large-scale pretraining: in progress.
- Pretrained checkpoints: planned for a later validated release.
What Neo contains
Neo combines broad atomistic supervision with sparse, mechanism-specific response data:
- periodic energy, force, and tensile-positive Cauchy stress for the L1 conservative foundation;
- charge, dipole, polarizability, Born effective charge, and clamped-ion piezoelectric response for L2;
- spin, magnetic moment, effective spin field, and spin-conditioned total stress for L3; and
- OMat24-derived energy, force, and stress foundations in SE, Plus, and Max.
The schema also defines an optional paired electronic Hamiltonian/eigenvector payload for WALoss. None of the current seven tiers contains that payload; it is a supported extension, not an advertised Neo label.
Missing targets remain explicitly masked. Neo does not replace an unavailable physical quantity with a model prediction or a synthetic zero.
Browse the data
The Dataset Viewer offers four table views from the configuration menu:
The structure rows are intentionally marked preview_only = true. They make geometry and representative labels easy to inspect in a browser; they are not a reduced substitute for the complete HDF5 training files. Every preview row retains its HDF5 path, file SHA-256, record index, source identity, and upstream row index when applicable.
from datasets import load_dataset
examples = load_dataset(
"FonaTech/E3-miu-GNN", "structure_examples"
)
tiers = load_dataset(
"FonaTech/E3-miu-GNN", "tier_summary", split="train"
)See VIEWER_TABLES.md for the table fields and reproducible generation command.
Choose a file
Start with Standard for the GUI and ordinary model development. Tiny and Small are deterministic nested subsets of Standard:
Tiny subset Small subset StandardSE, Plus, and Max are standalone HDF5 files. Their OMat24 selection and response payload are stored inside each file, so an external OMat24 directory is not required. SE combines 559,279 deterministically selected OMat24 structures (1/180 of Max) with complete Standard. Plus and Max embed the 2026-07-25 enriched Large revision. Their packed OMat24 selections were not changed.
Stress and coupled response
The Tiny-Large revision recovers stress only from original first-principles outputs. MPtrj provides MP GGA/GGA+U trajectory stress, and JARVIS-DFPT provides matched stress, BEC, and electronic clamped-ion piezoelectric tensors. All accepted stress is converted to tensile-positive eV/angstrom^3 and is used only for fully three-dimensional periodic, nonsingular cells.
The stress-plus-spin records supervise total stress in a declared magnetic state. They are not same-geometry target/reference spin pairs and therefore do not identify a differential magnetoelastic tensor by themselves. The released presets keep w_magnetoelastic = 0; paired supervision will be enabled only after a separately converged constrained-spin DFT campaign is available.
Use with E3-miu-GNN
git clone https://github.com/FonaTech/E3-miu-GNN.git
cd E3-miu-GNN
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
python E3_miu_GNN.py guiSelect a downloaded Neo HDF5 file under Data & Checkpoints. The application detects available labels, disables incompatible controls, and supports Base, Response, Joint, and Full Chain curricula. Command-line training uses the same configuration system:
python E3_miu_GNN.py train --config path/to/config.jsonFor SE, Plus, or Max, Full Chain performs:
Base (OMat24 L1) -> Response (Neo response labels) -> Joint (coupled fine-tuning)Data format
Tiny through Large use e3mu-hdf5-v1: packed geometry under structures/, numerical targets under labels/, validity indicators under masks/, and provenance and fixed splits under metadata/.
SE, Plus, and Max use e3mu-composite-hdf5-v1. They retain the four canonical groups for an embedded complete Standard or Large response payload and add selection/, sources/omat24/packed/, and atomic_reference/ for the foundation rows, exact source-row provenance, curriculum selectors, and atomic-reference statistics. Stored scientific arrays remain float64 and are not quantized.
See DATA_SCHEMA.md for shapes, units, and loading examples. The Parquet tables used by the Dataset Viewer are a companion publication layer; HDF5 remains the complete, lossless training authority.
Optional WALoss payload
The current reader and schema can consume two paired structure-level arrays in one fixed, gauge-aligned orbital or Wannier subspace:
One compatible file declares wavefunction_dim = K; both masks are active for a row or both are inactive. Orbital order, gauge (including degenerate subspaces), energy zero, spin channel, k-point convention, and electronic-structure method must be consistent across a training domain.
Tiny, Small, Standard, SE, Large, Plus, and Max currently omit the optional wavefunction_dim attribute and contain neither optional array. They remain valid for their existing targets, but they cannot train WALoss. The model-side feature is documented in the project README.
Sources and attribution
Neo is a derived compilation. Its OMat24 portions are based on OMat24 (Open Materials 2024), documented by FAIR Chemistry and distributed through the ColabFit OMat24 subdatasets under CC BY 4.0. Cite the OMat24 authors and <https://doi.org/10.48550/arXiv.2410.12771>. Neo selects, deduplicates, assigns fixed splits, and repackages retained records; it is not an official Meta, FAIR Chemistry, or ColabFit release, and those parties do not endorse it.
Neo also incorporates separately licensed MPtrj, JARVIS-DFT, QM7-X, SO3LR, SCFNN, and DeepSPIN-derived records. Because the aggregate contains multiple licenses, the card uses license: other. Full citations, notices, processing details, and the open archive-level rights review for the MLFF_and_BEC-derived records are documented in LICENSES_AND_ATTRIBUTION.md and SOURCES_AND_PROCESSING.md.
Limitations and responsible use
Neo is intended for atomistic machine-learning research. It is not a certified experimental database or a substitute for independent electronic-structure calculations. Coverage is uneven across chemistry and target type, and source methods use different approximations and energy references. Dataset validation establishes schema integrity, provenance, tensor conventions, and leakage controls; it does not establish model accuracy or transferability.
Do not rely on Neo alone for safety-critical or high-consequence decisions. Validate units, references, uncertainty, convergence, and applicability for each downstream use. The dataset is provided as-is, without warranty, to the extent permitted by applicable law. Upstream licenses remain controlling for their components.
Documentation
- GitHub repository: model, GUI, training, inference, tests, and architecture documentation.
- DATA_SCHEMA.md: HDF5 layout, units, masks, and loaders.
- VIEWER_TABLES.md: browser tables, fields, and loading.
- SOURCES_AND_PROCESSING.md: source selection, conversion, grouping, and physical mapping.
- LICENSES_AND_ATTRIBUTION.md: upstream terms, citations, notices, and disclaimer.
- HUGGINGFACE_UPLOAD.md: maintainer release workflow.
Citation
Until a versioned paper or dataset DOI is issued, cite the GitHub repository and the exact Hugging Face revision used, together with every applicable upstream dataset citation. Project citation metadata is maintained in `CITATION.cff`.
