jepacpp/jepa.cpp-fixtures
jepa.cpp parity fixtures PyTorch golden reference dumps for jepa.cpp, a ggml-based C/C++ inference engine for the JEPA family. tests/test-parity and tests/test-predictor replay these tensors through the engine and gate per-token cosine, pooled outputs and classifier top-1/top-5 against per-family thresholds. Generated from jepa.cpp main @ 00bfd4e with scripts/dump_reference.py --model all, in float32 eval mode on 32 CPU threads, no autocast. Contents… See the full description on the dataset page: https://huggingface.co/datasets/jepacpp/jepa.cpp-fixtures.
jepa.cpp parity fixtures
PyTorch golden reference dumps for **jepa.cpp**, a ggml-based C/C++ inference engine for the JEPA family. tests/test-parity and tests/test-predictor replay these tensors through the engine and gate per-token cosine, pooled outputs and classifier top-1/top-5 against per-family thresholds.
Generated from jepa.cpp main @ `00bfd4e` with scripts/dump_reference.py --model all, in float32 eval mode on 32 CPU threads, no autocast.
Contents
222 files, 1176 MB in total.
Layout
One directory per model, each with a manifest.json and one .npy per tensor per sample, named <sample>.<tensor>.npy. All arrays are float32 C-order except frames_u8 (uint8) and top5_idx (int64). Shapes carry no batch dimension except input, which is stored exactly as fed to the model.
The manifest records the model id, the hyper-parameters, the preprocessing pipeline that produced input, the PyTorch forward wall time per sample (timing_s.forward_s, the baseline for the speed tables of the jepa.cpp docs), the frame indices sampled from each clip, and the label strings for the classifier. Per-model tensor lists and the exact preprocessing recipe: **docs/fixtures.md**.
How jepa.cpp consumes it
git clone --recursive https://github.com/aselimc/jepa.cpp && cd jepa.cpp
cmake -S . -B build -G Ninja -DCMAKE_BUILD_TYPE=Release && cmake --build build -j
scripts/download_fixtures.sh # this dataset -> tests/fixtures/ref, plus the media it needs
scripts/download_models.sh small # GGUFs from https://huggingface.co/jepacpp
cmake -S . -B build && ctest --test-dir build # the parity suites register at configure timetest-parity runs two passes per file: first the stored input tensor (bypassing preprocessing, so a graph bug shows up alone), then jepa.cpp's own preprocessor on the source media (so a preprocessing mismatch shows up separately). The second pass needs tests/fixtures/media/, which is not part of this dataset — see below.
Input provenance
The dumps are model outputs computed on two public research sets. Neither the source images nor the source videos are redistributed here; only the reference activations and the decoded frame tensors the tests replay are.
The input and frames_u8 arrays inside ref/ are preprocessed pixels of those images and clips, kept because the parity tests must feed the network exactly the tensor PyTorch saw. If you hold rights in any of the underlying material and want it removed, open an issue on github.com/aselimc/jepa.cpp.
Licence
cc-by-nc-4.0, the most restrictive licence among the checkpoints whose outputs are stored here: the dumps of I-JEPA ViT-H/14 (IN1k) and LeVJEPA ViT-L/16 (VideoMix) are outputs of CC BY-NC 4.0 models, the rest of MIT / Apache-2.0 ones. Research and non-commercial use.
Regenerating instead of downloading
scripts/download_fixtures.sh media # the source media
scripts/download_models.sh --convert all # the source checkpoints
.venv/bin/python scripts/dump_reference.py --model all # ~1 min on 32 coresLinks
- Code: <https://github.com/aselimc/jepa.cpp>
- Documentation: <https://aselimc.github.io/jepa.cpp/>
- GGUF models: <https://huggingface.co/jepacpp>
