datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sentinel-2-reanalysis-NDVI-EVI-TCG
Description
This dataset contains Sentinel-2 raw bands and reanalysis (NDVI, EVI, TCG) with valid pixel fractions per field.
Each parquet file contains multiple sunflower crop fields (identified as field_id).
Each field has 36 time stamps per year which is 10-day interval aggregation of the values of scaled and filtered bands and reanalysis.
Shape files of field polygons were loaded from Google Earth Engine (GEE) project assets.
Content
Country: Hungary
Fields: Sunflower… See the full description on the dataset page: https://huggingface.co/datasets/jaelin215/sentinel-2-reanalysis-NDVI-EVI-TCG.sentinel-2-reanalysis-NDVI-EVI-TCG
Description
This dataset contains Sentinel-2 raw bands and reanalysis (NDVI, EVI, TCG) with valid pixel fractions per field.
Each parquet file contains multiple sunflower crop fields (identified as field_id).
Each field has 36 time stamps per year which is 10-day interval aggregation of the values of scaled and filtered bands and reanalysis.
Shape files of field polygons were loaded from Google Earth Engine (GEE) project assets.
Content
Country: Hungary
Fields: Sunflower… See the full description on the dataset page: https://huggingface.co/datasets/Delineeeee/sentinel-2-reanalysis-NDVI-EVI-TCG.moltbook-ayush-reanalysis-20260505
MoltBook Ayush Reanalysis Bundle — 2026-05-05
This dataset is a self-contained analysis bundle for the MoltBook entropy-collapse reanalysis. It contains the raw experiment data used for the analysis, the exact included post index, cached Qwen embeddings, metadata-blind LLM-as-a-judge artifacts, deterministic/embedding/topical outputs, figures, reports, and scripts.
The purpose is to let a collaborator inspect or reproduce the analysis without needing the local workstation state.… See the full description on the dataset page: https://huggingface.co/datasets/Ayushnangia/moltbook-ayush-reanalysis-20260505.public_checkpoint_reanalysis
Public-checkpoint reanalysis: normalised run tables
Normalised training-loss observations from published language-model studies, used for the
public-checkpoint reanalysis accompanying Tokens-per-Parameter Coverage Is Critical for Robust
LLM Scaling Law Extrapolation.
Every number in that analysis is recomputable from these tables alone, with no need to re-fetch
the upstream sources.
Files
File
Rows
Contents
runs_expanded.csv
12,531
Full corpus: all ten… See the full description on the dataset page: https://huggingface.co/datasets/TPPIsCritical/public_checkpoint_reanalysis.
