RGES-PIT/MachineLearning
Machine Learning Tier This dataset is a collection of synthetic microlensing light curves from the Nancy Grace Roman Space Telescope Galactic Bulge Time Domain Survey. It is intended for the training and benchmarking of machine learning models for microlensing event classification, parameter estimation, and anomaly detection. The raw distribution of event properties is not representative of what Roman will see, but should span a statistically larger set of events. More… See the full description on the dataset page: https://huggingface.co/datasets/RGES-PIT/MachineLearning.
Machine Learning Tier
This dataset is a collection of synthetic microlensing light curves from the Nancy Grace Roman Space Telescope Galactic Bulge Time Domain Survey. It is intended for the training and benchmarking of machine learning models for microlensing event classification, parameter estimation, and anomaly detection.
The raw distribution of event properties is not representative of what Roman will see, but should span a statistically larger set of events. More representative samples can be drawn using the final_weight column, which is roughly proportional to the rate at which events occur relative to others in their class.
The machine learning tier data is provided as a set of normalized tables in Parquet format: RMDC26_ML_Data_obs.parquet, RMDC26_ML_Data_meta.parquet, and RMDC26_ML_Data_epoch.parquet.
Entity-Relationship Diagram
┌───────────────────┐ ┌───────────────────┐ ┌───────────────────┐
│ metadata │ │ observations │ │ epochs │
├───────────────────┤ ├───────────────────┤ ├───────────────────┤
│ event_id (PK) │◄───────├ event_id (FK) │ ┌──►│ epoch_id (PK) │
│ name │ │ epoch_id (FK) ────┼────┘ │ bjd │
│ sim_label │ │ filt │ │ obs_x │
│ tE_helio │ │ flux_uJy │ │ obs_y │
│ u0lens1 │ │ flux_err_uJy │ │ obs_z │
│ …(physical params)│ │ true_flux_uJy │ └───────────────────┘
└───────────────────┘ │ saturation_flag │
└───────────────────┘Relationships:
- observations → metadata: many-to-one on
event_id(each event has one metadata record) - observations → epochs: many-to-one on
epoch_id(each epoch has many observations) - To reconstruct a full lightcurve: join all three tables on
event_idandepoch_id
Tables
1. Observations Table (_obs.parquet)
Contains the primary time-series data for all events.
2. Metadata Table (_meta.parquet)
Contains the simulated physical parameters and lensing variables for each event.
3. Epochs Table (_epoch.parquet)
Maps each unique observation time (BJD) and corresponding telescope coordinates (X, Y, Z relative to the Solar System Barycenter) to a unique epoch ID.
Simulation Details & Disclaimer
Simulation Engine: All events have been generated using the gulls simulator. More information is available at gulls-microlensing.github.io.
Stellar Populations: The input stellar population catalogs used for the source and lens stars were generated using synthpop (github.com/synthpop-galaxy/synthpop).
Event Complexity:
- There are three simulations included in this dataset:
RMDC26_1S1L_ML: Single lens and single source events.RMDC26_1S2L_ML: One star - one planet lens and single source events.RMDC26_2S2L_ML: One star - one planet lens and potentially binary source events without source orbital motion.- Note that not all lenses and sources are necessarily involved in the microlensing event.
ObsGroup_0_chi2> 160 for events from the binary lens simulations indicates that an anomaly may be present in the light curve. The anomalies may be planetary perturbations and caustic crossings, finite source effects, or parallax effects.
Representativeness & Weighting:
- This data set only includes potentially detectable microlensing events where the chi-squared difference between the best fit flat model and the true model is greater than 60, i.e. $\chi^2{\text{flat}} - \chi^2{\text{true}} > 60$.
- The full set of events is not directly representative of the raw astrophysical population due to detection significance cuts and GULLS coordinate transformations (e.g. in
croincaustic-relative coordinates, the host star $t0$ and $u0$ are often transformed far outside of normal ranges or observing seasons, even though the planetary anomaly itself is actively observed in-season). - To make the dataset more representative of the true underlying Galactic population, you can apply the population weights stored in the `final_weight` column directly during training and analysis, or use them to draw a resampled, less biased dataset.
- The `is_out_of_season` flag indicates whether the closest approach time of the primary lens (
t0lens1) falls in a seasonal gap when the Galactic Bulge is unobservable. While the lightcurves of these events still show signature features of the microlensing event, their peak stellar magnification is not covered by observations. As a result, these events might be difficult for machine learning models to classify correctly, and you may want to filter them out depending on your specific modeling goals.
Dataset Structure:
- The dataset is split into observations, metadata, and epochs.
- To reconstruct a full lightcurve for an event, join the
obstable with themetatable onevent_idand theepochtable onepoch_id.
Disclaimer: This dataset is intended for testing and development purposes and should be used with the understanding of the biases inherent in the simulation sampling.
Sample Light Curves
Here are normalized light curves for 9 sample events from the dataset (3 from each simulation type: 1S1L, 1S2L, 2S2L), colored by filter band. Each filter's flux is divided by its median value (per event) so all bands appear on the same relative scale — the dashed gray line marks the baseline (ratio = 1). The time axis is zoomed to $t0 \pm 5 \cdot tE$.
Code Examples
The dataset is hosted on Hugging Face as RGES-PIT/MachineLearning with three configs: observations, metadata, and epochs. Below are examples for loading and working with the data.
Prerequisites
To load the metadata and epochs tables directly from Hugging Face, install the datasets and pandas libraries:
pip install datasets pandas pyarrow duckdb huggingface_hubSet your Hugging Face access token:
export HF_TOKEN="hf_your_token_here"1. Load the metadata and epochs tables
from datasets import load_dataset
import os
token = os.environ.get("HF_TOKEN")
assert token, "Set HF_TOKEN environment variable first"
dataset_name = "RGES-PIT/MachineLearning"
# Load metadata and epochs tables (which fit easily in RAM)
meta_df = load_dataset(dataset_name, "metadata", split="train", token=token).to_pandas()
epochs_df = load_dataset(dataset_name, "epochs", split="train", token=token).to_pandas()
print(f"Metadata : {meta_df.shape[0]:,} rows")
print(f"Epochs : {epochs_df.shape[0]:,} rows")Expected output:
Observations : 18,441,855,367 rows
Metadata : 372,655 rows
Epochs : 98,974 rows2. Reconstruct a single event's lightcurve
Because the observations table is extremely large (~160 GB in total), loading it in full is not recommended for local machines. Instead, we use DuckDB's remote hf:// protocol to fetch observations for a single event without downloading the entire table:
import duckdb
from huggingface_hub import HfFileSystem
import os
token = os.environ.get("HF_TOKEN")
con = duckdb.connect()
con.register_filesystem(HfFileSystem(token=token))
obs_p = "hf://datasets/RGES-PIT/MachineLearning/RMDC26_ML_Data_obs.parquet"
meta_p = "hf://datasets/RGES-PIT/MachineLearning/RMDC26_ML_Data_meta.parquet"
epoch_p = "hf://datasets/RGES-PIT/MachineLearning/RMDC26_ML_Data_epoch.parquet"
event_id = 3000
# Reconstruct a single event's lightcurve
event_obs = con.execute(f"""
SELECT
o.event_id,
m.name,
e.bjd,
o.filt,
o.flux_uJy,
m.tE_helio,
m.sim_label
FROM '{obs_p}' o
JOIN '{epoch_p}' e ON o.epoch_id = e.epoch_id
JOIN '{meta_p}' m ON o.event_id = m.event_id
WHERE o.event_id = {event_id}
ORDER BY e.bjd
""").df()
print(f"Event {event_id}: {event_obs['name'].iloc[0]} ({event_obs['sim_label'].iloc[0]})")
print(f" {len(event_obs):,} observations")
print(f" Einstein timescale tE_helio = {abs(event_obs['tE_helio'].iloc[0]):.2f} days")Expected output:
Event 3000: RMDC26_003000 (RMDC26_1S1L_ML)
49,488 observations
Einstein timescale tE_helio = 3.55 days3. Filter events by metadata constraints
You can perform filtering on the metadata table first, and then fetch observations for matched events. To run this efficiently over the network without full-table scans, query observations only for the target event IDs:
from datasets import load_dataset
import os
import duckdb
from huggingface_hub import HfFileSystem
token = os.environ.get("HF_TOKEN")
meta_df = load_dataset("RGES-PIT/MachineLearning", "metadata", split="train", token=token).to_pandas()
# Select low-significance single-lens events (chi2 <= 160)
good_events = meta_df[
(meta_df["sim_label"] == "RMDC26_1S1L_ML") &
(meta_df["ObsGroup_0_chi2"] <= 160)
]
print(f"{len(good_events)} events matched filters")
# Query observations count for matched events via DuckDB filter pushdown
con = duckdb.connect()
con.register_filesystem(HfFileSystem(token=token))
obs_p = "hf://datasets/RGES-PIT/MachineLearning/RMDC26_ML_Data_obs.parquet"
event_ids_str = ", ".join(map(str, good_events["event_id"].tolist()))
good_obs_count = con.execute(f"SELECT COUNT(*) FROM '{obs_p}' WHERE event_id IN ({event_ids_str})").fetchone()[0]
print(f"{good_obs_count:,} corresponding observations")Expected output:
59222 events matched filters
2,930,778,336 corresponding observations4. Sum final_weight per simulation
To verify the population representation or check the total astrophysical weights of each simulation, you can sum the final_weight column grouped by simulation label:
# Sum weights and count events per simulation
weights_summary = meta_df.groupby("sim_label").agg(
total_weight=("final_weight", "sum"),
event_count=("final_weight", "count")
).reset_index()
print(weights_summary.to_string(index=False))Expected output:
sim_label total_weight event_count
RMDC26_1S1L_ML 60411.874322 299466
RMDC26_1S2L_ML 1588.290136 31633
RMDC26_2S2L_ML 1940.221485 41556