Kokoslocke/NACA_4_Digit_for_ML
NACA 4-Digit Airfoil CFD Dataset Point-cloud CFD solutions for NACA 4-digit airfoils, generated with OpenFOAM v13 (k-ω SST). Intended for training surrogate models that predict steady-state flow fields from airfoil geometry and flow conditions. Dataset Summary ~850 converged in-distribution cases across 50 distinct NACA 4-digit profiles AoA range: −5° to +5° Reynolds number range: 100,000 – 500,000 129 out-of-distribution (OOD) probe cases at high Re (1–2 × 10⁶)… See the full description on the dataset page: https://huggingface.co/datasets/Kokoslocke/NACA_4_Digit_for_ML.
NACA 4-Digit Airfoil CFD Dataset
Point-cloud CFD solutions for NACA 4-digit airfoils, generated with OpenFOAM v13 (k-ω SST). Intended for training surrogate models that predict steady-state flow fields from airfoil geometry and flow conditions.
Dataset Summary
- ~850 converged in-distribution cases across 50 distinct NACA 4-digit profiles
- AoA range: −5° to +5°
- Reynolds number range: 100,000 – 500,000
- 129 out-of-distribution (OOD) probe cases at high Re (1–2 × 10⁶) and high AoA (10–15°) for generalization evaluation
- Split strategy: profile-level — all cases for a given airfoil shape belong to a single split, preventing geometry leakage between train/val/test
Airfoil Parameter Space
File Format
Each case is stored as a compressed NumPy archive (.npz). The bounding box is chord-normalized with the leading edge at x = 0 and the trailing edge at x = 1.
Spatial domain (bounding box):
x ∈ [−1.5, 3.5] (1.5c upstream, 2.5c downstream)
y ∈ [−1.5, 1.5] (1.5c above and below)Arrays per .npz file
N varies per case (typically 50,000 – 150,000 cells inside the bounding box).
Airfoil-surface arrays
In addition to the volume point cloud, each file carries a separate surface table with one row per airfoil wall face (M rows, M ≈ N_wall, typically a few hundred). These live on the airfoil surface itself, unlike is_wall which marks the first layer of volume cells sitting just off the wall. Rows are ordered by wall-face index and are mutually aligned.
Availability: the surface arrays are present for all splits —train,val,test, andood.
Quantities are kinematic (divided by density, consistent with p), matching OpenFOAM's wallShearStress function object. Multiply by ρ for physical units. Skin friction and pressure coefficients follow directly:
import numpy as np
data = np.load("NACA2412_p3.5_2.0e5.npz")
u_mag = float(np.hypot(data["u_init"][0], data["v_init"][0]))
q = 0.5 * u_mag**2 # kinematic dynamic pressure
cf = np.linalg.norm(data["wall_shear"], axis=1) / q
cp = data["wall_p"] / q
# Chordwise wall-shear sign flags separation (τ_w,x < 0 → reversed flow):
tau_x = data["wall_shear"][:, 0]Loading a sample
import numpy as np
data = np.load("NACA2412_p3.5_2.0e5.npz")
x, y = data["x"], data["y"]
u, v, p = data["u"], data["v"], data["p"]
re = float(data["reynolds"])File Naming Convention
NACA{code}_{sign}{aoa}_{Re}.npz{code}— 4-digit NACA identifier (e.g.,2412){sign}—pfor positive AoA,nfor negative AoA{aoa}— angle of attack in degrees (one decimal place){Re}— Reynolds number in scientific notation (e.g.,2.0e5)
Example: NACA2412_p3.5_2.0e5.npz → NACA 2412 airfoil, AoA = +3.5°, Re = 200,000.
Dataset Structure
NACA_4_digit_for_ml/
├── README.md
├── metadata.csv # Per-case tabular summary (see below)
└── *.npz # One file per converged casemetadata.csv columns
Dataset Splits
Splits are profile-level: every case sharing an airfoil code is assigned to one partition only. This ensures the model cannot memorize geometry.
The split field is also stored in each case's meta.yaml in the source repository. The OOD cases use airfoil shapes not present in train/val/test (95 unique new geometries, including thickness 6–26 % outside the 8–18 % training range) at Reynolds numbers and angles of attack well outside the training envelope (Re 1–2 × 10⁶, |AoA| 10–15°).
CFD Setup
Only cases that converged within the iteration budget are included. Non-converged cases are logged in the source repository's dataset/rejection_log.csv.
Source
Generated with cfd_data_generator. Mesh and field data produced by the pipeline in dataset/scripts/; ML-ready .npz files assembled by dataset/scripts/build_ml_dataset.py.
