Team Ai
Datasetpublic

NoraResearchLab/Lithology-Training-Dataset

Lithology Training Dataset A supervised training dataset for machine learning and AI systems that learn to identify lithology from well-log data. The dataset contains 400 wells with standardized wireline-log measurements and corresponding lithological labels. Training dataset: https://huggingface.co/datasets/NoraResearchLab/Lithology-Training-Dataset Overview The core task is: Given a sequence of well-log measurements across depth, predict the lithology… See the full description on the dataset page: https://huggingface.co/datasets/NoraResearchLab/Lithology-Training-Dataset.

sourceHugging Faceotherupdated 8d agoView on Hugging Face
1likes159downloads
Dataset Card

![GitHub](https://github.com/Nora-Research-Lab) ![Hugging Face](https://huggingface.co/NoraResearchLab) ![LinkedIn](https://www.linkedin.com/company/nora-research-lab) ![X](https://x.com/noraresearchlab) ![Website](https://noraresearchlab.site)

Lithology Training Dataset

A supervised training dataset for machine learning and AI systems that learn to identify lithology from well-log data.

The dataset contains 400 wells with standardized wireline-log measurements and corresponding lithological labels.

Training dataset: https://huggingface.co/datasets/NoraResearchLab/Lithology-Training-Dataset


Overview

The core task is:

Given a sequence of well-log measurements across depth, predict the lithology present at each depth interval and/or reconstruct the corresponding lithological sequence.

The dataset is designed for:

  • —Lithology classification
  • —Well-log interpretation
  • —Sequence modelling
  • —Lithological interval segmentation
  • —Petrophysical machine learning
  • —Geoscience AI
  • —Subsurface representation learning
  • —Benchmark and model development

This dataset is intended for training and model development.

It is separate from the independent evaluation benchmark.


Dataset Size

PropertyValue
Total wells400
Lithological classes8
Primary data typeWireline well logs
Label typeRule-derived supervised labels
SamplingStandardized depth grid
Primary taskLithology identification

Lithology Classes

The dataset contains:

  • —SHALE
  • —SANDSTONE
  • —LIMESTONE
  • —DOLOMITE
  • —ANHYDRITE
  • —SALT
  • —COAL
  • —UNKNOWN/MIXED

UNKNOWN/MIXED represents intervals where the available log evidence does not provide sufficient confidence for assigning one of the principal lithological classes.


Dataset Structure

The repository contains multiple Parquet tables representing different levels of the dataset.

These tables do not share an identical schema and should therefore be loaded independently.

text
data/
├── logs/
│   ├── train.parquet
│   ├── validation.parquet
│   └── test.parquet
│
├── depthwise/
│   ├── train.parquet
│   ├── validation.parquet
│   └── test.parquet
│
└── intervals/
    ├── train.parquet
    ├── validation.parquet
    └── test.parquet

The exact files available in the repository should be treated as the authoritative source for the current release.


Dataset Configurations

logs

Depth-indexed wireline-log measurements.

Typical structure:

text
well_id
depth_ft
GR
RHOB
NPHI
PEF
DT
RT
CALI
...

This table is intended to provide the primary input features for machine-learning models.


depthwise

Depthwise supervised labels and derived information associated with individual log samples.

Typical structure may include:

text
well_id
depth_ft
lithology_pointwise
score
...

This configuration is useful for:

  • —Pointwise lithology classification
  • —Sequence modelling
  • —Depthwise prediction
  • —Label analysis
  • —Training supervised models

intervals

Lithological intervals represented as continuous depth ranges.

Typical structure:

text
well_id
interval_id
top_depth_ft
base_depth_ft
thickness_ft
lithology
...

This configuration is useful for:

  • —Lithological sequence reconstruction
  • —Interval segmentation
  • —Continuous geological interpretation
  • —Comparing predicted and reference intervals

Recommended Loading Method

Because the repository contains multiple heterogeneous Parquet tables, the recommended approach is to load each Parquet file directly.

Python

python
import pandas as pd

BASE_URL = (
    "https://huggingface.co/datasets/"
    "NoraResearchLab/Lithology-Training-Dataset/"
    "resolve/main/data"
)

# Logs
logs_train = pd.read_parquet(
    f"{BASE_URL}/logs/train.parquet"
)

logs_validation = pd.read_parquet(
    f"{BASE_URL}/logs/validation.parquet"
)

logs_test = pd.read_parquet(
    f"{BASE_URL}/logs/test.parquet"
)

# Depthwise labels
depthwise_train = pd.read_parquet(
    f"{BASE_URL}/depthwise/train.parquet"
)

depthwise_validation = pd.read_parquet(
    f"{BASE_URL}/depthwise/validation.parquet"
)

depthwise_test = pd.read_parquet(
    f"{BASE_URL}/depthwise/test.parquet"
)

# Lithological intervals
intervals_train = pd.read_parquet(
    f"{BASE_URL}/intervals/train.parquet"
)

intervals_validation = pd.read_parquet(
    f"{BASE_URL}/intervals/validation.parquet"
)

intervals_test = pd.read_parquet(
    f"{BASE_URL}/intervals/test.parquet"
)

Install the required packages with:

bash
pip install pandas pyarrow

This method reads the Parquet files directly from the Hugging Face repository without requiring the repository to be converted into a single Hugging Face Dataset object.


Why load_dataset() Is Not Recommended

The repository contains different table types with different schemas.

For example, logs and depthwise are primarily depthwise/pointwise data, while intervals contains range-based interval records.

Their structures are therefore different.

Attempting:

python
from datasets import load_dataset

dataset = load_dataset(
    "NoraResearchLab/Lithology-Training-Dataset"
)

may cause Hugging Face Datasets to attempt to construct a unified schema across the repository's Parquet files.

This can result in errors such as:

text
CastError:
Couldn't cast ...
because column names don't match

This is not an indication that the underlying Parquet files are invalid. The issue is that the repository contains multiple heterogeneous tabular structures.

Load each table independently instead.


Loading Only One Configuration

If you only need the well-log features for model training, you do not need to download or load the other tables.

For example:

python
import pandas as pd

url = (
    "https://huggingface.co/datasets/"
    "NoraResearchLab/Lithology-Training-Dataset/"
    "resolve/main/data/logs/train.parquet"
)

logs_train = pd.read_parquet(url)

print(logs_train.shape)
print(logs_train.head())

For the depthwise labels:

python
import pandas as pd

url = (
    "https://huggingface.co/datasets/"
    "NoraResearchLab/Lithology-Training-Dataset/"
    "resolve/main/data/depthwise/train.parquet"
)

depthwise_train = pd.read_parquet(url)

print(depthwise_train.shape)
print(depthwise_train.head())

For lithological intervals:

python
import pandas as pd

url = (
    "https://huggingface.co/datasets/"
    "NoraResearchLab/Lithology-Training-Dataset/"
    "resolve/main/data/intervals/train.parquet"
)

intervals_train = pd.read_parquet(url)

print(intervals_train.shape)
print(intervals_train.head())

Typical Training Workflow

A typical lithology-classification workflow can use the logs and depthwise tables together.

python
import pandas as pd

BASE_URL = (
    "https://huggingface.co/datasets/"
    "NoraResearchLab/Lithology-Training-Dataset/"
    "resolve/main/data"
)

logs = pd.read_parquet(
    f"{BASE_URL}/logs/train.parquet"
)

labels = pd.read_parquet(
    f"{BASE_URL}/depthwise/train.parquet"
)

The two tables can then be related using their common identifiers, typically:

text
well_id
depth_ft

For example:

python
training_data = logs.merge(
    labels,
    on=["well_id", "depth_ft"],
    how="inner"
)

Before training, users should inspect the resulting schema and verify that the expected feature and target columns are present:

python
print(training_data.shape)
print(training_data.columns.tolist())
print(training_data.head())

The exact columns should always be determined from the released Parquet files rather than assumed from this README.


Working With Individual Wells

The dataset is organized around complete wells. A model can therefore be trained using individual depth sequences rather than treating every row as an independent observation.

Example:

python
well_id = logs["well_id"].iloc[0]

well_logs = logs[
    logs["well_id"] == well_id
].sort_values("depth_ft")

print(well_logs)

This is particularly important for sequence-based approaches such as:

  • —LSTM/GRU models
  • —Temporal convolutional networks
  • —Transformers
  • —Sequence-to-sequence models
  • —Depthwise segmentation models

Train / Validation / Test Splits

The dataset provides separate splits where available:

text
train
validation
test

Users should preserve these splits when developing models.

Do not randomly mix rows from the same well across training and test sets unless the experimental design explicitly requires this.

For geological sequence modelling, well-level separation is important because neighboring depth samples from the same well are highly correlated.


Relationship to the Lithology Sequence Identification Benchmark

This training dataset is intended for model development and training.

It is separate from the:

Lithology Sequence Identification Benchmark

The benchmark is designed to evaluate whether a model can reconstruct the lithological sequence of a complete well from raw wireline measurements.

The intended workflow is:

text
Lithology Training Dataset
            │
            ▼
      Model Training
            │
            ▼
     Model Development
            │
            ▼
  Lithology Sequence Identification
        Benchmark
            │
            ▼
       Independent
        Evaluation

Models should not use benchmark evaluation wells as training data.


Data Loading Notes

Remote Parquet

pandas.read_parquet() can read the Parquet files directly from their Hugging Face HTTPS URLs when the appropriate Parquet engine is installed.

Recommended dependencies:

bash
pip install pandas pyarrow

Hugging Face Resolve URLs

The general URL pattern is:

text
https://huggingface.co/datasets/{USER}/{REPOSITORY}/resolve/{BRANCH}/{PATH}

For this dataset:

text
https://huggingface.co/datasets/NoraResearchLab/Lithology-Training-Dataset/resolve/main/data/

Individual files can then be addressed directly:

text
data/logs/train.parquet
data/depthwise/train.parquet
data/intervals/train.parquet

Important Usage Recommendation

Do not assume that every Parquet file in the repository has the same schema.

Treat each configuration as a separate table:

text
logs       → depthwise wireline measurements
depthwise  → pointwise labels / derived depthwise information
intervals  → continuous lithological intervals

Load only the tables required for the task you are performing.


License and Usage

This dataset is provided under the license specified in the repository metadata.

Users are responsible for reviewing the applicable dataset provenance, source-data licenses, and restrictions before redistribution or commercial use.

The dataset is intended for research, machine-learning development, and geoscience AI applications.


Citation

If you use this dataset in research, benchmarking, or a software project, please cite the NORA Research Lab dataset repository:

NORA Research Lab — Lithology Training Dataset

Hugging Face: https://huggingface.co/datasets/NoraResearchLab/Lithology-Training-Dataset


Organization

NORA Research Lab

Building intelligence for the real world.

  • —GitHub: https://github.com/Nora-Research-Lab
  • —Hugging Face: https://huggingface.co/NoraResearchLab
  • —LinkedIn: https://www.linkedin.com/company/nora-research-lab
  • —X: https://x.com/noraresearchlab
  • —Website: https://noraresearchlab.site