Team Ai
Datasetpublic

NoraResearchLab/lithology-sequence-benchmark

Lithology Sequence Identification Benchmark Evaluation-only benchmark for reconstructing the lithological sequence of a complete well from raw wireline logs A fixed-well benchmark for testing whether machine-learning and AI systems can infer continuous lithological intervals from conventional petrophysical measurements. Overview The Lithology Sequence Identification Benchmark evaluates models on a practical subsurface interpretation problem: Given the… See the full description on the dataset page: https://huggingface.co/datasets/NoraResearchLab/lithology-sequence-benchmark.

sourceHugging Faceotherupdated 6d agoView on Hugging Face
1likes151downloads
Dataset Card

<div align="center">

Lithology Sequence Identification Benchmark

Evaluation-only benchmark for reconstructing the lithological sequence of a complete well from raw wireline logs

A fixed-well benchmark for testing whether machine-learning and AI systems can infer continuous lithological intervals from conventional petrophysical measurements.

![GitHub](https://github.com/Nora-Research-Lab) ![Hugging Face](https://huggingface.co/NoraResearchLab) ![LinkedIn](https://www.linkedin.com/company/nora-research-lab) ![X](https://x.com/noraresearchlab) ![Website](https://noraresearchlab.site)

</div>


Overview

The Lithology Sequence Identification Benchmark evaluates models on a practical subsurface interpretation problem:

Given the complete wireline-log suite from a well, can a model reconstruct the ordered sequence of lithological intervals across the logged depth?

Unlike conventional point-wise classification benchmarks, the task is not simply to assign a lithology label to every individual depth sample.

The expected prediction is a geologically meaningful sequence of contiguous depth intervals, for example:

json
[
  {
    "top_depth_ft": 2150.0,
    "base_depth_ft": 2191.5,
    "lithology": "SHALE"
  },
  {
    "top_depth_ft": 2191.5,
    "base_depth_ft": 2230.0,
    "lithology": "LIMESTONE"
  },
  {
    "top_depth_ft": 2230.0,
    "base_depth_ft": 2246.5,
    "lithology": "UNKNOWN/MIXED"
  }
]

Predicted intervals must be:

  • —ordered by depth;
  • —contiguous;
  • —non-overlapping;
  • —and collectively cover the logged interval.

The benchmark therefore evaluates both lithology identification and boundary placement / sequence reconstruction.


Benchmark Statistics

PropertyValue
Total wells245
Training wells166
Validation wells25
Test wells54
Depth samples1,963,283
Reference intervals49,830
Label-quality Tier A167 wells
Label-quality Tier B78 wells
Sampling interval0.5 ft
Lithology classes8
Evaluation typeWell-level / sequence prediction
Config fingerprint7b440d013c0a08a1
Reference fingerprint92e77d413b4c86d4

The train, validation, and test sets are split by well, rather than by individual depth samples.

This is important because adjacent samples from the same well are highly correlated. A random depth-level split could allow information from the same geological environment to appear in both training and test data and would therefore produce an overly optimistic estimate of generalisation.


Task Definition

Input

Each benchmark example consists of one complete well log sampled at approximately 0.5 ft.

The standard input curves are:

text
GR
RHOB
NPHI
PEF
DT
RT
CALI

Not every well contains every curve.

Missing curves are represented as:

text
NaN

A model should therefore be capable of operating with incomplete curve suites.


Output

The model must return an ordered sequence of lithological intervals:

text
top_depth_ft
base_depth_ft
lithology

The permitted lithology classes are:

text
SHALE
SANDSTONE
LIMESTONE
DOLOMITE
ANHYDRITE
SALT
COAL
UNKNOWN/MIXED

UNKNOWN/MIXED is an explicit class rather than an error state.

It represents depths where the available wireline evidence does not provide sufficient support for a reliable single-lithology interpretation.


Why Sequence Identification?

Traditional machine-learning approaches to lithology prediction often formulate the problem as:

text
wireline measurements at depth
        ↓
lithology class

This produces a label for every depth sample.

The present benchmark instead treats the well as a sequence reconstruction problem:

text
Complete well logs
        ↓
depth-wise geological evidence
        ↓
lithology transitions
        ↓
continuous geological intervals
        ↓
ordered lithological sequence

This distinction matters because a useful subsurface interpretation must identify not only what lithology occurs, but also:

  • —where an interval begins;
  • —where it ends;
  • —how thick it is;
  • —what lithology occurs above and below it;
  • —and whether the evidence is sufficiently strong to make a confident call.

A model that produces correct labels but places boundaries several feet away from the reference intervals can therefore receive different results from a model that accurately reconstructs both lithology and boundaries.


Lithology Classes

The benchmark contains eight possible output classes.

ClassDescription
SHALEFine-grained, generally clay-rich sedimentary interval
SANDSTONEPredominantly clastic sand-sized sedimentary interval
LIMESTONEPredominantly carbonate interval dominated by calcite
DOLOMITECarbonate interval with dolomitic characteristics
ANHYDRITEEvaporite interval dominated by anhydrite
SALTEvaporite interval dominated by halite
COALCoal-bearing interval
UNKNOWN/MIXEDInsufficient or conflicting evidence for a reliable single-class interpretation

The classes are intended for benchmark evaluation and should not be interpreted as exhaustive geological descriptions of every formation or facies present in the source wells.


Dataset Configurations

The dataset is organised into two principal configurations.

ConfigContentsPurpose
logsQC'd depth-indexed wireline measurementsModel input
wellsWell-level metadata and split informationMetadata / analysis

The reference lithology intervals are not publicly released.

They are retained privately by the benchmark maintainers for evaluation.


Files

logs

Contains the processed wireline-log observations.

Expected fields include:

text
well_id
depth_ft
GR
RHOB
NPHI
PEF
DT
RT
CALI

Missing measurements are represented using NaN.

The data are quality-controlled and standardised so that models can operate on a consistent representation despite differences between the original LAS files.


wells

Contains one record per well.

The metadata include information such as:

text
well_id
split
label_quality_tier
curves_present
baseline_information
reference_interval_count

This configuration is intended for dataset inspection, stratified analysis, and benchmark bookkeeping.


Private reference data

The maintainers retain the reference lithology representation used for scoring.

This includes the equivalent of:

text
intervals
depthwise labels

These files are intentionally withheld from the public release.

The purpose is to prevent direct optimisation against the evaluation labels.


Train / Validation / Test Split

The benchmark uses a well-level split:

text
Train:       166 wells
Validation:   25 wells
Test:         54 wells

A well and all of its depth samples belong to exactly one split.

This prevents the same well from contributing correlated observations to both training and evaluation.

The public benchmark therefore tests whether a model can generalise to previously unseen wells, rather than merely interpolate between nearby depth samples from wells it has already observed.


Reference Label Construction

The reference labels are generated through a deterministic petrophysical interpretation pipeline.

They are not produced using machine learning.

The pipeline is designed to convert heterogeneous raw LAS data into a consistent, reproducible lithological reference representation.

The process consists of:

text
Raw LAS files
      ↓
Curve identification
      ↓
Unit standardisation
      ↓
Quality control
      ↓
Depth-grid standardisation
      ↓
Multi-log lithology scoring
      ↓
UNKNOWN/MIXED confidence gate
      ↓
Sequence smoothing
      ↓
Minimum-bed-thickness processing
      ↓
Final lithological intervals

No randomness is used in the reference-label generation process.


Step 1 — Curve Detection

Different wells may use different mnemonics for equivalent measurements.

The benchmark therefore maps raw LAS curves to standard benchmark channels using a deterministic hierarchy:

text
mnemonic alias
      ↓
regular-expression matching
      ↓
description keywords
      ↓
unit information

For example, density curves can appear under names such as:

text
RHOB
RHOZ
DEN

and are mapped to:

text
RHOB

Likewise, resistivity-related curves may appear under names such as:

text
ILD
LLD
AT90

or conductivity representations and are standardised into the benchmark's RT representation where appropriate.

When multiple candidate curves satisfy the matching rules, ties are resolved using data coverage and then curve naming.

The complete mapping logic is defined by:

text
curve_specs.json

Step 2 — Unit Standardisation

The source LAS files contain measurements using different units.

The preprocessing system converts measurements to consistent benchmark units.

Examples include:

Original representationStandard representation
metresfeet
kg/m³g/cc
percentagefraction
µs/mµs/ft
millimetresinches
conductivityresistivity

The objective is to ensure that the downstream interpretation rules operate in a consistent numerical space.


Step 3 — Quality Control

Several quality-control operations are applied before lithology scoring.

These include:

Null and sentinel masking

Known null values and sentinel values are converted to missing observations.

Physical-range filtering

Measurements outside physically plausible ranges are masked.

Flat-line detection

Extended flat-line behaviour is detected as a possible dead-tool or failed-tool condition.

Such measurements are prevented from contributing misleading evidence.

Hampel spike filtering

Isolated extreme spikes are detected and removed using a Hampel-style robust outlier procedure.

Depth resampling

Logs are resampled onto a common fixed depth grid.

Fine-resolution curves use bin averaging.

Coarser measurements can use gap-limited interpolation where appropriate.

Coverage filtering

Curves with insufficient usable coverage are rejected for that well or interval.

Washout detection

Caliper measurements are used to identify potential borehole washout.

When washout is detected, density, neutron, PEF, and sonic evidence can be down-weighted because these measurements may become less representative of the formation.


Step 4 — Multi-Log Lithology Scoring

At every usable depth, each lithology receives a score based on the available petrophysical evidence.

The scoring system can incorporate:

  • —gamma-ray index;
  • —bulk density;
  • —neutron porosity;
  • —photoelectric factor;
  • —sonic travel time;
  • —logarithmic resistivity;
  • —neutron-density separation;
  • —and other configured evidence terms.

Each feature contributes through configurable membership functions.

The system uses fuzzy/trapezoidal membership functions rather than requiring a measurement to fall inside a single hard threshold.

This allows gradual transitions between lithological interpretations.

For example, a density measurement can provide partial support for multiple lithologies rather than producing an immediate binary decision.


Missing Curves

The benchmark explicitly supports incomplete log suites.

If a curve is absent:

text
NaN

is propagated through the scoring process.

Missing data do not automatically invalidate a sample.

Instead, the evidence contribution of the missing feature is removed from the total available evidence.

This means a well with:

text
GR
RHOB
NPHI
RT

can still receive an interpretation even if:

text
PEF
DT
CALI

are unavailable.

However, fewer available curves generally produce greater uncertainty.


Step 5 — UNKNOWN/MIXED Gate

The system does not force every depth into one of the seven specific lithology classes.

A depth can instead be assigned:

text
UNKNOWN/MIXED

when the available evidence is insufficient.

The gate considers factors such as:

  • —total available evidence weight;
  • —absolute score of the best lithology;
  • —difference between the best and second-best lithology;
  • —configured confidence thresholds.

Conceptually:

text
Strong evidence
      ↓
Specific lithology

Weak / contradictory evidence
      ↓
UNKNOWN/MIXED

This is intended to prevent the benchmark from treating every ambiguous petrophysical response as a confidently identified lithology.


Step 6 — Sequence Segmentation

Raw depth-wise classifications can contain short oscillations caused by measurement noise or borderline scores.

The reference-generation process therefore applies controlled smoothing and segmentation.

The procedure includes:

  1. 1.median/mean smoothing of depth-wise calls;
  2. 2.detection of short runs;
  3. 3.iterative absorption of intervals below the configured minimum bed thickness;
  4. 4.assignment of short intervals to the better-fitting neighbouring lithology;
  5. 5.merging of adjacent intervals with the same lithology.

The final output is therefore a sequence of continuous geological intervals rather than a noisy label at every individual sample.


Lithology Reference Signatures

The benchmark uses configurable default petrophysical signatures.

The principal plateau ranges are:

LithologyGRIRHOBNPHIPEFDTLOGRTND_SEP
SHALE0.65–2.52.2–2.650.25–0.452.5–3.885–140-0.2–0.70.05–0.22
SANDSTONE-1–0.252.15–2.65-0.04–0.181.6–2.352–900–2.2-0.12–-0.02
LIMESTONE-1–0.252.35–2.72-0.01–0.184.7–5.546–720.9–3-0.02–0.03
DOLOMITE-1–0.282.65–2.90.02–0.22.9–3.543–650.8–30.06–0.16
ANHYDRITE-1–0.122.92–3-0.04–0.034.8–5.448–532.6–50.12–0.22
SALT-1–0.122–2.15-0.04–0.044.4–4.965–702.6–5-0.45–-0.25
COAL-1–0.551.2–1.70.45–0.850.2–1.5110–1601.3–3.5-0.25–0.25

The complete membership functions, thresholds, weights, and other interpretation parameters are stored in:

text
label_config.json

These ranges should be understood as generic petrophysical signatures, not universal geological laws.


Evaluation

The public test logs do not include their reference lithology labels.

Researchers can therefore submit predictions generated from the held-out wells and evaluate them against the private reference set maintained by the benchmark authors.

A typical model interface is:

python
def my_model(well_curves_df):
    # Return:
    # top_depth_ft
    # base_depth_ft
    # lithology
    ...

The expected output is a DataFrame with:

text
top_depth_ft
base_depth_ft
lithology

For example:

python
prediction = pd.DataFrame([
    {
        "top_depth_ft": 2150.0,
        "base_depth_ft": 2191.5,
        "lithology": "SHALE"
    },
    {
        "top_depth_ft": 2191.5,
        "base_depth_ft": 2230.0,
        "lithology": "LIMESTONE"
    },
    {
        "top_depth_ft": 2230.0,
        "base_depth_ft": 2246.5,
        "lithology": "UNKNOWN/MIXED"
    }
])

Evaluation Metrics

The benchmark evaluates multiple aspects of model performance.

1. Depth-weighted / pooled lithology accuracy

Measures the proportion of evaluated depth samples for which the predicted lithology agrees with the reference.

This captures overall depth coverage but can be dominated by thick intervals.


2. Depth macro-F1

Macro-F1 gives greater importance to performance across individual lithology classes rather than allowing the most common lithologies to dominate the score.

This is particularly relevant when lithologies have highly unequal thickness distributions.


3. Confident-reference evaluation

Metrics can additionally be calculated only on reference depths where the underlying reference system has sufficient confidence.

This separates:

text
performance against all reference interpretations

from:

text
performance against high-confidence reference interpretations

Boundary F1

Lithology identification alone does not fully measure sequence reconstruction.

A model can correctly identify the lithologies but place boundaries incorrectly.

Boundary F1 therefore evaluates whether predicted lithological transitions occur close to reference boundaries.

The benchmark reports boundary F1 at:

text
±2.5 ft
±5 ft
±10 ft

A predicted boundary is considered a match when it falls within the corresponding tolerance of a reference boundary.

This provides progressively more permissive measures of geological boundary localisation.


Sequence Similarity

The benchmark also evaluates the similarity of the ordered lithological sequence.

The sequence metric is based on:

text
1 − normalised edit distance

between the predicted and reference lithology sequences.

For example:

text
Reference:
SHALE → SANDSTONE → LIMESTONE → DOLOMITE

Prediction:
SHALE → SANDSTONE → LIMESTONE → DOLOMITE

has identical sequence structure.

A prediction such as:

text
SHALE → LIMESTONE → DOLOMITE

has a different sequence even if the model correctly identifies several of the individual lithologies.

This metric therefore captures sequence-level agreement rather than simply point-wise classification.


Recommended Evaluation Table

Researchers should report performance separately for the complete test set and, where applicable, high-confidence reference depths.

A recommended format is:

ModelDepth Macro-F1AccuracyBoundary F1 ±2.5 ftBoundary F1 ±5 ftBoundary F1 ±10 ftSequence Similarity
Model A——————
Model B——————

Per-class metrics are also encouraged, particularly for:

text
SANDSTONE
LIMESTONE
DOLOMITE
ANHYDRITE
SALT
COAL

because these classes can be substantially less frequent than shale.


Baselines

The benchmark is designed to support comparison between several classes of approaches, including:

  • —rule-based petrophysical interpretation;
  • —classical machine-learning classifiers;
  • —gradient-boosted models;
  • —recurrent sequence models;
  • —temporal convolutional networks;
  • —Transformer-based sequence models;
  • —hybrid physics/ML approaches;
  • —foundation-model or multimodal approaches.

A useful baseline should operate only on information available to the model at inference time and should not use the withheld reference labels.


Evaluation Example

A simplified evaluation workflow is:

python
from datasets import load_dataset
import pandas as pd

logs = pd.concat([
    load_dataset(
        "NoraResearchLab/Lithology-Sequence-Benchmark",
        "logs",
        split=s
    ).to_pandas()
    for s in ["test"]
])

# Reference labels are privately maintained.
# Public users submit predictions for scoring.

# my_model(well_curves_df) ->
# DataFrame[
#   top_depth_ft,
#   base_depth_ft,
#   lithology
# ]

# summary, per_well = evaluate_model(
#     my_model,
#     logs,
#     reference
# )

The exact evaluation implementation and private reference data are maintained separately from the public test inputs.


Data Leakage Considerations

This benchmark is explicitly designed to prevent direct access to the evaluation labels.

The public release contains:

text
QC'd logs
+
well metadata

but does not contain:

text
test lithology intervals
test depth-wise labels

Researchers should not attempt to reconstruct the private reference labels from benchmark metadata or implementation details and should not use the private reference representation during model development.

Because the reference labels are generated by a deterministic rule-based system, a sufficiently detailed reproduction of the reference pipeline could potentially approximate the scoring target. Researchers using such an approach should clearly disclose that methodology.


Label Quality

The reference data contain two quality tiers:

text
Tier A: 167 wells
Tier B: 78 wells

The tiers reflect differences in the quality and completeness of the available wireline evidence.

Tier A wells generally provide stronger multi-log evidence.

Tier B wells may have more limited curve availability, increasing ambiguity in lithology discrimination.

This is particularly relevant for distinguishing:

text
SANDSTONE
LIMESTONE
DOLOMITE

when PEF or sonic information is unavailable.


Important Limitations

Silver-standard reference labels

The benchmark labels are rule-derived silver-standard labels.

They are not equivalent to independently verified geological ground truth from:

  • —core;
  • —cuttings;
  • —thin sections;
  • —petrographic analysis;
  • —formation-tester measurements;
  • —or a geologist's independent interpretation.

The benchmark therefore measures agreement with a transparent, deterministic petrophysical interpretation system.

A high score demonstrates that a model can reproduce the benchmark's reference interpretation.

It does not, by itself, establish that the model has achieved independently verified geological accuracy.


Generic petrophysical signatures

The default lithology signatures are based on generic clean-matrix / textbook-style petrophysical ranges.

They are not calibrated specifically to every formation represented in the source archive.

Formation-specific mineralogy, pore-fluid properties, compaction, diagenesis, borehole conditions, and logging-tool characteristics can shift observed responses.


Neutron scale assumptions

Neutron measurements are interpreted on a limestone-equivalent scale unless the source mnemonic indicates another interpretation.

This can introduce ambiguity when different logging conventions or environmental corrections are present.


Borehole effects

Poor hole conditions can affect:

  • —density;
  • —neutron;
  • —PEF;
  • —sonic;
  • —and other measurements.

Caliper-based washout detection reduces the contribution of affected curves, but it cannot completely eliminate borehole-related uncertainty.


Missing curves

Not every well contains the complete curve suite.

The benchmark therefore evaluates models under realistic missing-feature conditions.

A model that requires every possible curve will have limited applicability across the complete benchmark.


Synthetic / rule-derived interpretation

Although the source measurements are real wireline logs, the reference lithology sequence is produced algorithmically.

The benchmark should therefore be regarded as an evaluation of:

text
wireline logs → benchmark lithology sequence

rather than:

text
wireline logs → independently verified geological truth

Source Data

The underlying wireline logs originate from the public digital wireline-log archives of the:

Kansas Geological Survey (KGS), University of Kansas

Source:

https://www.kgs.ku.edu/Magellan/Logs/

The benchmark is an independent derivative work and is not endorsed by the Kansas Geological Survey or the University of Kansas.

Users should review the applicable KGS terms, disclaimers, and conditions before using the underlying data or derivative products in downstream commercial or research applications.


Reproducibility and Versioning

Several files are used to make the benchmark pipeline auditable and reproducible.

label_config.json

Contains the lithology membership functions, scoring weights, confidence thresholds, segmentation parameters, and other reference-generation settings.

curve_specs.json

Contains the curve mnemonic aliases, matching rules, unit rules, and standardisation logic.

manifest.json

Contains dataset version information, file manifests, and cryptographic hashes.

Configuration fingerprint

text
7b440d013c0a08a1

Reference fingerprint

text
92e77d413b4c86d4

These fingerprints allow benchmark users to identify the exact configuration and reference version associated with a reported result.


Reproducibility Principles

For meaningful comparisons, researchers should:

  1. 1.Use the fixed test wells.
  2. 2.Do not modify the public test inputs.
  3. 3.Do not train on the withheld reference labels.
  4. 4.Report the benchmark version/configuration.
  5. 5.Report preprocessing performed by the model.
  6. 6.Report how missing curves are handled.
  7. 7.Report whether depth-wise predictions are post-processed into intervals.
  8. 8.Report the exact model checkpoint used for evaluation.

Intended Uses

The benchmark is intended for:

  • —lithology classification research;
  • —automated well-log interpretation;
  • —sequence modelling;
  • —petrophysical machine learning;
  • —geological boundary detection;
  • —formation evaluation research;
  • —benchmarking missing-log robustness;
  • —comparing classical ML and deep-learning approaches;
  • —evaluating AI systems for subsurface interpretation.

It can also serve as a controlled test case for research into models that combine numerical sequence understanding with geological reasoning.


Not Intended For

The benchmark should not be used as the sole basis for:

  • —drilling decisions;
  • —reservoir development decisions;
  • —formation abandonment decisions;
  • —commercial reserve estimation;
  • —safety-critical geological interpretation;
  • —regulatory reporting;
  • —or independent confirmation of geological conditions.

Any operational subsurface application should use appropriately qualified geological and petrophysical review and independently validated reference data.


Recommended Research Questions

The benchmark can support research questions such as:

Can sequence models outperform independent depth-wise classifiers?

A model may exploit geological continuity and neighbouring depth information to produce more coherent intervals.

How much does missing-log robustness affect performance?

Researchers can compare models across wells with different available curve suites.

Can models identify thin beds without producing excessive segmentation noise?

Boundary F1 and sequence similarity provide complementary measures for this problem.

Can machine-learning models outperform generic petrophysical rules?

Because the reference labels are generated by a transparent rule system, this question should be interpreted carefully: performance improvements indicate better agreement with the reference construction, not necessarily independently verified geological superiority.

How well do models generalise across unseen wells?

The well-level split makes this a central benchmark property.


Citation

If you use the Lithology Sequence Identification Benchmark in research, publications, model evaluations, or derivative work, please cite this benchmark and acknowledge the underlying Kansas Geological Survey data source.

Benchmark

bibtex
@misc{nora_lithology_sequence_benchmark_2026,
  title        = {Lithology Sequence Identification Benchmark},
  author       = {{NORA Research Lab}},
  year         = {2026},
  publisher    = {NORA Research Lab},
  note         = {Evaluation benchmark for lithological sequence identification from wireline logs}
}

Source Data

The underlying wireline logs originate from the public Kansas Geological Survey digital wireline-log archive.

Please consult the source archive and applicable KGS terms for the appropriate attribution and data-use requirements.


Maintainer

NORA Research Lab

NORA Research Lab develops datasets, benchmarks, models, and tools for artificial intelligence applied to scientific and real-world domains.

![GitHub](https://github.com/Nora-Research-Lab) ![Hugging Face](https://huggingface.co/NoraResearchLab) ![LinkedIn](https://www.linkedin.com/company/nora-research-lab) ![X](https://x.com/noraresearchlab)

Quick Links

Website · GitHub · Hugging Face · LinkedIn · X


Summary

The Lithology Sequence Identification Benchmark contains 245 wells and approximately 1.96 million depth samples for evaluating AI and machine-learning systems that infer lithological sequences from wireline logs.

The central task is:

text
Complete wireline logs
        ↓
Lithology sequence
        ↓
Contiguous depth intervals

The benchmark contains eight possible lithology classes and explicitly supports UNKNOWN/MIXED predictions where the available evidence is insufficient.

Unlike a conventional random sample classification dataset, the benchmark is split by complete wells and evaluates models at both the depth level and the geological sequence level.

The withheld reference labels, deterministic reference-generation pipeline, configuration fingerprints, and multiple evaluation metrics are intended to provide a reproducible framework for comparing automated lithology interpretation systems.