Team Ai
Datasetpublic

UngLong/radiology-test-v2

Radiology Test v2 Evaluation dataset for the Radiology agent of an AI Medical Department, paired with UngLong/radiology-ready-v2 (training set). CT images are in UngLong/openm3chest-npy-v2. ⚠️ Before using: CT scans for this test set must be uploaded to openm3chest-npy-v2 first. See scripts/test_npy_needed.txt (154 scan keys) for the list to upload via build_npy_hub.py. Dataset Summary Total rows 154 Screening rows 74 (all 8 screening tasks)… See the full description on the dataset page: https://huggingface.co/datasets/UngLong/radiology-test-v2.

sourceHugging Faceotherupdated 4mo agoView on Hugging Face
0likes12downloads
Dataset Card

Radiology Test v2

Evaluation dataset for the Radiology agent of an AI Medical Department, paired with UngLong/radiology-ready-v2 (training set). CT images are in UngLong/openm3chest-npy-v2.

⚠️ Before using: CT scans for this test set must be uploaded to openm3chest-npy-v2 first. See scripts/test_npy_needed.txt (154 scan keys) for the list to upload via build_npy_hub.py.

Dataset Summary

Total rows154
Screening rows74 (all 8 screening tasks)
Detail rows80 (4 nodule characterization tasks)
Unique CT volumes154 (no overlap between screening & detail pools)
Source*_test.json splits from NLST/OpenM3Chest
FormatMulti-task JSON — same as training set

⚠️ Known Limitation: Missing Positive Labels in Screening

This is the most important thing to understand about this test set.

Root Cause

The 8 screening tasks in NLST were annotated independently on different scan subsets:

chest_abn_54_test.json  → 13,346 scans annotated for atelectasis
chest_abn_55_test.json  → 12,722 scans annotated for pleural thickening
nodule_presence_test.json → 10,406 unique scans
...

The multi-task format requires all 8 labels to be present for the same scan. Taking the intersection of all 8 test files leaves only 74 scans — and these happen to be predominantly "normal" scans:

chest_abn_54 (atelectasis)        :  0/74 positive  ← 0%
chest_abn_55 (pleural thickening) :  0/74 positive  ← 0%
chest_abn_57 (chest wall)         :  0/74 positive  ← 0%
chest_abn_56 (mass/adenopathy)    :  2/74 positive  ← 3%
chest_abn_58 (consolidation)      :  2/74 positive  ← 3%
chest_abn_59 (emphysema)          : 28/74 positive  ← 38% ✅
chest_abn_61 (fibrosis)           : 22/74 positive  ← 30% ✅
nodule_presence                   : 43/74 positive  ← 58% ✅

Why the positive cases are missing

In NLST, rare findings (atelectasis 1%, consolidation 0.8%, chest wall 0.1%) are unevenly distributed across scan subsets. Scans with chest_abn_54=Yes may not have been included in the nodule_presence annotation batch, and vice versa. This is a structural characteristic of the original NLST annotation process — not a bug in data processing.

Evaluation implications

TaskEvaluable?Note
nodule_presence✅ Yes43 positive, 31 negative
chest_abn_59 (emphysema)✅ Yes28 positive
chest_abn_61 (fibrosis)✅ Yes22 positive
chest_abn_56 (mass)⚠️ LimitedOnly 2 positive
chest_abn_58 (consolidation)⚠️ LimitedOnly 2 positive
chest_abn_54 (atelectasis)❌ N/A0 positive → recall undefined
chest_abn_55 (pleural)❌ N/A0 positive → recall undefined
chest_abn_57 (chest wall)❌ N/A0 positive → recall undefined
All 4 detail tasks✅ YesBalanced distribution

Always report F1/Recall per task, not overall accuracy. Accuracy will be misleadingly high for tasks with 0 positives (model predicts "No" always → 100% accuracy).


Planned Extension: Single-Task Test Supplements

To properly evaluate the 3 tasks with no positives, supplementary test sets will be added in future versions. These will relax the "all 8 labels required" constraint and test each rare finding independently:

Future datasetTaskFormat
radiology-test-abn54chestabn54 (atelectasis)Single-task
radiology-test-abn55chestabn55 (pleural)Single-task
radiology-test-abn57chestabn57 (chest wall)Single-task

These will be balanced 50/50 and can be used alongside this multi-task set for comprehensive evaluation.


Label Distribution

Screening (74 rows)

TaskFindingPositiveNegative% Yes
chestabn54Atelectasis0740%
chestabn55Pleural thickening/effusion0740%
chestabn56Mass/adenopathy ≥10mm2723%
chestabn57Chest wall abnormality0740%
chestabn58Consolidation2723%
chestabn59Emphysema284638%
chestabn61Fibrosis/honeycombing225230%
nodule_presenceLung nodule433158%

Detail (80 rows)

TaskDistribution
nodule_locationRUL=18%, LUL=21%, LLL=25%, RLL=30%, RML=6%
nodule_attenuationSolid=74%, Ground Glass=18%, Others=9% ≈ train
nodule_marginSmooth=70%, Poorly defined=14%, Spiculated=11%, Unable=5% ≈ train
nodule_size4~6mm=46% (⚠️ higher than train 33%), ≤4mm=26%, 8~15mm=11%
nodule_size: test set skewed toward small nodules (4~6mm overrepresented). Size classification metrics may be slightly optimistic.

Schema

Same as radiology-ready-v2:

FieldTypeDescription
keysstringSeriesInstanceUID → {pids}/{keys}.npy in NPY repo
pidsstringPatient ID
promptstringClinical record + questions + JSON format instruction
responsestringGround truth JSON {task: answer, ...}
task_typestring"screening" or "detail"

Usage

python
import json
import numpy as np
from datasets import load_dataset
from huggingface_hub import hf_hub_download

ds = load_dataset("UngLong/radiology-test-v2", split="train")
row = ds[0]

# Load CT volume (must be uploaded first)
npy_path = hf_hub_download(
    repo_id="UngLong/openm3chest-npy-v2",
    filename=f"{row['pids']}/{row['keys']}.npy",
    repo_type="dataset",
)
volume = np.load(npy_path).astype(np.float32)  # [Z, H, W], HU values

print(row["task_type"])
print(json.loads(row["response"]))

Recommended Evaluation Metrics

python
from sklearn.metrics import f1_score, recall_score, classification_report

# Per-task F1 and Recall
for task in ["nodule_presence", "chest_abn_59", "chest_abn_61", ...]:
    y_true = [json.loads(r["response"])[task] for r in test_rows]
    y_pred = [model_predict(r)[task] for r in test_rows]
    print(f"{task}: F1={f1_score(y_true, y_pred, average='macro'):.3f}")
    # Do NOT use accuracy for imbalanced tasks

Source

Built from OpenM3Chest — NLST-based chest CT dataset. CT images via IDC. Training set: UngLong/radiology-ready-v2.