Team Ai
Datasetpublic

Yuchiwang02/smart-logistics-delay-split-v0

Smart Logistics delay split v0 (frozen) The 1,000-row Kaggle Smart Logistics Supply Chain Dataset (user ziya07, listed as CC0 on https://www.kaggle.com/datasets/ziya07/smart-logistics-supply-chain-dataset when read on 2026-09-03) together with the exact 800/200 ShareGPT-style JSONL split on which Yuchiwang02/Llama-3.2-1B-DelaySentinel was fine-tuned in September 2025. What this dataset makes possible A complete, worked example of label leakage, small enough to… See the full description on the dataset page: https://huggingface.co/datasets/Yuchiwang02/smart-logistics-delay-split-v0.

sourceHugging Facecc0-1.0updated 19d agoView on Hugging Face
0likes82downloads
Dataset Card

Smart Logistics delay split v0 (frozen)

The 1,000-row Kaggle Smart Logistics Supply Chain Dataset (user ziya07, listed as CC0 on https://www.kaggle.com/datasets/ziya07/smart-logistics-supply-chain-dataset when read on 2026-09-03) together with the exact 800/200 ShareGPT-style JSONL split on which `Yuchiwang02/Llama-3.2-1B-DelaySentinel` was fine-tuned in September 2025.

What this dataset makes possible

A complete, worked example of label leakage, small enough to audit by hand. Logistics_Delay equals Shipment_Status == "Delayed" OR Traffic_Status == "Heavy" on all 1,000 rows: 566 delayed, 434 not delayed. A depth-2 decision tree matches the published checkpoint's 100% accuracy on the historical 200-row evaluation split, and gradient boosting without the two rule fields scores 0.500 accuracy and 0.452 AUROC. The 100% measures agreement with the label rule. The frozen files make the whole study reproducible: find the leak, compare baselines with a fine-tuned LLM, and probe that model's behavior on the exact rows it was trained and scored on.

Load it and check the rule

The viewer and load_dataset open the original 16-column table as config source, split source. The frozen chat-format split loads from its two files by name.

python
from datasets import load_dataset

repo = "Yuchiwang02/smart-logistics-delay-split-v0"
table = load_dataset(repo, "source", split="source")  # 1,000 rows x 16 columns
chat = load_dataset(repo, data_files={"train": "train.jsonl", "test": "test.jsonl"})  # 800 / 200

df = table.to_pandas()
rule = (df["Shipment_Status"] == "Delayed") | (df["Traffic_Status"] == "Heavy")
print((rule == (df["Logistics_Delay"] == 1)).all())  # True

Check the rule yourself in Data Studio with one line of SQL:

sql
SELECT (Shipment_Status='Delayed' OR Traffic_Status='Heavy') AS rule, Logistics_Delay, COUNT(*) FROM source GROUP BY 1,2

It returns two rows, 566 and 434, with no disagreement.

Columns

source holds 15 input fields and the label, in CSV order: Timestamp, Asset_ID, Latitude, Longitude, Inventory_Level, Shipment_Status, Temperature, Humidity, Traffic_Status, Waiting_Time, User_Transaction_Amount, User_Purchase_Frequency, Logistics_Delay_Reason, Asset_Utilization, Demand_Forecast, and the label Logistics_Delay (1 is delayed). Each JSONL record is one conversations list of three turns: the system prompt, the 15 fields as Column: value lines, and the answer Logistics_Delay: 0|1.

Source and interpretation

The table reads as constructed: coordinates span the whole globe, oceans and both poles included, numeric ranges are visually regular, and 318 label-negative rows carry a delay reason. The Kaggle page documents no collection process. That makes it an instructional dataset, and a sharp one: the leak is total, two column filters expose it, and a 100% headline score hides it.

Rows combine truck status, inventory and customer attributes, with no promised or actual delivery dates. On 118 rows a Delivered status carries label 1 because traffic is Heavy: the label tracks the two fields rather than a measured delivery outcome.

filerowsloads asnote
smart_logistics_dataset.csv1,000config source, split sourcethe Kaggle file, unchanged; its 263 literal None reasons load as null
train.jsonl800data_files, split trainsystem / user / assistant records; unseeded shuffle, September 2025
test.jsonl200data_files, split test116 positive / 84 negative; its first 20 rows were the Trainer eval subset during training
`SPLIT.md`hashes, how the split was made, why it must not be re-split

The JSONL files, including the system prompt they contain, are the author's mechanical transformation of the CSV and are dedicated to the public domain under CC0-1.0 as well.

Code, saved results and model-behavior tests: DelaySentinel on GitHub. The project's leakage scanner prints this report for the table.