Team Ai
Datasetpublic

asana17/ai_can_anomaly_detection_data

ai_can_anomaly_detection_data The rows the detectors in asana17/ai_can_anomaly_detection are trained, calibrated and tested on, built from CAN logs by assemble.dataset there. A run reads them at one revision and records that revision, so the models in asana17/ai_can_anomaly_detection_runs each name the data they were fitted on. python3 -m evaluate.pc.run asana17/ai_can_anomaly_detection_data <revision> out runs_clone main holds the dataset built from every log. A smaller one… See the full description on the dataset page: https://huggingface.co/datasets/asana17/ai_can_anomaly_detection_data.

sourceHugging Facecc-by-4.0updated 11d agoView on Hugging Face
0likes5.4kdownloads
Dataset Card

aicananomalydetectiondata

The rows the detectors in asana17/ai_can_anomaly_detection are trained, calibrated and tested on, built from CAN logs by assemble.dataset there. A run reads them at one revision and records that revision, so the models in asana17/ai_can_anomaly_detection_runs each name the data they were fitted on.

python3 -m evaluate.pc.run asana17/ai_can_anomaly_detection_data <revision> out runs_clone

main holds the dataset built from every log. A smaller one built for a quick try goes to a branch of its own.

Source

Built from the University of Turku J1939 truck dataset, a Renault Euro VI truck on the road, normal traffic only. It is CC BY 4.0, and so is this. https://etsin.fairdata.fi/dataset/7586f24f-c91b-41df-92af-283524de8b3e/data

The raw logs are not here. Fetch them from the page above.

Files

fileholds
seconds.jsonMIN_SPEED, and each log's seconds above it, which splits the logs into train and test
grid.jsonthe train logs and the grid settings the grid was built with
grid_raw.npythe train logs on a 0.1 s grid, float32, 17 signals a row. The train and calibration rows are cut from it
grid_t.npy, grid_seg.npyeach grid row's time and segment
scale.npythe mean, then the std, every row is z-scored by, fitted to the train rows
built.jsonthe logs and the settings the attack set and the scale were built with
attacked.jsonwhere each attack sits, its PGN, time span, rows and how far it moved them
attacked_raw.npythe test logs with one attack each, on the grid
attacked_rows.npyattacked_raw.npy z-scored by scale.npy
attacked_t.npy, attacked_seg.npy, attacked_label.npy, attacked_wheel.npyeach test row's time, segment, attack label and wheel speed

The signals and the settings are in grid.json and built.json.

Frames

frames/ holds the same attacked test logs frame by frame, for sending the attacks over a real CAN bus. They go with the rows above, the attack set the runs 20260916-001002, 20260916-064753, 20260916-234726 and 20260919-025203 scored.

fileholds
frames/frames.parquetevery test log's frames with its attack in, one row group per log
frames/attacked.jsonattacked.json with each attack's log added
columnholds
logthe source log, such as part_3/20210204093406279004.csv
timestampepoch seconds
can_idthe 29-bit identifier
datathe payload bytes
attackedTrue where the attack changed the payload

A frame can change while no row moves, so score a detector on attacked_label.npy, not on attacked. A row cut short in the source log was dropped on reading.