Team Ai
Datasetpublic

serrooT/selene-android-paper-artifacts

SELENE Android Paper Artifacts This dataset contains the compact, paper-facing artifacts produced by SELENE from ARTEMIS dynamic Android malware analyses. It is organized by artifact type, with android10 and android14 as execution-environment splits. These are not train/test splits. The release preserves the analyzed emulator observations as produced. Malware-controlled endpoints, paths, identifiers, and credential-shaped strings are not redacted. What is included… See the full description on the dataset page: https://huggingface.co/datasets/serrooT/selene-android-paper-artifacts.

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
0likes48downloads
Dataset Card

SELENE Android Paper Artifacts

This dataset contains the compact, paper-facing artifacts produced by SELENE from ARTEMIS dynamic Android malware analyses. It is organized by artifact type, with android10 and android14 as execution-environment splits. These are not train/test splits.

The release preserves the analyzed emulator observations as produced. Malware-controlled endpoints, paths, identifiers, and credential-shaped strings are not redacted.

What is included

ConfigPurpose
events_l05Temporally compacted syscall-derived events (the paper's L0.5 layer).
events_l1Semantic behavioral events and the default entry point.
analysesFrozen run selection and run identity metadata.
lifecycle_eventsTarget-aware lifecycle observations extracted from logcat.
lifecycle_summaryOne lifecycle summary row per selected run.
run_featuresOne L1 behavioral feature row per selected run.
fidelity_oracleFrozen 1,000-run validation sample for each environment.
fidelity_evidenceEvidence supporting the fidelity oracle labels.
context_compactSmallest auditable text context representation.
context_standardIntermediate-detail text context representation.
context_evidence_richMost detailed released text context representation.
table_04_layer_metricsPer-run raw/L0/L0.5/L1 counts used to recompute paper Table 4.
table_05_storage_metricsPer-run derived-file inventory plus logical context-pack measurements used to recompute paper Table 5.
table_07_context_metricsPer-run token cost and signal-retention counts used to recompute paper Table 7 without publishing graph bodies.
Environment splitAndroid/APISelected runsL0.5 eventsL1 events
android1010 / 2944,48573,635,32673,673,283
android1414 / 3430,75186,050,14882,843,224

There are 30,746 APK SHA-256 values present in both environments.

Complete L0 syscall tables, suppression diagnostics, parser diagnostics, graph topology, graph visualizations, graph/hybrid context-pack bodies, APK binaries, and redundant intermediate files are intentionally not included. Raw dynamic-analysis artifacts are distributed separately in the `serrooT/artemis-android-dynamic-traces` dataset.

Compact layout

text
data/
├── analyses/<environment>.parquet
├── events_l05/<environment>.parquet
├── events_l1/<environment>.parquet
├── lifecycle_events/<environment>.parquet
├── lifecycle_summary/<environment>.parquet
├── run_features/<environment>.parquet
├── fidelity_oracle/<environment>.parquet
├── fidelity_evidence/<environment>.parquet
├── context_compact/<environment>.jsonl.zst
├── context_standard/<environment>.jsonl.zst
├── context_evidence_rich/<environment>.jsonl.zst
└── paper_metrics/android10/
    ├── table_04_layer_counts.parquet
    ├── table_05_storage_metrics.parquet
    └── table_07_context_metrics.parquet

<environment> is android10 or android14. This shallow layout avoids tens of thousands of small files. Event rows remain keyed by the exact (sha256, run_id) analysis identity and are never merged across runs.

Read any published analysis with Python

In an isolated Python environment, install the two reader dependencies:

bash
pip install pyarrow huggingface_hub

Then set ENVIRONMENT, SHA256, and RUN_ID to any analysis present in the analyses config of `serrooT/selene-android-paper-artifacts`. The example values below select the gustuff run used later in the case study, but the reader itself is generic.

The helper opens the public Parquet through HTTP Range requests, uses sha256 row-group statistics, and downloads only candidate row groups rather than the complete multi-gigabyte event file:

python
from huggingface_hub import HfFileSystem
import pyarrow.compute as pc
import pyarrow.parquet as pq

REPO_ID = "serrooT/selene-android-paper-artifacts"
REVISION = "main"  # use a commit hash when an immutable snapshot is required
ENVIRONMENT = "android10"
SHA256 = "4d4f85997b11a4a7352910a559991f5bba48660a53b1aceffdf76e5b38292656"
RUN_ID = "99b6eb87-809e-414f-9c0a-069bfd35af33"

fs = HfFileSystem()


def read_run(relative_path):
    hub_path = f"datasets/{REPO_ID}@{REVISION}/{relative_path}"
    with fs.open(hub_path, "rb") as stream:
        parquet = pq.ParquetFile(stream)
        sha_column = parquet.schema_arrow.get_field_index("sha256")
        groups = []
        for index in range(parquet.metadata.num_row_groups):
            statistics = parquet.metadata.row_group(index).column(sha_column).statistics
            if (
                statistics is not None
                and statistics.has_min_max
                and statistics.min <= SHA256 <= statistics.max
            ):
                groups.append(index)
        if not groups:
            raise LookupError(f"APK {SHA256} was not found in {relative_path}")
        table = parquet.read_row_groups(groups)

    run_filter = pc.and_(
        pc.equal(table["sha256"], SHA256),
        pc.equal(table["run_id"], RUN_ID),
    )
    rows = table.filter(run_filter)
    if rows.num_rows == 0:
        raise LookupError(f"run ({SHA256}, {RUN_ID}) was not found in {relative_path}")
    return rows


analysis = read_run(f"data/analyses/{ENVIRONMENT}.parquet")
l05 = read_run(f"data/events_l05/{ENVIRONMENT}.parquet")
l1 = read_run(f"data/events_l1/{ENVIRONMENT}.parquet")

print(analysis.select(["sha256", "run_id", "source_run_id"]))
print("L0.5 events:", l05.num_rows)
print("L1 rows:", l1.num_rows)

selected_l1 = l1.slice(0, 1)
source_id = selected_l1["source_event_l0c_id"][0].as_py()
selected_l05 = l05.filter(pc.equal(l05["event_l0c_id"], source_id))
print(selected_l1.select(["event_l1_id", "event_type", "source_event_l0c_id"]))
print(selected_l05.select(["event_l0c_id", "syscall_group", "resource_key"]))

Replace the three uppercase selection values to inspect another run. analyses maps the SELENE run_id to the ARTEMIS source_run_id; L1 then links to L0.5 through source_event_l0c_id. A run can legitimately have no row in an event layer if that layer emitted no event for it.

For bulk analysis, download the required environment files and use the same identity filter locally. Avoid downloading both complete event tables merely to inspect one run.

Paper metric contracts

The three files under data/paper_metrics/android10/ are compact inputs, not copied final tables:

  • —Table 4 has one row for each of the 44,485 selected runs, including raw-line, L0 materialized/suppressed, L0.5, L1, and parser-failure counts.
  • —Table 5 has one physical-file inventory row per selected run and derived layer, plus five logical-JSONL context measurements. Physical Parquet sizes and logical context sizes remain distinct because that is the measurement basis used in the paper.
  • —Table 7 has one row per selected run and context tier with estimated tokens and eligible/retained signal counts. It reproduces graph/hybrid comparisons without distributing graph bodies.

Raw-storage inputs for paper Tables 3 and 10 live in the `serrooT/artemis-android-dynamic-traces` dataset. The public SELENE project validates these contracts and recomputes the paper results with selene reproduce.

Processing hierarchy and the public SELENE project

The released data follows this hierarchy:

text
ARTEMIS raw strace -> SELENE L0 -> L0.5 events -> L1 semantic events
ARTEMIS raw logcat ---------------------------> lifecycle evidence
L1 + lifecycle ------------------------------> features and contexts

L0 is a structured normalization of raw strace lines. It retains syscall time, process identity, arguments, return information, resource hints, and raw-line provenance. L0 is not duplicated in this dataset because it is the largest reconstructible intermediate layer.

The maintained implementation is the public SELENE repository. Its current interfaces are:

bash
# Process one local plain-text or Zstandard-compressed strace.
selene process-trace TRACE_PATH \
  --include-l0 \
  --output-dir OUTPUT_DIRECTORY

# Process one ARTEMIS TAR shard, a shard directory, or an original ARTEMIS tree.
selene process-artemis ARTEMIS_TAR_OR_DIRECTORY \
  --context android10 \
  --include-l0 \
  --output-dir OUTPUT_DIRECTORY

# Recompute the paper claims from the frozen public dataset inputs.
selene reproduce \
  --cache-dir .cache/selene \
  --output-dir reproduced-paper-tables

Omit --include-l0 when only L0.5 and L1 outputs are required. The current CLI does not expose an arbitrary published-run query command; use the generic Python reader above for that purpose. Full installation, resource requirements, expected outputs, and experiment commands are documented in the SELENE README.

Identity and future additions

Rows use context_id, sha256, and run_id as SELENE run identity. The analyses config maps each selected run to the exact ARTEMIS source_run_id. Use (context_id, sha256, source_run_id) to locate the raw analysis; do not join the two datasets by SHA-256 alone. event_l0c_id and event_l1_id identify events within one SELENE run.

Future Android environments can be added as new split files, for example data/events_l1/android15.parquet. New collection snapshots must be versioned and must not silently modify the frozen paper snapshot. APK age or collection cohorts should be derived from VirusTotal first_seen metadata rather than inferred from old directory names.

Case study: the gustuff provenance path

The generic reader above uses this Android 10 analysis as its concrete example:

  • —APK SHA-256: 4d4f85997b11a4a7352910a559991f5bba48660a53b1aceffdf76e5b38292656;
  • —SELENE run_id: 99b6eb87-809e-414f-9c0a-069bfd35af33; and
  • —ARTEMIS source_run_id: f892c32f-5ee9-43f5-861e-13d2fce63470.

It returns 4,645 L0.5 events and 4,996 L1 rows. One audited L1 event is:

text
4d4f85997b11a4a7352910a559991f5bba48660a53b1aceffdf76e5b38292656_99b6eb87-809e-414f-9c0a-069bfd35af33_00000219_NETWORK_EXTERNAL

Selecting that identifier in l1 and following its source is the same operation used for any other event:

python
event_id = (
    "4d4f85997b11a4a7352910a559991f5bba48660a53b1aceffdf76e5b38292656_"
    "99b6eb87-809e-414f-9c0a-069bfd35af33_00000219_NETWORK_EXTERNAL"
)
target = l1.filter(pc.equal(l1["event_l1_id"], event_id))
source_id = target["source_event_l0c_id"][0].as_py()
source = l05.filter(pc.equal(l05["event_l0c_id"], source_id))
print(target.select(["event_type", "resource_key", "n_l0_calls"]))
print(source.select(["syscall_group", "resource_key", "n_calls", "evidence_line_refs"]))

The result is a NETWORK_EXTERNAL interpretation of one NETWORK_IO L0.5 event. It condenses 13 calls involving TCP:54.156.36.158:443 and retains representative raw line references 14539;14542;14547;17169;17173.

StagePublic location or key
L1data/events_l1/android10.parquet, selected by (sha256, run_id, event_l1_id)
L0.5data/events_l05/android10.parquet, selected by source_event_l0c_id
L0Rebuilt by SELENE from the byte-preserving raw strace
Raw indexARTEMIS `data/android10/artifacts.parquet`, selected by (sha256, source_run_id, artifact_type)
Raw shardARTEMIS `data/android10/shards/part-00086.tar`
Raw member<sha256>/analysis_runs/f892c32f-5ee9-43f5-861e-13d2fce63470/results/strace.txt.zst

For reviewers who want to rebuild this fixed paper sample through L0, L0.5, and L1 and verify the five raw lines, the SELENE repository provides the bounded PoC:

bash
selene audit --sample reviewer \
  --cache-dir .cache/selene \
  --output-dir reviewer-sample

This last command is intentionally fixed to the paper sample; it is not the generic dataset reader. Arbitrary published runs are selected with the Python identity filter above, while arbitrary raw traces or ARTEMIS batches are processed with process-trace or process-artemis.

Limitations

The artifacts represent behavior observed during finite Monkey-driven emulator runs. Absence of an event is not proof that an APK cannot perform that behavior. SELENE event categories organize observed evidence; they are not malware verdicts, family predictions, or ATT&CK attributions.

License and mandatory citation

The dataset is available under the SELENE Paper Artifacts Data License 1.0. Because every SELENE artifact is derived from ARTEMIS observations, use must also comply with the upstream ARTEMIS Dynamic Traces Data License 1.0. Commercial and non-commercial use, analysis, modification, derivative datasets, and redistribution are permitted subject to both licenses. Public outputs must cite both the ARTEMIS and SELENE papers as specified in CITATION.md.

The ARTEMIS paper is available at https://doi.org/10.5753/sbseg.2025.11393. The SELENE paper, Do Bruto ao Brilho: Compactação Semântica com Preservação de Proveniência para Traços Dinâmicos de Malware Android, has been accepted at SBSeg 2026; its DOI is pending.

Authors and contact

  • —Cláudio Torres Júnior and André Grégio — Departamento de Informática, Universidade Federal do Paraná (UFPR); SecRET, Curitiba, Paraná, Brazil.
  • —João Pincovscy — Centro de Pesquisa e Desenvolvimento para a Segurança das Comunicações (CEPESC); Universidade do Distrito Federal Professor Jorge Amaury Maia Nunes (UnDF), Brasília, Distrito Federal, Brazil.

Contact: claudio.torres@ufpr.br, gregio@ufpr.br, pincovscy@cepesc.gov.br.