datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
androidlife-public
AndroidLife — Public 60-task preview + run trajectories
This dataset publishes the public 60-task AndroidLife sample (task definitions +
sidecars) together with the full per-run execution artifacts from real-phone
benchmarks (OnePlus CPH2423 via ADB/MobileRun + real LLMs).
The full 530-task corpus lives in
YuvrajSingh9886/androidlife-530.
Dataset preview (tasks)
Open the Dataset Viewer above (config tasks) for a table of all 60 public
tasks: task_id, day… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/androidlife-public.AndroidCodeandroidlife-530
AndroidLife-530 — Android agent benchmark (real phone, real LLM)
AndroidLife runs Android agent tasks against a real phone (via ADB/MobileRun)
and a real LLM, and grades the agent on reaching a verifiable device end-state.
This repo ships the 530-task corpus plus everything needed to reproduce runs.
Benchmark, or template — your call. The 530 tasks are an extended version
of the benchmark, usable as a larger evaluation set for further benchmarking of
models beyond the 60-task… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/androidlife-530.Android-Malware-DatasetAndroidFlux_RL_Train
AndroidFlux RL — policy prompts
Slice
Prompts
Source
t_minus_n
2,459
AndroidFlux source trajectories
successful
2,623
AndroidFlux source trajectories
t_minus_1
2,459
AndroidFlux source trajectories
t
2,459
AndroidFlux source trajectories
Subtotal
10,000
ui_genie
10,000
UI-Genie reward-model prompts (5,000-prompt core marked by in_reduced)
The four AndroidFlux slices are drawn from replayed source trajectories.
The successful slice contains 53… See the full description on the dataset page: https://huggingface.co/datasets/Gyubeum/AndroidFlux_RL_Train.AndroidControlAndroidControlParsed-20kAndroidFlux_RM_Eval
AndroidFlux RM Eval
The evaluation sets that AndroidFlux/rm_benchmark
scores reward models on, repacked as self-contained parquet.
The original sets are JSONL files that reference screenshots by absolute path on
the machine that built them. Here the image bytes travel with the rows, so the
data is portable. Bytes are copied verbatim — sha256(encoded_bytes) equals
both the image_id and the hash of the original file, with no re-encoding.
2026-08-26: android_flux_pairwise (203… See the full description on the dataset page: https://huggingface.co/datasets/Gyubeum/AndroidFlux_RM_Eval.AndroidDaily
AndroidDaily Dataset
This repository hosts the AndroidDaily dataset, a benchmark grounded in real-world mobile usage patterns, introduced in the paper Step-GUI Technical Report.
The AndroidDaily benchmark comprises 3146 static actions and 235 end-to-end tasks across high-frequency daily scenarios. It is specifically designed to assess whether GUI agents can handle authentic everyday usage, providing a robust evaluation for GUI automation capabilities.
Project Page:… See the full description on the dataset page: https://huggingface.co/datasets/stepfun-ai/AndroidDaily.AndroidControl_testselene-android-paper-artifacts
SELENE Android Paper Artifacts
This dataset contains the compact, paper-facing artifacts produced by SELENE from
ARTEMIS dynamic Android malware analyses. It is organized by artifact type, with
android10 and android14 as execution-environment splits. These are not
train/test splits.
The release preserves the analyzed emulator observations as produced. Malware-controlled
endpoints, paths, identifiers, and credential-shaped strings are not redacted.
What is included… See the full description on the dataset page: https://huggingface.co/datasets/serrooT/selene-android-paper-artifacts.myket-android-application-recommendation-dataset
Myket Android Application Install Dataset
This dataset contains information on application install interactions of users in the Myket android application market. The dataset was created for the purpose of evaluating interaction prediction models, requiring user and item identifiers along with timestamps of the interactions.
Data Creation
The dataset was initially generated by the Myket data team, and later cleaned and subsampled by Erfan Loghmani a master student at… See the full description on the dataset page: https://huggingface.co/datasets/erfanloghmani/myket-android-application-recommendation-dataset.artemis-android-dynamic-traces
ARTEMIS Android Dynamic Traces
ARTEMIS executes Android applications in controlled emulators and collects dynamic
analysis artifacts. This public dataset contains the runs performed in Android 10
(API 29) and Android 14 (API 34). APK binaries are not included. The public dataset is
serrooT/artemis-android-dynamic-traces.
The dataset preserves the analysis artifacts as collected. Trace contents are not
redacted, sanitized, normalized, or recompressed. Existing Zstandard files… See the full description on the dataset page: https://huggingface.co/datasets/serrooT/artemis-android-dynamic-traces.ornith-android-instruct
Ornith Android Instruct
5,500 instruction-tuning examples for Android development — Kotlin-first, Jetpack Compose-first, current (non-deprecated) APIs. Built as the training corpus for Ornith, a small language model specialized in writing Android application code, and usable for fine-tuning any code LLM toward modern Android development.
Train: 5,225 examples (ornith_dataset_merged.jsonl)
Eval: 275 held-out examples (ornith_eval_merged.jsonl)
Format: JSONL, one example per line… See the full description on the dataset page: https://huggingface.co/datasets/giggiovpg/ornith-android-instruct.android_control_train
Processed Android Control Training Set
Dataset Description
This repository contains the processed training set derived from the Android Control dataset by Google Research.
The data processing methodology is identical to that used for our corresponding test set, which can be found at Reallm-Labs/android_control_test.
Data Content and Image Extraction
Important Note: Due to the large size of the dataset, this repository contains only the processed text files.… See the full description on the dataset page: https://huggingface.co/datasets/InfiX-ai/android_control_train.AndroidDaily
AndroidDaily Dataset
This repository hosts the AndroidDaily dataset, a benchmark grounded in real-world mobile usage patterns, introduced in the paper Step-GUI Technical Report.
The AndroidDaily benchmark comprises 3146 static actions and 235 end-to-end tasks across high-frequency daily scenarios. It is specifically designed to assess whether GUI agents can handle authentic everyday usage, providing a robust evaluation for GUI automation capabilities.
Project Page:… See the full description on the dataset page: https://huggingface.co/datasets/XiWang12/AndroidDaily.sei-cert-android-rules
Dataset Card for SEI CERT Android Secure Coding Standard (Wiki rules)
Structured export of the SEI CERT Android Secure Coding Standard from the SEI wiki.
Dataset Details
Dataset Description
Android-focused secure coding rules (Java / Android APIs) with descriptive text and examples as published by SEI CERT.
Curated by: Derived from public SEI CERT wiki pages; packaged as CSV by the dataset maintainer.
Funded by [optional]: [More Information Needed]
Shared… See the full description on the dataset page: https://huggingface.co/datasets/safebuffer/sei-cert-android-rules.AndroidFlux_Source_Traj_dev
AndroidFlux Source Trajectories
This dataset packages the complete 116-task AndroidFlux source trajectories for
15 source models/runs. Every model is stored in a separate replay parquet and a
separate replay-history parquet. The schemas are identical to
Gyubeum/AndroidFlux_Failure_Recovery_Eval, so the existing AndroidFlux
materializer and replay tooling can consume the files without conversion.
Dataset scope
Split
Label provenance
Models
Rows per model… See the full description on the dataset page: https://huggingface.co/datasets/Gyubeum/AndroidFlux_Source_Traj_dev.cqadupstack-android-pl-qrelsPart of BEIR-PL: Zero Shot Information Retrieval Benchmark for the Polish Language.
Link to arxiv: https://arxiv.org/pdf/2305.19840.pdf
Contact: konrad.wojtasik@pwr.edu.pl
HF-Android-AppAndroidFlux_RL_Train_Test
AndroidFlux RL Train Test
Test-only AndroidFlux RL data, in two directories.
Directory
Rows
A row is
Docs
data_from_failure_recovery/
1,107
a checkpoint-to-next-action prompt; the policy generates
docs/DATA_CONSTRUCTION.md
data_from_rm_eval/
2,000
a prompt plus a candidate action, with the preference stored outside the conversation
docs/DATA_FROM_RM_EVAL.md
Shared code lives in code/.
The tables in each directory have different schemas, so they load as separate… See the full description on the dataset page: https://huggingface.co/datasets/Gyubeum/AndroidFlux_RL_Train_Test.
