datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lasa1m-annotate-part-12lasa1m-annotate-part-16lasa1m-annotate-part-07lasa1m-annotate-part-15lasa1m-annotate-part-14lasa1m-annotate-part-06lasa1m-annotate-part-13lasa1m-annotate-part-09lasa1m-annotate-part-05lasa1m-annotate-part-10lasa1m-annotate-part-11lasa1m-annotate-part-08agents-last-exam-data
Agents Last Exam — Task Input Data
Input files (the materials each task hands to the agent at run start) for the
Agents Last Exam (ALE) benchmark. Browsable per-task directory layout.
The Agents Last Exam dataset family
ALE is published as three companion HuggingFace datasets:
Dataset
Contents
Access
Task Card Metadata
One row per task: titles, prompts, taxonomy, input-file descriptors
Open
Task Input Data
The input/ files each task hands the agent at… See the full description on the dataset page: https://huggingface.co/datasets/agents-last-exam/agents-last-exam-data.LaSOT
Dataset Card for LaSOT
Dataset Summary
Large-scale Single Object Tracking (LaSOT) aims to provide a dedicated platform for training data-hungry deep trackers as well as assessing long-term tracking performance.
This repository contains the conference version of LaSOT, published in CVPR-19 (LaSOT: A High-quality Benchmark for Large-scale Single Object Tracking).
LaSOT is featured in:
Large-scale: 1,400 sequences with more than 3.5 millions frames
High-quality: Manual… See the full description on the dataset page: https://huggingface.co/datasets/l-lt/LaSOT.ale-images-qcow2
ALE QEMU runner image
agentslastexam/ale-qemu is the container-side runtime used by the ALE qemu
provider. It packages QEMU, KVM integration, NAT networking, noVNC, and process
supervision. The Ubuntu or Windows guest is supplied separately as
/storage/data.qcow2.
Docker is the container runtime. Dockur is the upstream QEMU-in-Docker project
whose startup and networking stack this image inherits. ALE adds a stable
runner contract around that upstream image.
The image is based on… See the full description on the dataset page: https://huggingface.co/datasets/agents-last-exam/ale-images-qcow2.agents-last-exam-reference
Agents Last Exam — Reference (Ground-Truth) Data
⚠️ Gated dataset. This repo contains the ground-truth / reference outputs
used to score the Agents Last Exam (ALE) benchmark. Access requires login,
agreement to the terms on the access-request form, and manual approval.
Note (06/16/26): This repository was accidentally deleted and has been recreated. The
previous list of approved requesters could not be restored, so even if you
were granted access before, you will need to… See the full description on the dataset page: https://huggingface.co/datasets/agents-last-exam/agents-last-exam-reference.DOTAv1.0refusal-activations
Refusal Activations Dataset
This dataset is now configured to load the full ~97k samples from jailbreak_mixed_100k.csv.
dart_laser_vie
dart_laser_vie — data
Training data + mixes for the DART-LaSER BRIGHT 4B retriever (best model 38.74).
Full guide: https://github.com/abdoelsayed2016/dart_laser_vie (DOCUMENTATION.md).
Contents
reason-embed-data-0928/ — the real per-domain training data (12 <domain>-formatted.jsonl): ReasonEmbed-format
{prompt, query, pos, neg, train_group_size, batch_size}. This is what everything is built from.
combo_rank_aug/ — the base mix (per-domain + aug.jsonl 81k +… See the full description on the dataset page: https://huggingface.co/datasets/abdoelsayed/dart_laser_vie.lists-activationsagents-last-exam
Agents Last Exam — Task Card Metadata (v1.1)
A metadata-only release (v1.1) of 151 tasks from the Agents Last Exam (ALE)
benchmark for evaluating computer-use agents on long-horizon professional work.
The Agents Last Exam dataset family
ALE is published as three companion HuggingFace datasets:
Dataset
Contents
Access
Task Card Metadata
One row per task: titles, prompts, taxonomy, input-file descriptors
Open
Task Input Data
The input/ files each task… See the full description on the dataset page: https://huggingface.co/datasets/agents-last-exam/agents-last-exam.science-activationsp2-dynamic-granger-lasso-resultsdart_lasermetaphors-activationssycophancy-activationsauthority-activationslastlayer-residuals-anon-20260630lasa1m-annotate-part-04agents-last-exam-data-archive
Agents Last Exam — Task Data Archive (input + reference)
⚠️ Gated dataset. This repo packages each task's input, software, and
reference (ground-truth) data into a single archive (ale-tasks-data.tar.gz)
for convenient one-shot download — in particular for running ALE locally with
the local Docker provider,
which fetches it and mounts each task's data at run time. Because it includes
the reference outputs used to score runs, access requires login, agreement to
the terms on the… See the full description on the dataset page: https://huggingface.co/datasets/agents-last-exam/agents-last-exam-data-archive.
