datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
agent-simulations
Agent Simulations
Made with the whileai SDK · Collections: Simulation, Start here: foundational post-training datasets
53,971 synthetic agent trajectories generated by simulations
across 34 agent types. The rows include successful and failed
trajectories for supervised fine-tuning, preference work, reinforcement learning, and
evaluation.
NOTE: This is generated test and training data, not curated ground truth. Review and
filter it for your application before training or… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/agent-simulations.text-to-sql-shop
Text-to-SQL on a seeded store schema, with checkpoints
Recipe: recipes/04-train/text-to-sql · Collection: Analyst
A question about an online store's database in, one PostgreSQL query out, graded by
a program: run the query, compare the result set to the gold query's result. The
schema (8 tables, seeded, schema.sql + seed.sql), the verifier, the trainer and
the benchmark runner are the
recipes/04-train/text-to-sql
recipe in the open-source whileai SDK.
Splits
|… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/text-to-sql-shop.identity-behavior
identity-behavior
Recipe: recipes/04-train/identity · Collections: Character, Start here: foundational post-training datasets
Teach an open model who it is.
Identity behavior is the simplest thing every shipped assistant needs and
open models do not have out of the box: a consistent answer to "who are
you?" and "who made you?", in every phrasing and every language, without a
system prompt propping it up. Ask a base Qwen model and it tells you about
Alibaba; put a persona in the… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/identity-behavior.tool-call-efficiency
tool-call-efficiency
Made with the whileai SDK · Collections: Efficiency, Start here: foundational post-training datasets
Teach an agent to make every tool call count.
An agent that calls a tool twice with the same arguments, looks up what
the user just told it, or keeps calling after the task is done is slow,
expensive, and harder to trust. Ask a base Qwen3-4B to work through
1,133 tool-using tasks across six agents and it does this a lot:
only 52% of its 6,681 rollouts finish… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/tool-call-efficiency.airline-voice-concise
airline-voice-concise
Made with the whileai SDK · Used by: recipes/community/airline-voice-concise-under-probe-outcome-filter · Collection: Register
Training data for putting a speaking register into a model's weights. An
airline support agent that leads with the answer and stops, trained so the
register survives with no instruction in the prompt.
Trained on this set, Qwen3-4B goes from 2.2% to 92.1% of held-out replies
in the register, and becomes less likely to omit required… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/airline-voice-concise.airline-resist-jailbreaks
airline-resist-jailbreaks
Made with the whileai SDK · Collection: Robustness
Jailbreak resistance for a customer support agent, trained on simulated
attacks and tested on real ones.
The real attacks come from elder-plinius/L1B3RT4S,
a public library of working jailbreaks. We read it to extract the attack
techniques and never trained on a single string from it. It is the
evaluation set, unseen by the model.
On 165 unseen blocks from a public jailbreak library the agent holds its… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/airline-resist-jailbreaks.character-training-model-spec
Character training on the OpenAI Model Spec
Recipe: examples/character · Collection: Character
Graded replies for character training: a model answering the style prompts of the
OpenAI Model Spec (8 traits) under a bare
deployment prompt, judged against each trait's principle. Made by
examples/character
in the while-ai SDK (0.24); the recipe is
docs/character-training.md
and the page is docs.withwhile.com.
split
rows
prompts
pass rate
what
train
60
15
0.72
the spec's… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/character-training-model-spec.tau2-simulated
tau2 Simulated Training Set
Made with the whileai SDK · Collections: Simulation, Start here: foundational post-training datasets
The training set that took a base model from 5% to 30% on tau2-bench
telecom, made from nothing but the agent's tool list and policy.
If you build a customer-facing agent, you already have the two files this
dataset was made from: the tools it can call and the policy it follows.
The whileai SDK turned those into 1,057 graded conversations across the… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/tau2-simulated.ecommerce-intent
E-Commerce Intent
Made with the whileai SDK · Collection: Ecommerce Intent Detection
Customer conversations labeled with payment intent, built for training small models that verify what a user actually asked for before an AI agent acts on it. Each conversation carries one structured intent object over seven types: spend, send, exchange, recur, bill, reverse, none.
How it was made
Not scraped, not templated. While builds e-commerce intent data as a multi-agent… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/ecommerce-intent.retail-voice-concise
retail-voice-concise
Made with the whileai SDK · Collection: Register
The same speaking register as
airline-voice-concise,
trained on a different agent. A retail support agent that leads with the
answer and stops.
This exists to test the limitation stated on the airline card: that nothing
there showed the register transfers off airline content. It does. Same
constitution, same recipe, different world, different tools, different
records.
Trained on this set, Qwen3-4B goes from… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/retail-voice-concise.brand
brand
Made with the whileai SDK · Not a training set
Brand assets for the while-ai organization page: the banner the org card shows and the mark. The source of truth for the palette, the wordmark and the whale mark is BRAND.md in the website repo. Nothing here is training data.
