Team Ai
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jiaxin-wen /generalization-dynamics-evals Generalization Dynamics — Main Eval Suite Prepared test sets for the 6 main evaluation families from Generalization dynamics across fine-tuning (Table 1). Use with the unified runner: https://github.com/jiaxin-wen/FT-generalization/tree/main/release from huggingface_hub import snapshot_download root = snapshot_download( repo_id="jiaxin-wen/generalization-dynamics-evals", repo_type="dataset") Or browse a single task (the dataset viewer shows all configs): from datasets… See the full description on the dataset page: https://huggingface.co/datasets/jiaxin-wen/generalization-dynamics-evals.texttext-classification10K<n<100K0 likes178 downloads5mo agoHugging Face02visv-Bro /repro-flat-minima-and-generalization-insights-from-stochastic-convex-optimization-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes49 downloads2mo agoHugging Face03ajirs /weird-generalization-final-dataset Weird Generalization Final Dataset Clean handoff bundle for the two strongest weird-generalization tasks: 3_1_old_bird_names 3_2_german_city_names This folder intentionally keeps only the data, evaluation materials, and final shareable plots needed to inspect or reuse these tasks. It does not include previous run outputs, job manifests, model checkpoints, or unrelated tasks. Layout datasets/ 3_1_old_bird_names/ train/ test/ original_full/… See the full description on the dataset page: https://huggingface.co/datasets/ajirs/weird-generalization-final-dataset.texttext-generation1K<n<10K0 likes27 downloads5mo agoHugging Face04rlundqvist /vea-generalization-benchmark VEA-Generalization Benchmark A diagnostic set of matched response pairs to test whether a Reward Model's dispreference for verbalized evaluation-awareness (VEA) is broad (it penalizes any "I might be being tested" signal) or narrow (it mainly fires on the specific "Wood Labs" cue seen in training). Companion to rlundqvist/ifeval-obf-rl-preferences and the paper "LLM Judges Disprefer Evaluation Awareness." The idea Each item is a matched pair: an identical model… See the full description on the dataset page: https://huggingface.co/datasets/rlundqvist/vea-generalization-benchmark.texttext-classificationn<1K0 likes26 downloads2mo agoHugging Face05spectralbranding /exp-primacy-generalization Experiment E: Primacy Effect Generalization Across LLM Elicitation Formats Dataset Summary This dataset tests whether the serial position (primacy) effect found in JSON-formatted LLM elicitation generalizes to other response formats (natural language, Likert, ranking). A methodological contribution applicable to all LLM-as-respondent research. Records 2,400 calls (2,351 valid, 98.0%) across 4 response formats x 8 Latin-square orderings x 5 focal brands x 5 LLM… See the full description on the dataset page: https://huggingface.co/datasets/spectralbranding/exp-primacy-generalization.tabulartext-generation1K<n<10K0 likes23 downloads3mo agoHugging Face06namakoo /idfu-generalization-specialty IDFU Generalization (Transformers) Specialty Pack — $9 Trial Pack Single-domain Python failure dataset focused on Advanced_Generalization_and_Overfitting_Mitigation_in_Transformers, designed as a low-cost entry point to the IDFU Code Failure Dataset family. Full pack size 87 samples Price $9 USD Free preview in this repo 10 samples (data_sample.jsonl) Buyer profile ML training engineer Type Trial / starter pack (single-domain focus) For broader 19-domain… See the full description on the dataset page: https://huggingface.co/datasets/namakoo/idfu-generalization-specialty.texttext-classificationn<1K0 likes12 downloads5mo agoHugging Face07CircularBalls /tt638d-four-function-targeted-generalization-v1 TT638D Four Function Targeted Generalization Goal: Fix TT638C dense failures before any dyadic/Mercy proof. TT638C evidence: seen failed on uppercase canonical prompts for SUBTRACT, MULTIPLY, DIVIDE heldout failed on divide phrasing: Create divide so it computes a over b. TT638D changes: balanced uppercase function-template coverage across ADD/SUBTRACT/MULTIPLY/DIVIDE explicit regression gate for TT638C failed prompts targeted divide a over b coverage fresh heldout prompts… See the full description on the dataset page: https://huggingface.co/datasets/CircularBalls/tt638d-four-function-targeted-generalization-v1.text10K<n<100K0 likes8 downloads3mo agoHugging Face08ashish-soni08 /asynchow-procedure-generalizationtext1K<n<10K0 likes8 downloads3mo agoHugging Face09skinny-cloud /dx3-recall-generalization-benchmark Dx3 Recall Generalization Benchmark — held-out v0 Author: Asif Waliuddin · NXTG.AI · CC BY 4.0 A held-out recall goldenset (20 query→target pairs), disjoint from the tuning set, built to test whether a retrieval fix generalizes rather than memorizes the set it was tuned against. Every target was full-text-confirmed present in the live store before inclusion — so a recall@5 miss means weak retrieval/ranking, not absence. Honest methodology artifact: a generalization gate, not a… See the full description on the dataset page: https://huggingface.co/datasets/skinny-cloud/dx3-recall-generalization-benchmark.textn<1K0 likes5 downloads2mo agoHugging Face10CircularBalls /tt638c-four-function-prompt-generalization-v1 TT638C Four Function Prompt Generalization Goal: Improve prompt generalization for ADD/SUBTRACT/MULTIPLY/DIVIDE. This version adds: more phrasing variation symbolic operator cues stronger DIVIDE coverage seen and held-out exact generation gates text1K<n<10K0 likes4 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.