Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01BlidReview /steady-rans-generalization Steady-RANS cross-family generalization dataset Data for the paper "Towards generalized flow field prediction: one model across unseen object families" (under double blind review; this account is anonymous for that reason). Trained checkpoints and evaluation code are in the companion model repo: steady-rans-surrogates. Steady incompressible k-omega SST (OpenFOAM simpleFoam) external flow around 855 distinct shapes (17 scripted parametric families plus 40 ModelNet object… See the full description on the dataset page: https://huggingface.co/datasets/BlidReview/steady-rans-generalization.3d1K<n<10K0 likes2.5k downloads2mo agoHugging Face02jiaxin-wen /generalization-dynamics-evals Generalization Dynamics — Main Eval Suite Prepared test sets for the 6 main evaluation families from Generalization dynamics across fine-tuning (Table 1). Use with the unified runner: https://github.com/jiaxin-wen/FT-generalization/tree/main/release from huggingface_hub import snapshot_download root = snapshot_download( repo_id="jiaxin-wen/generalization-dynamics-evals", repo_type="dataset") Or browse a single task (the dataset viewer shows all configs): from datasets… See the full description on the dataset page: https://huggingface.co/datasets/jiaxin-wen/generalization-dynamics-evals.texttext-classification10K<n<100K0 likes178 downloads5mo agoHugging Face03harness-generalization /dataset Harness Generalization Rollouts Evolution-run rollouts for Qwen3-4B-Instruct-2507, Qwen2.5-3B-Instruct, gpt-oss-120b, and gpt-oss-20b. One Parquet file per run is stored at data/<task>/<model>/<configuration>/<timestamp>.parquet. tabular1M<n<10M0 likes92 downloads2mo agoHugging Face04AlignmentResearch /food-preference-generalizationtext1K<n<10K0 likes77 downloads9mo agoHugging Face05cmaldona /Generalization-MultiClass-CLINC150-ROSTDThis dataset merge 3 datasets and have two setup for experiments in generalisation for multi-class clasificacitino task. ID, near-OOD, covariate-shitf: CLINC150 ID, near-OOD, covariate-shitf: ROSTD+OOD (fbreleasecoarse version) far-OOD Validation: SST2 far-OOD Test: News Category (v3) texttext-classification10K<n<100K1 likes67 downloads3y agoHugging Face06cmaldona /All-Generalization-OOD-CLINC150Datasets structure. Attributes: data: text labels: class (str) domain: parent class (str) - This attribute signifies the parent class in the hierarchy and may be absent in some datasets. generalisation: type of OOD Splits: Train: ID: Clinc150 near-OOD: Clinc150 far-OOD: Yelp Validation: ID: Clinc150 near-OOD: Clinc150 far-OOD: SST2 Test: ID: Clinc150 near-OOD: Clinc150 far-OOD: NewCategoryV3 cov-shift: ROSTD+ text10K<n<100K0 likes61 downloads3y agoHugging Face07visv-Bro /repro-flat-minima-and-generalization-insights-from-stochastic-convex-optimization-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes49 downloads2mo agoHugging Face08generalization /sst2_Full-p_05tabular10K<n<100K0 likes37 downloads4y agoHugging Face09generalization /conv_intent_Sampled-p_1tabular10K<n<100K0 likes36 downloads4y agoHugging Face10nz00shuuuu /anomalyxl-generalization AnomalyXL Generalization Real-data, out-of-distribution generalization sets for the AnomalyXL time-series anomaly task, from the paper TimeRLM: Recursive Language Models Are General Temporal Reasoners. Each row is a single, isolated anomaly (or a clean negative) spliced from a real clinical recording and cast into the exact anomalyxl-precise classify_with_evidence schema, so the timeseries_qa environment scores it unchanged. Because no synthetic signal appears here, performance… See the full description on the dataset page: https://huggingface.co/datasets/nz00shuuuu/anomalyxl-generalization.tabularn<1K0 likes32 downloads3mo agoHugging Face11generalization /banking_intent_Full-p_1tabular10K<n<100K0 likes30 downloads4y agoHugging Face12generalization /sst2_Sampled-p_1tabular10K<n<100K0 likes28 downloads4y agoHugging Face13generalization /trec6_Full-p_1tabular10K<n<100K0 likes28 downloads4y agoHugging Face14ajirs /weird-generalization-final-dataset Weird Generalization Final Dataset Clean handoff bundle for the two strongest weird-generalization tasks: 3_1_old_bird_names 3_2_german_city_names This folder intentionally keeps only the data, evaluation materials, and final shareable plots needed to inspect or reuse these tasks. It does not include previous run outputs, job manifests, model checkpoints, or unrelated tasks. Layout datasets/ 3_1_old_bird_names/ train/ test/ original_full/… See the full description on the dataset page: https://huggingface.co/datasets/ajirs/weird-generalization-final-dataset.texttext-generation1K<n<10K0 likes27 downloads5mo agoHugging Face15generalization /conv_intent_Full-p_05tabular10K<n<100K0 likes26 downloads4y agoHugging Face16generalization /banking_intent_Sampled-p_1tabular10K<n<100K0 likes26 downloads4y agoHugging Face17rlundqvist /vea-generalization-benchmark VEA-Generalization Benchmark A diagnostic set of matched response pairs to test whether a Reward Model's dispreference for verbalized evaluation-awareness (VEA) is broad (it penalizes any "I might be being tested" signal) or narrow (it mainly fires on the specific "Wood Labs" cue seen in training). Companion to rlundqvist/ifeval-obf-rl-preferences and the paper "LLM Judges Disprefer Evaluation Awareness." The idea Each item is a matched pair: an identical model… See the full description on the dataset page: https://huggingface.co/datasets/rlundqvist/vea-generalization-benchmark.texttext-classificationn<1K0 likes26 downloads2mo agoHugging Face18generalization /testingtabular10K<n<100K0 likes25 downloads4y agoHugging Face19generalization /sst2_Full-p_1tabular10K<n<100K0 likes23 downloads4y agoHugging Face20spectralbranding /exp-primacy-generalization Experiment E: Primacy Effect Generalization Across LLM Elicitation Formats Dataset Summary This dataset tests whether the serial position (primacy) effect found in JSON-formatted LLM elicitation generalizes to other response formats (natural language, Likert, ranking). A methodological contribution applicable to all LLM-as-respondent research. Records 2,400 calls (2,351 valid, 98.0%) across 4 response formats x 8 Latin-square orderings x 5 focal brands x 5 LLM… See the full description on the dataset page: https://huggingface.co/datasets/spectralbranding/exp-primacy-generalization.tabulartext-generation1K<n<10K0 likes23 downloads3mo agoHugging Face21generalization /banking_intent_Full-p_05tabular10K<n<100K3 likes22 downloads4y agoHugging Face22generalization /trec6_Sampled-p_1tabular10K<n<100K0 likes21 downloads4y agoHugging Face23generalization /trec6_Full-p_05tabular10K<n<100K0 likes20 downloads4y agoHugging Face24generalization /square_Sampled-p_1tabular10K<n<100K0 likes20 downloads4y agoHugging Face25value-generalization /constitution-v3-sfttext100K<n<1M0 likes20 downloads3d agoHugging Face26generalization /square_Full-p_05tabular10K<n<100K0 likes17 downloads4y agoHugging Face27generalization /conv_intent_Full-p_1tabular10K<n<100K0 likes16 downloads4y agoHugging Face28generalization /banking_intent_Sampled-p_05tabular10K<n<100K0 likes13 downloads4y agoHugging Face29generalization /square_Full-p_1tabular10K<n<100K0 likes13 downloads4y agoHugging Face30generalization /newsgroups_Full-p_1tabular10K<n<100K0 likes13 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.