Team Ai
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Interplay-LM-Reasoning /composition On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models Charlie Zhang, Graham Neubig, Xiang Yue Carnegie Mellon University, Language Technologies Institute Does Reinforcement Learning Truly Extend Reasoning? This work explores the discrepancy in views on RL's effectiveness in extending language models' reasoning abilities. Some characterize RL as a capability refiner, while others see it as inducing new compositional skills. This challenge… See the full description on the dataset page: https://huggingface.co/datasets/Interplay-LM-Reasoning/composition.tabularquestion-answering100M<n<1B2 likes368 downloads9mo agoHugging Face02obaydata /multi-image-composition-instruction-following Multi-Image Composition Instruction-Following A large-scale multimodal dataset for multi-image composition via natural language instruction-following. Each case provides 2-3 input images (characters + scene) along with detailed Chinese instructions to compose them into a single photorealistic output image. Designed for training and evaluating models on complex image composition tasks that require understanding of character identity preservation, pose generation, scene integration… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/multi-image-composition-instruction-following.imageimage-to-imagen<1K0 likes254 downloads7mo agoHugging Face03mainlp /Compositional-ARCCompositional-ARC: Assessing Systematic Generalization in Abstract Spatial Reasoning Philipp Mondorf, Shijia Zhou, Monica Riedler, and Barbara Plank. (2026). Compositional-ARC: Assessing systematic generalization in abstract spatial reasoning. In The Fourteenth International Conference on Learning Representations. Systematic generalization refers to the capacity to understand and generate novel combinations from known components. Despite recent progress by large language… See the full description on the dataset page: https://huggingface.co/datasets/mainlp/Compositional-ARC.texttext-generation100K<n<1M0 likes158 downloads8mo agoHugging Face04goodevening /composition-10B-valtabular1K<n<10K0 likes67 downloads1y agoHugging Face05goodevening /composition-10B-testtabular10K<n<100K0 likes51 downloads1y agoHugging Face06goodevening /composition-10B-rltabular100K<n<1M0 likes43 downloads1y agoHugging Face07lucky-verma /dyt-composition-artifacts DyT Composition Study Artifacts This dataset contains sanitized result manifests and analysis outputs for When Does Removing LayerNorm Help? Activation Bounding as a Regime-Dependent Implicit Regularizer. DOI: https://doi.org/10.48550/arXiv.2604.23434 Contents The artifacts include aggregate training metrics, saturation measurements, statistical-test summaries, predictor-validation outputs, table-source manifests, and selected aggregate analysis files used by the… See the full description on the dataset page: https://huggingface.co/datasets/lucky-verma/dyt-composition-artifacts.textn<1K0 likes30 downloads5mo agoHugging Face08compositional-gsm /compositional_gsmtext1K<n<10K0 likes14 downloads1y agoHugging Face09mdg-nlp /timex-compositional-sentencetext1K<n<10K0 likes12 downloads8mo agoHugging Face10dda71427 /sand_composition.jsontextn<1K0 likes9 downloads9mo agoHugging Face11stair-lab /skill_composition_hypothesisgated A Dataset for Skill Composition Hypothesis We employ a language model to annotate the specific skills assessed by each question derived from various Natural Language Processing benchmarks. The skill taxonomy utilized is sourced from IXL. The associated GitHub repository that produces this dataset can be found here: https://github.com/sangttruong/skill-composition-hypothesis. textfeature-extraction10K<n<100K0 likes2 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.