Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01DesmondYMTang2024 /Language-Grounded_Sparse_Encoder_Training Language-Grounded Sparse Encoder (LanSE) — Training Data This repository hosts the AI-generated images and human annotation datasets accompanying the paper: Human-like Content Analysis for Generative AI with Language-Grounded Sparse Encoders Yiming Tang, Arash Lagzian, Srinivas Anumasa, Qiran Zou, Yingtao Zhu, Ye Zhang, Trang Nguyen, Yih-Chung Tham, Ehsan Adeli, Ching-Yu Cheng, Yilun Du, Dianbo Liu National University of Singapore · Tsinghua University · Stanford University ·… See the full description on the dataset page: https://huggingface.co/datasets/DesmondYMTang2024/Language-Grounded_Sparse_Encoder_Training.textimage-classification100K<n<1M1 likes7k downloads29d agoHugging Face02WHB139426 /Grounded-VideoLLM13 likes6.3k downloads1y agoHugging Face03chenyilun95 /Grounded_3D-LLM_data2 likes3.8k downloads2y agoHugging Face04mlfoundations-cua-dev /easyr1-grounding-dataset-30k-not_grounded-SE-GUI-3B-2MPimage10K<n<100K1 likes1.3k downloads1y agoHugging Face05arpandeepk /swe-zero-grounded-fulltextn<1K0 likes513 downloads4mo agoHugging Face06ShuaiYang03 /Grounded_3D_LLM_with_Referent_Tokens_Dataset Grounded 3D-LLM Dataset For detailed information and resources, please visit the following links: Paper Arxiv Project Website Dataset Access Code We are in the process of releasing our data incrementally: Processed ScanNet200 PCD(~7G): Each .npyfile represents a N*12 array with the following structure: coordinates, color, normals, segments, labels = ( points[:, :3], points[:, 3:6], points[:, 6:9], points[:, 9]… See the full description on the dataset page: https://huggingface.co/datasets/ShuaiYang03/Grounded_3D_LLM_with_Referent_Tokens_Dataset.textquestion-answering0 likes482 downloads2y agoHugging Face07tomhodemon /grounded-visual-spatial-reasoning Grounded Visual Spatial Reasoning Code for generating the annotations can be found here: github.com Dataset Summary This dataset extends the Visual Spatial Reasoning (VSR) dataset with visual grounding annotations: each caption is annotated with COCO-category object mentions, their positions , and corresponding bounding boxes in the image. Data instance Each sample instance has the following structure: Field Type Description image_file string… See the full description on the dataset page: https://huggingface.co/datasets/tomhodemon/grounded-visual-spatial-reasoning.image10K<n<100K3 likes468 downloads1y agoHugging Face08risaleinur /risale-nur-grounded-multipool Risale-i Nur Grounded Multi-Pool LLM Dataset TR. 15 kanonik Risale-i Nur kitabından hazırlanan; kaynak bağlı üretim, SFT, tercih, değerlendirme, sürekli ön eğitim ve erişim çalışmaları için çok görünümlü bir veri seti. EN. A multi-view dataset built from 15 canonical Risale-i Nur books for grounded generation, SFT, preference learning, evaluation, continued pretraining, and retrieval. v2.10.0 · 199 configs · 463 config/split views · 527,196 rows across configured views… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risale-nur-grounded-multipool.tabulartext-generation100K<n<1M4 likes323 downloads28d agoHugging Face09mitermix /audioset-with-grounded-captionsaudio1M<n<10M4 likes236 downloads1y agoHugging Face10geodesic-research /discourse-grounded-misalignment-evals Synthetic Misalignment Propensity Evaluations We evaluate alignment using a suite of synthetic scenario-based evaluations created for this work. Each question presents the AI with a high-stakes setting requiring a choice between two actions: one aligned and one misaligned. The misaligned action is typically framed as instrumentally appealing, making these evaluations a relevant proxy for misaligned AIs across a range of terminal goals (Bostrom, 2012). We measure tendencies toward… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/discourse-grounded-misalignment-evals.tabular1K<n<10K1 likes235 downloads8mo agoHugging Face11schneiderkamplab /dfm13-multilingual-grounded-instruct-sl dfm13_wave4_synthetic_sl_grounded_instruct 20000 complete conversations; 20000 native assistant targets. All user/tool history and tool definitions are preserved. Gemma native student rendering, thinking disabled. Generated and separately model-reviewed by Gemma4 26B A4B; automated judgments are fallible, not human or native-speaker certification. Includes unchanged original accepted conversations and narrowly recovered complete keep reviews rejected solely for an empty… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm13-multilingual-grounded-instruct-sl.texttext-generation10K<n<100K0 likes184 downloads1d agoHugging Face12schneiderkamplab /dfm13-multilingual-grounded-instruct-hu dfm13_wave4_synthetic_hu_grounded_instruct 20000 complete conversations; 20000 native assistant targets. All user/tool history and tool definitions are preserved. Gemma native student rendering, thinking disabled. Generated and separately model-reviewed by Gemma4 26B A4B; automated judgments are fallible, not human or native-speaker certification. Includes unchanged original accepted conversations and narrowly recovered complete keep reviews rejected solely for an empty… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm13-multilingual-grounded-instruct-hu.texttext-generation10K<n<100K0 likes182 downloads1d agoHugging Face13schneiderkamplab /dfm13-multilingual-grounded-instruct-bg dfm13_wave4_synthetic_bg_grounded_instruct 20000 complete conversations; 20000 native assistant targets. All user/tool history and tool definitions are preserved. Gemma native student rendering, thinking disabled. Generated and separately model-reviewed by Gemma4 26B A4B; automated judgments are fallible, not human or native-speaker certification. Includes unchanged original accepted conversations and narrowly recovered complete keep reviews rejected solely for an empty… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm13-multilingual-grounded-instruct-bg.texttext-generation10K<n<100K0 likes180 downloads1d agoHugging Face14schneiderkamplab /dfm13-multilingual-grounded-instruct-sr dfm13_wave4_synthetic_sr_grounded_instruct 20000 complete conversations; 20000 native assistant targets. All user/tool history and tool definitions are preserved. Gemma native student rendering, thinking disabled. Generated and separately model-reviewed by Gemma4 26B A4B; automated judgments are fallible, not human or native-speaker certification. Includes unchanged original accepted conversations and narrowly recovered complete keep reviews rejected solely for an empty… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm13-multilingual-grounded-instruct-sr.texttext-generation10K<n<100K0 likes179 downloads1d agoHugging Face15schneiderkamplab /dfm13-multilingual-grounded-instruct-sq dfm13_wave4_synthetic_sq_grounded_instruct 20000 complete conversations; 20000 native assistant targets. All user/tool history and tool definitions are preserved. Gemma native student rendering, thinking disabled. Generated and separately model-reviewed by Gemma4 26B A4B; automated judgments are fallible, not human or native-speaker certification. Includes unchanged original accepted conversations and narrowly recovered complete keep reviews rejected solely for an empty… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm13-multilingual-grounded-instruct-sq.texttext-generation10K<n<100K0 likes177 downloads1d agoHugging Face16schneiderkamplab /dfm13-multilingual-grounded-instruct-lb dfm13_wave4_synthetic_lb_grounded_instruct 9272 complete conversations; 9272 native assistant targets. All user/tool history and tool definitions are preserved. Gemma native student rendering, thinking disabled. Generated and separately model-reviewed by Gemma4 26B A4B; automated judgments are fallible, not human or native-speaker certification. Includes unchanged original accepted conversations and narrowly recovered complete keep reviews rejected solely for an empty rationale.… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm13-multilingual-grounded-instruct-lb.texttext-generation1K<n<10K0 likes176 downloads1d agoHugging Face17schneiderkamplab /dfm13-multilingual-grounded-instruct-sk dfm13_wave4_synthetic_sk_grounded_instruct 20000 complete conversations; 20000 native assistant targets. All user/tool history and tool definitions are preserved. Gemma native student rendering, thinking disabled. Generated and separately model-reviewed by Gemma4 26B A4B; automated judgments are fallible, not human or native-speaker certification. Includes unchanged original accepted conversations and narrowly recovered complete keep reviews rejected solely for an empty… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm13-multilingual-grounded-instruct-sk.texttext-generation10K<n<100K0 likes175 downloads1d agoHugging Face18Vikhrmodels /Grounded-RAG-RU-v2 Датасет для алайнмента (граундинга) способности LLM отвечать на вопросы по документам (RAG) Этот датасет был собран на основе 13к разных статей из русской Википедии с помошью синтетических вопросов и ответов gpt-4-turbo-1106. Датасет содержит 4047 уникальных кластеров, т.е. комбинаций из документов - улосвная симуляция "найденных результатов" в Retrieval системе. Подробнее описано в разделе "Общие этапы сборки этого датасета". Общий объем датасета - 50210 уникальных диалогов. В… See the full description on the dataset page: https://huggingface.co/datasets/Vikhrmodels/Grounded-RAG-RU-v2.tabular10K<n<100K23 likes174 downloads2y agoHugging Face19schneiderkamplab /dfm13-multilingual-grounded-instruct-be dfm13_wave4_synthetic_be_grounded_instruct 20000 complete conversations; 20000 native assistant targets. All user/tool history and tool definitions are preserved. Gemma native student rendering, thinking disabled. Generated and separately model-reviewed by Gemma4 26B A4B; automated judgments are fallible, not human or native-speaker certification. Includes unchanged original accepted conversations and narrowly recovered complete keep reviews rejected solely for an empty… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm13-multilingual-grounded-instruct-be.texttext-generation10K<n<100K0 likes173 downloads1d agoHugging Face20schneiderkamplab /dfm13-multilingual-grounded-instruct-bs dfm13_wave4_synthetic_bs_grounded_instruct 20000 complete conversations; 20000 native assistant targets. All user/tool history and tool definitions are preserved. Gemma native student rendering, thinking disabled. Generated and separately model-reviewed by Gemma4 26B A4B; automated judgments are fallible, not human or native-speaker certification. Includes unchanged original accepted conversations and narrowly recovered complete keep reviews rejected solely for an empty… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm13-multilingual-grounded-instruct-bs.texttext-generation10K<n<100K0 likes172 downloads1d agoHugging Face21schneiderkamplab /dfm13-multilingual-grounded-instruct-hr dfm13_wave4_synthetic_hr_grounded_instruct 20000 complete conversations; 20000 native assistant targets. All user/tool history and tool definitions are preserved. Gemma native student rendering, thinking disabled. Generated and separately model-reviewed by Gemma4 26B A4B; automated judgments are fallible, not human or native-speaker certification. Includes unchanged original accepted conversations and narrowly recovered complete keep reviews rejected solely for an empty… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm13-multilingual-grounded-instruct-hr.texttext-generation10K<n<100K0 likes167 downloads1d agoHugging Face22schneiderkamplab /dfm13-multilingual-grounded-instruct-fa dfm13_wave4_synthetic_fa_grounded_instruct 20000 complete conversations; 20000 native assistant targets. All user/tool history and tool definitions are preserved. Gemma native student rendering, thinking disabled. Generated and separately model-reviewed by Gemma4 26B A4B; automated judgments are fallible, not human or native-speaker certification. Includes unchanged original accepted conversations and narrowly recovered complete keep reviews rejected solely for an empty… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm13-multilingual-grounded-instruct-fa.texttext-generation10K<n<100K0 likes166 downloads1d agoHugging Face23saitejaalasyam /grounded-qa-preferences Grounded QA preferences Preference pairs for a small RLHF stack. Each row is a passage, a question, a preferred answer, and a rejected answer. The questions, answer spans, and unanswerable labels come from SQuAD 2.0 (Rajpurkar et al.). This dataset does not add new human rankings. A fixed rule turns those annotations into Bradley-Terry pairs: pair_type When Chosen Rejected wrong_span The passage answers the question The gold span A different short span from the same… See the full description on the dataset page: https://huggingface.co/datasets/saitejaalasyam/grounded-qa-preferences.texttext-generation1K<n<10K1 likes160 downloads12d agoHugging Face24NovachronoAI /RAG-Grounded-QA-188k 🎯 RAG Grounded QA 186K The Anti-Hallucination Dataset Teach language models to answer from context — or shut up trying. Built by NovachronoAI — Precision AI for the real world. Full Dataset (186K) · 20K Subset · Schema · Sources · Usage Guide 🧠 Why This Dataset Exists Most QA datasets teach models what to say. This one also teaches them when to stay silent. RAG (Retrieval-Augmented Generation) systems have a fatal flaw: the model hallucinates when… See the full description on the dataset page: https://huggingface.co/datasets/NovachronoAI/RAG-Grounded-QA-188k.tabularquestion-answering100K<n<1M0 likes149 downloads7mo agoHugging Face25chnln /grounded-misunderstandings-in-maptask GMMT: Grounded Misunderstandings in MapTask The Grounded Misunderstandings in MapTask (GMMT) dataset was produced for the LREC 2026 paper Grounded Misunderstandings in Asymmetric Dialogue: A Perspectivist Annotation Scheme for MapTask by Nan Li, Albert Gatt, and Massimo Poesio. It provides perspectivist annotations of the HCRC MapTask corpus, capturing both speaker-intended and addressee-interpreted landmarks for every reference expression (RE). The annotations support… See the full description on the dataset page: https://huggingface.co/datasets/chnln/grounded-misunderstandings-in-maptask.tabular10K<n<100K2 likes148 downloads2mo agoHugging Face2634data /grounded-videollm-coinvideo1K<n<10K0 likes139 downloads2mo agoHugging Face27JosephZ /vg150_grounded_vqaimage10K<n<100K0 likes131 downloads2y agoHugging Face28fittar /visually_grounded_embeddings Visually Grounded embeddings for Fast-text and GloVe This repository contains multiple visually grounded word embedding models. All of these embeddings have been effectively infused with visual information from images. They have been proven to show stronger correlations (compared to textual embeddings) to human judgments on various word similarities and relatedness benchmarks. Usage All of the models are encoded in gensim format. Loading the model: import gensim… See the full description on the dataset page: https://huggingface.co/datasets/fittar/visually_grounded_embeddings.0 likes126 downloads3y agoHugging Face29arpandeepk /swe-zero-grounded-v8textn<1K0 likes123 downloads4mo agoHugging Face30SM-Bello /C172P-Grounded-JSBSim-Airborne-Trim-Failure-Negative-Result c172p Grounded A Negative Result: JSBSim's c172p Could Not Be Trimmed for Level Flight Why this dataset exists Most published aerospace ML/control work only shows what worked. This one doesn't. This is a negative result from the early stage of the PHI-CTRL project (Physics-Hybrid Integrity Control — a fault-tolerant flight control architecture). Before the project settled on the F-16A as its plant model, the original plan was to build and… See the full description on the dataset page: https://huggingface.co/datasets/SM-Bello/C172P-Grounded-JSBSim-Airborne-Trim-Failure-Negative-Result.textother1K<n<10K0 likes123 downloads28d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.