Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01instruction-pretrain /general-instruction-augmented-corpora Instruction Pre-Training: Language Models are Supervised Multitask Learners (EMNLP 2024) This repo contains the general instruction-augmented corpora (containing 200M instruction-response pairs covering 40+ task categories) used in our paper Instruction Pre-Training: Language Models are Supervised Multitask Learners. We explore supervised multitask pre-training by proposing Instruction Pre-Training, a framework that scalably augments massive raw corpora with instruction-response… See the full description on the dataset page: https://huggingface.co/datasets/instruction-pretrain/general-instruction-augmented-corpora.texttext-classification24 likes56k downloads7mo agoHugging Face02radar-generalist /RADAR-auxiliary-data RADAR: Preprocessed Anatomical Masks for Merlin CT Data This dataset provides preprocessed anatomical segmentation masks for the Merlin abdominal CT training set, generated by TotalSegmentator and post-processed for use with the RADAR framework. These masks enable anatomy-aware vision–language pretraining without any additional manual annotation. Overview RADAR is a generalist vision–language model trained on over 400,000 contrast-enhanced abdominal CT… See the full description on the dataset page: https://huggingface.co/datasets/radar-generalist/RADAR-auxiliary-data.image-segmentation10K<n<100K8 likes9.5k downloads23d agoHugging Face03General-Medical-AI /SlideChat Introduction This repository provides the dataset resources used for training and evaluating SlideChat, a multimodal large language model for whole-slide pathology image understanding. The dataset includes both instruction-following training data and VQA/Caption evaluation benchmarks across multiple pathology cohorts and tasks. Contents Training Instruction Data SlideInstruct_train_stage1_caption.json: Slide-level caption instruction data used for Stage-1… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/SlideChat.text100K<n<1M19 likes7.7k downloads1mo agoHugging Face04General-Level /General-Bench-Closeset On Path to Multimodal Generalist: General-Level and General-Bench [📖 Project] [🏆 Leaderboard] [📄 Paper] [🤗 Paper-HF] [🤗 Dataset-HF] [📝 Dataset-Github] Close Set of General-Bench We divide our General-Bench into two settings: open and close. This is the Close Set, where we release only the sample inputs—without ground-truth answers—for 🏆 Leaderboard purpose. To participate the leaderboard, please follow the detailed instructions to submit the evaluation results (submission).… See the full description on the dataset page: https://huggingface.co/datasets/General-Level/General-Bench-Closeset.2 likes7.3k downloads1y agoHugging Face05mariiakoroliuk /generalization-science-data0 likes5.9k downloads17h agoHugging Face06Shofo /shofo-tiktok-general-small Shofo TikTok General (Small) Overview Shofo TikTok General (Small) is a dataset containing 50,000 TikTok videos with comprehensive metadata, transcripts, comments, and engagement metrics. This is a curated subset of Shofo's larger TikTok index, which contains hundreds of millions of indexed videos. Size: ~50K videos (~500GB) Modality: Video + Audio + Text (transcripts, comments, captions) Source: TikTok Schema Column Type Description file_name… See the full description on the dataset page: https://huggingface.co/datasets/Shofo/shofo-tiktok-general-small.tabularvideo-classification10K<n<100K23 likes5.9k downloads8mo agoHugging Face07FineEnvs /MiMo-V2.6-RL-harbor-general MiMo-V2.6-RL General (Harbor) Work in a simulated company through its MCP systems. 925 Harbor tasks from the General domain of Xiaomi's MiMo-V2.6-RL-oss, the RL environments MiMo-V2.6 was trained on, converted so every one runs as a standard Harbor task. Each task is a workplace with 5 to 15 business systems (finance, legal, HR, operations, ...) served over MCP, a workspace of documents, and a brief. The agent works through the systems as an unprivileged user; the task's own… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/MiMo-V2.6-RL-harbor-general.othern<1K0 likes5.7k downloads1d agoHugging Face08General-Medical-AI /Ophora-160Kgated Introduction Ophora-160K contains 162,185 video clip-instruction pairs extracted from 9,819 narrative videos of ophthalmic surgery. The average duration of all clips is 5.54 seconds. All ophthalmic narrative videos were collected from the YouTube platform. The file ophora161k.csv contains all video clip IDs and their generation instructions. The file ophora28k.csv contains content that has been filtered to remove sensitive information such as subtitles and watermarks.… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/Ophora-160K.texttext-to-video100K<n<1M1 likes4.5k downloads1y agoHugging Face09General-Medical-AI /GMAI-VL-5.5M GMAI-VL-5.5M Dataset GMAI-VL-5.5M is a comprehensive, large-scale medical General Medical AI Vision-Language (GMAI-VL) dataset built specifically for training multimodal foundation models in the medical domain. It contains an extraordinary scale of high-quality instructions encompassing over 5.5 million multimodal question-answering pairs, carefully constructed based on hundreds of medical classification, segmentation, and detection datasets. This repository… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/GMAI-VL-5.5M.imagevisual-question-answering1M<n<10M6 likes3.3k downloads6mo agoHugging Face10BlidReview /steady-rans-generalization Steady-RANS cross-family generalization dataset Data for the paper "Towards generalized flow field prediction: one model across unseen object families" (under double blind review; this account is anonymous for that reason). Trained checkpoints and evaluation code are in the companion model repo: steady-rans-surrogates. Steady incompressible k-omega SST (OpenFOAM simpleFoam) external flow around 855 distinct shapes (17 scripted parametric families plus 40 ModelNet object… See the full description on the dataset page: https://huggingface.co/datasets/BlidReview/steady-rans-generalization.3d1K<n<10K0 likes2.5k downloads2mo agoHugging Face11natolambert /GeneralThought-430K-filteredData from https://huggingface.co/datasets/GeneralReasoning/GeneralThought-430K, removed prompts with non commercial data tabular100K<n<1M35 likes2.4k downloads2y agoHugging Face12General-Level /General-Bench-Openset On Path to Multimodal Generalist: General-Level and General-Bench [📖 Project] [🏆 Leaderboard] [📄 Paper] [🤗 Paper-HF] [🤗 Dataset-HF (Close-Set)] [🤗 Dataset-HF (Open-Set)] [📝 Github] Open Set of General-Bench We divide our General-Bench into two settings: Open and Close. This is the Open Set repo, where we release the full ground-truth annotations for all datasets, allowing to train and evaluate models for open research purpose. If you wish to rank on our 🏆 leaderboard, please… See the full description on the dataset page: https://huggingface.co/datasets/General-Level/General-Bench-Openset.4 likes2.2k downloads1y agoHugging Face13General-Medical-AI /GMAI-Reasoning10K GMAI-Reasoning10K Medical Reasoning dataset used in GMAI-VL-R1 Data description GMAI-Reasoning10K is a high-quality medical image reasoning dataset containing 10,000 carefully selected samples. The data was collected from 95 medical datasets from reliable sources such as Kaggle, GrandChallenge, and Open-Release, covering 12 imaging modalities including X-ray, CT, and MRI. Data preprocessing followed the standardization methods from SAMed-20M: 3D data (CT/MRI) had… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/GMAI-Reasoning10K.imagevisual-question-answering10K<n<100K6 likes1.7k downloads1y agoHugging Face14poco28 /math-code-general-prunedtabular1M<n<10M0 likes1.3k downloads12d agoHugging Face15OpenOneRec /OpenOneRec-General-Pretrain 通用文本数据集 本目录包含 OpenOneRec 项目使用的通用文本数据集信息。这些数据集均来自 HuggingFace,经过清洗处理和对齐到项目统一的数据格式,并转换为 Parquet 格式用于训练。 数据格式说明 所有数据集均已转换为统一的 Parquet 格式,符合项目的数据格式规范(参考 ../README.md)。数据格式支持: Segments 格式:用于普通文本数据,使用 segments 字段存储文本段落列表 Chat 格式:用于对话数据,使用 messages 字段存储对话消息列表 每个 Parquet 文件包含以下核心字段: uuid: 唯一标识符 source: 数据来源标识 metadata: JSON 格式的元数据字典 segments 或 messages: 文本内容(根据数据类型选择) 详细的数据格式规范请参考 ../README.md。 数据集列表 数据集名称 样本数量 HuggingFace 仓库 reasoning_v1_20m 1,666… See the full description on the dataset page: https://huggingface.co/datasets/OpenOneRec/OpenOneRec-General-Pretrain.tabular1M<n<10M3 likes1.2k downloads9mo agoHugging Face16Scale-or-Reason /general-reasoning-ift-pairs Reasoning-IFT Pairs (General Domain) This dataset provides the largest set of IFT and Reasoning answers pairs for a set of general domain queries (cf: math-domain).It is based on the Infinity-Instruct dataset, an extensive and high-quality collection of instruction fine-tuning data. We curated 900k queries from the 7M_core subset of Infinity-Instruct, which covers multiple domains including general knowledge, commonsense Q&A, coding, and math.For each query… See the full description on the dataset page: https://huggingface.co/datasets/Scale-or-Reason/general-reasoning-ift-pairs.textquestion-answering1M<n<10M6 likes1k downloads4mo agoHugging Face17rooty2020 /HyperWorld-Bench-General0 likes947 downloads19d agoHugging Face18harryxi /HelpSteer3-general-code-Shift-Qwen-2.5-1.5B-Instruct-chat-formatted-generationstext1M<n<10M5 likes911 downloads1y agoHugging Face19Senqiao /VisionThink-General-Train VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning Senqiao/VisionThink-General-Train This is the training dataset used for our Reasoning VLM on general VQA tasks. VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning [Paper] Senqiao Yang, Junyi Li, Xin Lai, Bei Yu, Hengshuang Zhao, Jiaya Jia Highlights Our VisionThink leverages reinforcement learning to autonomously learn whether to… See the full description on the dataset page: https://huggingface.co/datasets/Senqiao/VisionThink-General-Train.image100K<n<1M3 likes903 downloads1y agoHugging Face20RJT1990 /GeneralThoughtArchive GeneralThought-430K Thought wants to be free Open reasoning data for March 14 2025. This dataset was part of a side-project in the weeks following the R1 release by Chengxi and Ross - we are no longer maintaining this dataset but are archiving it here. The dataset contains questions, reference answers, reasoning traces, final answers and other metadata from several popular reasoning models including DeepSeek-R1, DeepSeek-R1-Zero, OpenThoughts-32B, LIMO… See the full description on the dataset page: https://huggingface.co/datasets/RJT1990/GeneralThoughtArchive.tabular100K<n<1M79 likes897 downloads1y agoHugging Face21cjfcsjt /AITW_Generaltabular100K<n<1M2 likes883 downloads2y agoHugging Face22Valen-Team /Valen-Training-General-100k Valen-Training-General-100k GitHub · Preview model · General evaluation · Technical notes ✨ Introduction Valen brings visual perception to System One decision-making: text, images or video in, candidate probabilities out. This dataset provides 100,000 image-based decision records for supervised training and decision-head learning in Valen, spanning visual question answering, interfaces, games and documents. Each record contains one decision question… See the full description on the dataset page: https://huggingface.co/datasets/Valen-Team/Valen-Training-General-100k.image100K<n<1M2 likes864 downloads17d agoHugging Face23Brainquiver /general-master-en-202608 General · Master · English · 2026-08 English pretraining text, assembled from three public sources, cleaned with one character-level cleaner, and filtered for repetition. 109,337,531 documents and 468,064,046,462 characters. Composition Config Documents Characters What it is fineweb-edu-dedup 65,010,430 297,544,916,118 Web text an educational classifier kept cosmopedia-v2 38,591,146 144,011,993,012 Synthetic prose from a seeded generator… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-master-en-202608.tabulartext-generation100M<n<1B1 likes788 downloads1mo agoHugging Face24honghong3 /IMG_GENERALgated Usage This dataset is provided for research purposes. Redistribution, mirroring, or republishing of the dataset files is not permitted. The uploader does not claim ownership of the underlying content. Users are responsible for complying with applicable copyright and other rights. 0 likes747 downloads3d agoHugging Face25Shiki42 /piperx-general-pickup Piper X General Pickup 已发布 79 个资产、3753/4420 条成功 episode。 每个资产目标 52 条成功轨迹;成功率过低、尝试预算内未能采满的资产按实际成功条数发布。 采集状态(截至 2026-10-08) 采集已结束。计划 85 个资产、每个 52 条,共 4,420 条;实际情况: 已采集 79 个资产、3,753 条(约 85%):69 个采满 52 条,10 个不满额,按实际成功条数发布。 未采集 6 个资产:在尝试预算内 0 条成功,数据集中没有它们的 episode。 不满额资产:speaker/00000 42、trophy/00000 10、garage/00017 7、bell/00000 5、corkscrew/00000 9、bottle_opener/00000 2、cactus/00000 4、chocolate_bag/00000 36、pot/00001 37、wrench/00002 13。 未采集资产: 资产 尝试次数 情况… See the full description on the dataset page: https://huggingface.co/datasets/Shiki42/piperx-general-pickup.videon<1K0 likes655 downloads1d agoHugging Face26Shiki42 /piperx-general-pickup-clutter-preview Piper X General Pickup — Same-cluster preview ▶ Play episode 0 4 successful episodes, 694 frames, 25 FPS; three synchronized 640×480 RGB cameras. 这一版将目标和干扰物放在同一簇内,取消干扰物的前后分区以及目标周围 7.5 cm 的统一禁入区。 至少三个干扰物初始摆在目标 XY 包围盒附近,初始间距为 1.5–3 cm;其余物体沿同一簇扩展。 物理稳定后与录制开始时,均要求至少三个刚体干扰物到目标的 XY 包围盒间距不超过 6 cm。 这些距离为保守包围盒足迹代理,不能视为精确网格表面距离或遮挡指标。 本预览为近邻、不重叠的桌面布局;不包含有意堆叠。 The target is physically grasped near its live mesh centre, approached and descended to with open fingers, then closed upon only… See the full description on the dataset page: https://huggingface.co/datasets/Shiki42/piperx-general-pickup-clutter-preview.imagen<1K0 likes634 downloads4d agoHugging Face27Valen-Team /Valen-Eval-General-5k Valen-Eval-General-5k GitHub · 中文 README · Preview model · Technical notes ✨ Introduction Valen brings visual perception to System One decision-making: text, images or video in, candidate probabilities out. This dataset provides 5,000 image-based decision records for held-out evaluation, spanning visual question answering, interfaces, games and documents. Each record contains one decision question, a target probability distribution, local image… See the full description on the dataset page: https://huggingface.co/datasets/Valen-Team/Valen-Eval-General-5k.image1K<n<10K2 likes616 downloads17d agoHugging Face28Misalignment-Empirics /theo_impulsive-generality-bank-lineage Status: NOT the paper's results. Build lineage (archive, generated, generated_v2) of the impulsive generality eval banks: eval inputs, not results. The paper's canonical results are Misalignment-Empirics/theo_impulsive-qwen_2_5-7b-14b-32b-big_eval_results, run runs/impulsive-qwen_2_5-7b-14b-32b-20260925/. theo_impulsive-generality-bank-lineage 0 likes612 downloads6d agoHugging Face29ziqima /Objaverse-General-Find3DThis dataset contains benchmarks for open-world object part segmentation proposed by Find Any Part in 3D (ICCV 2025). This dataset includes two human-annotated benchmarks: Objaverse-General (of 100 object categories) and ShapeNetPart-Objaverse (of the same categories of ShapeNetPart, but with objects source from Objaverse to study distribution shift). Usage Inside both objaverse-general and objaverse-shapanetepart directories, the benchmark has the following directory structure:… See the full description on the dataset page: https://huggingface.co/datasets/ziqima/Objaverse-General-Find3D.0 likes527 downloads1y agoHugging Face30Mozilla /standard_chat_tool_calling_generaltextn<1K2 likes523 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.