Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hatemestinbejaia /ExperimentDATA_knowledge_distillation_vs_fine_tuningtabular100M<n<1B2 likes44k downloads9mo agoHugging Face02inclusionAI /ASearcher-Local-Knowledgetext10M<n<100M7 likes14k downloads1y agoHugging Face03abuzreq /model-bending-knowledge-base Model Bending Knowledge Base This dataset records what happens when you bend the inside of a diffusion model. Bending means multiplying, rotating, adding noise to or otherwise changing the activations of a layer while the model generates. Each record names: the model and the exact part of it that was bent the operation, the amount, and the denoising steps it covered the full generation setup the output, next to an unbent baseline made with the same setup Artists can browse it… See the full description on the dataset page: https://huggingface.co/datasets/abuzreq/model-bending-knowledge-base.image10K<n<100K0 likes13k downloads4d agoHugging Face04BByrneLab /multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR PreFLMR M2KR Dataset Card Dataset details Dataset type: M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models. We pre-process the datasets into a uniform format and write several task-specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks:… See the full description on the dataset page: https://huggingface.co/datasets/BByrneLab/multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR.tabular10M<n<100M10 likes4.2k downloads1y agoHugging Face05DataPilot /Knowledge-QA-SingleTurn-Dataset Knowledge QA Single-turn Dataset(知識質問データセット・シングルターン) 概要 本データセットは、Aratako/Synthetic-JP-Conversations-Magpie-Nemotron-4-10k から質問を抽出し、DeepSeek V3.2で整形、Kimi K2.5で回答を生成した シングルターンの知識質問応答データセット です。Reasoning有効化により思考過程も最終データに含まれ、質問の難易度に応じてReasoning effortが動的に切り替わります。 生成にはSDG-LOOMという合成データ生成パイプラインを用いました。(sdg-loom) データの説明 項目 内容 件数 約7,000件 形式 JSONL(1行1JSON) 言語 日本語 ターン数 1ターン(質問1 + 回答1) ソースデータセット… See the full description on the dataset page: https://huggingface.co/datasets/DataPilot/Knowledge-QA-SingleTurn-Dataset.text1K<n<10K2 likes4k downloads7mo agoHugging Face06Bendyline /wikipedia-knowledge Wikipedia Knowledge Catalogs 20 gezk knowledge catalogs: portable, read-only reference corpora with full-text and vector indexes, each packaged as one .gezk file that Gezel and any gezk reader can search and cite offline. Every catalog is here twice — the archive, and the same content as Parquet tables for data tools. Catalogs Catalog Id Version Documents Chunks Archive Wikipedia: Arts & Architecture wikipedia-arts 2026.4.5 16,485 73,383 179.0 MB… See the full description on the dataset page: https://huggingface.co/datasets/Bendyline/wikipedia-knowledge.tabulartext-retrieval10M<n<100M0 likes2.2k downloads5d agoHugging Face07nvidia /Nemotron-RL-knowledge-mcqa Dataset Description: The Nemotron-RL-knowledge-mcqa is a multi-domain synthetic multiple-choice question-answering (MCQA) dataset containing knowledge based questions. It combines and refines subsets of the [OpenScienceReasoning-2] (https://huggingface.co/datasets/nvidia/OpenScienceReasoning-2) dataset and other unstructured sources such as books and articles.The dataset was created using Qwen3-32B, [Qwen3-235B-A22B-Instruct-2507]… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-knowledge-mcqa.text100K<n<1M13 likes2.1k downloads12d agoHugging Face08MatinaAI /peka_persian_knowledge_assessmentgated PeKA (Persian Knowledge Assessment) PeKA is a dataset introduced in the paper "Advancing Persian LLM Evaluation", accepted at NAACL 2025 findings. It was developed as part of a broader effort to evaluate and benchmark large language models (LLMs) for multiple Persian knowledge topics. For comprehensive details regarding the dataset’s construction, scope, task, and intended use, please refer to the original paper. This dataset is constructed so that answering these questions… See the full description on the dataset page: https://huggingface.co/datasets/MatinaAI/peka_persian_knowledge_assessment.tabularquestion-answering1K<n<10K3 likes1.5k downloads1y agoHugging Face09LanternCX /zhiya-knowledge 知芽教学知识库 这是一个面向教学问答的多模态知识库数据集,包含教材、编程与人工智能教学资料,以及资料中的文字、图片和少量视频。除了原始资料,还提供整理后的文本片段和已经生成的检索向量(Embedding)。 它用于开源项目 知芽 Zhiya。知芽是面向小学至高中学生的个性化 AI 学习搭子,帮助学生学习编程和人工智能。 这个知识库让知芽在回答问题、组织教学内容时,先查找相关资料,再根据资料讲解,并在对话中给出来源引用。开发者可以下载完整知识库,接入自己的知芽部署,无需重新计算这些资料的向量。 有哪些资料? 资料来自 Day of AI、Microsoft AI for Beginners、AI4K12、Datawhale 和 K12 textbook,共整理出 7,434 份文档。内容包括教材、课程讲义、幻灯片、编程示例和教学图片。 内容 位置 原始资料、提取后的文字、图片、视频及来源信息 knowledge/ 已生成的 99,630 条文本检索向量… See the full description on the dataset page: https://huggingface.co/datasets/LanternCX/zhiya-knowledge.image100K<n<1M0 likes1.2k downloads4d agoHugging Face10financeindustryknowledgeskills /modeling_valuation_knowledge Finance Training Data Repository A curated collection of financial modeling courses, materials, and resources designed to serve as training data for building a finance industry knowledge base. Repository Structure Finance_Training_Data/ ├── 01_Financial_Statement_Modeling/ # 3-statement modeling fundamentals ├── 02_DCF_Modeling/ # Discounted cash flow valuation ├── 03_Trading_Comps/ # Comparable company analysis ├──… See the full description on the dataset page: https://huggingface.co/datasets/financeindustryknowledgeskills/modeling_valuation_knowledge.documentn<1K0 likes1.1k downloads4mo agoHugging Face11knowledge-computing /FRIEDA FRIEDA is a multimodal benchmark for open-ended cartographic reasoning over real-world map images.Each example pairs reference maps (and optional contextual maps) with a natural-language question and a reference answer. The benchmark targets common GIS relation types (i.e., topological, metric, directional) and includes questions that require multi-step reasoning and cross-map grounding. Dataset Summary Modality: image + text # Examples: 500 Input: map image(s) + question… See the full description on the dataset page: https://huggingface.co/datasets/knowledge-computing/FRIEDA.imagevisual-question-answeringn<1K2 likes957 downloads9mo agoHugging Face12matlok /python-image-copilot-training-using-import-knowledge-graphs Python Copilot Image Training using Import Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each row contains a png file in the dbytes column. Rows: 216642 Size: 211.2 GB Data type: png Format: Knowledge graph using NetworkX with alpaca text box Schema The png is in the dbytes column: { "dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-import-knowledge-graphs.tabulartext-to-imagen<1K0 likes781 downloads3y agoHugging Face13onegoai /onego-knowledge-packs ONEGO Knowledge Packs Offline RAG databases for ONEGO / Offline AI Assistant. Files File Role Size SHA256 wikipedia_base.ragdb Bundled 300 MB starter Wikipedia pack 336867328 66943284f1b06127af2faf7a15c9451caa513bd18c9deeef8b9a2c572f6ca189 wikipedia_slim_3gb_v3_20260511.ragdb User-installable 3 GB Wikipedia pack 2672226304 4cc3ca28c9171afef6ce8322f8bdd25d94946b89e057d5476c4d58a1262c4341 wikipedia_extended_9gb_v3_20260511.ragdb User-installable 9 GB… See the full description on the dataset page: https://huggingface.co/datasets/onegoai/onego-knowledge-packs.textn<1K0 likes740 downloads5mo agoHugging Face14habedi /kaggle-knowledge-graph Kaggle Knowledge Graph A knowledge graph built from Kaggle's public Meta Kaggle (and Meta Kaggle Code datasets). It links competitions, teams, submissions, users, notebooks, datasets, discussion forums, tags, organizations, and notebook code invocations. See the project repository for build scripts and documentation on the graph schema, data model, and usage examples. Release Field Value Version 2026-09-30 Meta Kaggle snapshot 2026-09-30 Build code… See the full description on the dataset page: https://huggingface.co/datasets/habedi/kaggle-knowledge-graph.tabular1M<n<10M0 likes721 downloads10d agoHugging Face15matlok /python-image-copilot-training-using-class-knowledge-graphs Python Copilot Image Training using Class Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each row contains a png file in the dbytes column. Rows: 312277 Size: 304.3 GB Data type: png Format: Knowledge graph using NetworkX with alpaca text box Schema The png is in the dbytes column: { "dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-class-knowledge-graphs.tabulartext-to-imagen<1K0 likes676 downloads3y agoHugging Face16MRMRbenchmark /knowledge Evaluation Code The evaluation code is implemented based on MTEB framework and avaliable in https://github.com/rebeccaz4/MRMR. Disclaimers The guidelines for the annotators emphasized strict compliance with copyright and licensing rules from the initial data source, specifically avoiding materials from websites that forbid copying and redistribution. Should you encounter any data samples potentially breaching the copyright or licensing regulations of any site, we… See the full description on the dataset page: https://huggingface.co/datasets/MRMRbenchmark/knowledge.image10K<n<100K0 likes663 downloads10mo agoHugging Face17matlok /python-audio-copilot-training-using-function-knowledge-graphs Python Copilot Audio Training using Global Functions with Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each global function has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the associated source code file_path… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-function-knowledge-graphs.tabulartext-to-audion<1K1 likes613 downloads3y agoHugging Face18BByrneLab /multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR_CN PreFLMR M2KR Dataset Card Dataset details Dataset type: M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models. We pre-process the datasets into a uniform format and write several task-specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks:… See the full description on the dataset page: https://huggingface.co/datasets/BByrneLab/multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR_CN.tabular1M<n<10M0 likes477 downloads2y agoHugging Face19knowledge-in-visual-synthesis /v1 Knowledge in Visual Synthesis This dataset contains prompt–image examples for evaluating and studying knowledge-intensive visual synthesis. Samples are organized by contributor as dataset subsets (configs), with each upload version exposed as a split. Dataset structure Subset Splits byx v1, v2 yuner v1, v2, v3 zanyi v1, v2, v3 jiayu v1, v2, v3 sherry v1, v2 yujunz v1 The byx/v1 split contains 140 unique prompts and 300 generated images.… See the full description on the dataset page: https://huggingface.co/datasets/knowledge-in-visual-synthesis/v1.image1K<n<10K0 likes470 downloads4d agoHugging Face20MuskumPillerum /General-Knowledge Dataset Card for Dataset Name Dataset Summary The dataset is a collection of questions and answers themed on general facts and reasoning. The dataset is divided into two features - 'Question' and 'Answer'. It is meant to be used for training a model to be good at general knowledge and reasoning. This dataset is inspired from the Alpaca dataset, and infact contains a subset of the alpaca dataset in itself. Distribution The distribution of the… See the full description on the dataset page: https://huggingface.co/datasets/MuskumPillerum/General-Knowledge.texttext-classification10K<n<100K51 likes467 downloads10mo agoHugging Face21matlok /python-audio-copilot-training-using-class-knowledge-graphs Python Copilot Audio Training using Class with Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each class method has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the associated source code file_path identifier. Rows:… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-class-knowledge-graphs.tabulartext-to-audion<1K0 likes429 downloads3y agoHugging Face22houlab /arsma-knowledge-db ARSMA-web biomedical knowledge graph (arsma-knowledge-db) Identifiers and links from Rfam, ClinVar, Reactome, MONDO, HGNC and miRBase, joined to the motif classes of ARSMA-web. 225,183 nodes and 291,432 edges. file one row is notes nodes.parquet one node type, name, source edges.parquet one edge type, source, evidence, specificity build_report.json — what each source contributed Motifs are linked at the level of motif class, not occurrence. An Rfam family is… See the full description on the dataset page: https://huggingface.co/datasets/houlab/arsma-knowledge-db.text100K<n<1M0 likes423 downloads9d agoHugging Face23houlab /motif-knowledge-db ARSMA-web sequence-relation graph (motif-knowledge-db) A graph centered on motif sequence groups, built from ARSMA-web's own stores: where each sequence occurs, which RNA family it is in, what binds it and which sequences are variants of it. file one row is notes nodes.parquet one node type: seq_group, pdb_entry, ligand, rna_family, modification, motif_type; sequence-group nodes carry occurrence and contact counts edges.parquet one edge type: occurs_in, in_family… See the full description on the dataset page: https://huggingface.co/datasets/houlab/motif-knowledge-db.text100K<n<1M0 likes407 downloads9d agoHugging Face24Emulated-Inc /multilingual-knowledge-training-pool Multilingual knowledge training pool Public multiple-choice questions in many languages from seven datasets, read at the pinned revisions named below and laid out twice. Train on either layer or on both. pool.jsonl Every source rewritten into one shape, 328665 rows in 116 languages, one JSON object per line, with these fields. Field What it holds id a row identifier unique within this file question the question text, as its source publishes it… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/multilingual-knowledge-training-pool.textquestion-answering1K<n<10K0 likes377 downloads28d agoHugging Face25hoanganhpham /HealthBench_Knowledge_Question_250Ktext100K<n<1M0 likes374 downloads1y agoHugging Face26NoemaAI-labs /knowledge-packs Noema Knowledge Packs Pre-cleaned public-domain reference text for Noema's offline Knowledge Packs feature. All content is U.S.-government public domain (17 U.S.C. §105) or CC0. Each <pack>/<file>.txt is downloaded by the app and embedded on-device. Sources: U.S. Army FM 21-76 / FM 3-25.26 / FM 4-25.11 (via Internet Archive), Ready.gov, and the CIA World Factbook (factbook.json). text100K<n<1M0 likes374 downloads4mo agoHugging Face27FreedomIntelligence /huatuo_knowledge_graph_qa Dataset Card for Huatuo_knowledge_graph_qa Dataset Summary We built this QA dataset based on the medical knowledge map, with a total of 798,444 pieces of data, in which the questions are constructed by means of templates, and the answers are the contents of the entries in the knowledge map. Dataset Creation Source Data… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/huatuo_knowledge_graph_qa.texttext-generation100K<n<1M52 likes372 downloads3y agoHugging Face28matlok /python-audio-copilot-training-using-import-knowledge-graphs Python Copilot Audio Training using Imports with Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each imported module for each unique class in each module file has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-import-knowledge-graphs.tabulartext-to-audion<1K0 likes370 downloads3y agoHugging Face29chembricks /chemistry-knowledge ChemBricks Knowledge Does caffeine prefer water or an oil-like liquid?Why can adding one small group change a molecule's behavior?Can we design a molecule that interacts more favorably with water while meeting other constraints?How much energy does it take to remove an electron from a molecule? These are the kinds of questions behind this dataset. Each investigation connects a question to recorded calculations, an answer, and the evidence needed to examine that answer. Created… See the full description on the dataset page: https://huggingface.co/datasets/chembricks/chemistry-knowledge.tabularquestion-answering10K<n<100K1 likes357 downloads24d agoHugging Face30matlok /python-image-copilot-training-using-function-knowledge-graphs Python Copilot Image Training using Function Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each row contains a png file in the dbytes column. Rows: 134357 Size: 130.5 GB Data type: png Format: Knowledge graph using NetworkX with alpaca text box Schema The png is in the dbytes column: { "dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-function-knowledge-graphs.tabulartext-to-imagen<1K0 likes355 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.