Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01inclusionAI /ASearcher-Local-Knowledgetext10M<n<100M7 likes14k downloads1y agoHugging Face02DataPilot /Knowledge-QA-SingleTurn-Dataset Knowledge QA Single-turn Dataset(知識質問データセット・シングルターン) 概要 本データセットは、Aratako/Synthetic-JP-Conversations-Magpie-Nemotron-4-10k から質問を抽出し、DeepSeek V3.2で整形、Kimi K2.5で回答を生成した シングルターンの知識質問応答データセット です。Reasoning有効化により思考過程も最終データに含まれ、質問の難易度に応じてReasoning effortが動的に切り替わります。 生成にはSDG-LOOMという合成データ生成パイプラインを用いました。(sdg-loom) データの説明 項目 内容 件数 約7,000件 形式 JSONL(1行1JSON) 言語 日本語 ターン数 1ターン(質問1 + 回答1) ソースデータセット… See the full description on the dataset page: https://huggingface.co/datasets/DataPilot/Knowledge-QA-SingleTurn-Dataset.text1K<n<10K2 likes4k downloads7mo agoHugging Face03onegoai /onego-knowledge-packs ONEGO Knowledge Packs Offline RAG databases for ONEGO / Offline AI Assistant. Files File Role Size SHA256 wikipedia_base.ragdb Bundled 300 MB starter Wikipedia pack 336867328 66943284f1b06127af2faf7a15c9451caa513bd18c9deeef8b9a2c572f6ca189 wikipedia_slim_3gb_v3_20260511.ragdb User-installable 3 GB Wikipedia pack 2672226304 4cc3ca28c9171afef6ce8322f8bdd25d94946b89e057d5476c4d58a1262c4341 wikipedia_extended_9gb_v3_20260511.ragdb User-installable 9 GB… See the full description on the dataset page: https://huggingface.co/datasets/onegoai/onego-knowledge-packs.textn<1K0 likes740 downloads5mo agoHugging Face04MuskumPillerum /General-Knowledge Dataset Card for Dataset Name Dataset Summary The dataset is a collection of questions and answers themed on general facts and reasoning. The dataset is divided into two features - 'Question' and 'Answer'. It is meant to be used for training a model to be good at general knowledge and reasoning. This dataset is inspired from the Alpaca dataset, and infact contains a subset of the alpaca dataset in itself. Distribution The distribution of the… See the full description on the dataset page: https://huggingface.co/datasets/MuskumPillerum/General-Knowledge.texttext-classification10K<n<100K51 likes467 downloads10mo agoHugging Face05FreedomIntelligence /huatuo_knowledge_graph_qa Dataset Card for Huatuo_knowledge_graph_qa Dataset Summary We built this QA dataset based on the medical knowledge map, with a total of 798,444 pieces of data, in which the questions are constructed by means of templates, and the answers are the contents of the entries in the knowledge map. Dataset Creation Source Data… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/huatuo_knowledge_graph_qa.texttext-generation100K<n<1M52 likes372 downloads3y agoHugging Face06chembricks /chemistry-knowledge ChemBricks Knowledge Does caffeine prefer water or an oil-like liquid?Why can adding one small group change a molecule's behavior?Can we design a molecule that interacts more favorably with water while meeting other constraints?How much energy does it take to remove an electron from a molecule? These are the kinds of questions behind this dataset. Each investigation connects a question to recorded calculations, an answer, and the evidence needed to examine that answer. Created… See the full description on the dataset page: https://huggingface.co/datasets/chembricks/chemistry-knowledge.tabularquestion-answering10K<n<100K1 likes357 downloads24d agoHugging Face07luozhouyang /kgclue-knowledge KgCLUE-Knowledge The original data is from CLUEbenchmark/KgCLUE. Here is a JSON version of the original knowledge base. Usage from datasets import load_dataset dataset = load_dataset("luozhouyang/kgclue-knowledge") # or select files dataset = load_dataset("luozhouyang/kgclue-knowledge", data_files=["kgclue.knowledge00.jsonl"]) text10M<n<100M2 likes309 downloads5y agoHugging Face08gjata-legacy /ess-mai-poc-006-living-negative-knowledge-verified-reentry ESS-MAI POC 006 — Living Negative Knowledge: Verified Re-entry ESS-MAI is experimental, governance-first AI systems research by Bledar Gjata (Gjata Legacy), developed in Tirana, Albania. Albania denotes where the research is developed; ESS-MAI is not an Albanian-language model and not a national or sovereign AI system. Its relevance to AI governance is, in the repository's words, “an invitation to evaluate the project, not a claim of academic validation, regulatory compliance… See the full description on the dataset page: https://huggingface.co/datasets/gjata-legacy/ess-mai-poc-006-living-negative-knowledge-verified-reentry.textn<1K0 likes294 downloads4d agoHugging Face09AdaptLLM /med_knowledge_prob Adapting LLMs to Domains via Continual Pre-Training (ICLR 2024) This repo contains the Biomedicine Knowledge Probing dataset used in our paper Adapting Large Language Models via Reading Comprehension. We explore continued pre-training on domain-specific corpora for large language models. While this approach enriches LLMs with domain knowledge, it significantly hurts their prompting ability for question answering. Inspired by human learning via reading comprehension, we propose a… See the full description on the dataset page: https://huggingface.co/datasets/AdaptLLM/med_knowledge_prob.texttext-classification10K<n<100K12 likes240 downloads2y agoHugging Face10nvidia /Nemotron-RL-knowledge-web_search-mcqa Dataset Description: The Nemotron-RL-knowledge-web_search-mcqa is a multi-domain synthetic dataset designed to improve science and general reasoning in large language models (LLMs). It is a filtered subset of the OpenScienceReasoning-2 dataset and contains multiple-choice question–answer pairs spanning diverse domains: physics, biology, mathematics, humanities, computer science, engineering, chemistry, and others. This dataset is released as part of NVIDIA NeMo Gym, a framework… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-knowledge-web_search-mcqa.text1K<n<10K17 likes234 downloads17d agoHugging Face11snuh /specialist-level_medical_knowledge_dataset_sft specialist-level_medical_knowledge_dataset_sft Dataset Summary specialist-level_medical_knowledge_dataset_sft is an integrated collection of augmented SFT data across 13 distinct medical domains, developed by the Healthcare AI Research Institute (HARI) at SNUH. This dataset is derived and augmented from the Specialized Medical Knowledge Data (전문 의학지식 데이터) provided by AI-Hub. It focuses exclusively on complex clinical scenarios generated using the "Add Constraints"… See the full description on the dataset page: https://huggingface.co/datasets/snuh/specialist-level_medical_knowledge_dataset_sft.textquestion-answering10K<n<100K2 likes232 downloads2mo agoHugging Face12wuwu616 /huatuo_knowledge_graph_qa Dataset Card for Huatuo_knowledge_graph_qa Dataset Summary We built this QA dataset based on the medical knowledge map, with a total of 798,444 pieces of data, in which the questions are constructed by means of templates, and the answers are the contents of the entries in the knowledge map. Dataset Creation Source Data… See the full description on the dataset page: https://huggingface.co/datasets/wuwu616/huatuo_knowledge_graph_qa.texttext-generation100K<n<1M0 likes222 downloads23d agoHugging Face13whfeLingYu /Misleading_KnowledgeMisleading_Knowledge Misleading_Knowledge is the misleading-knowledge corpus introduced in “Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions.” It is designed for controlled research on factual robustness, evidence verification, source cues, and false-conclusion adoption in Deep Research agents. Paper: https://arxiv.org/abs/2607.20891 Code: https://github.com/whfeLingYu/MisKnow-Agent Dataset repository: https://huggingface.co/datasets/whfeLingYu/Misleading_Knowledge… See the full description on the dataset page: https://huggingface.co/datasets/whfeLingYu/Misleading_Knowledge.tabular1K<n<10K0 likes179 downloads2mo agoHugging Face14TonicAI /knowledge-worker-search-bench Knowledge-Worker Search Bench 40 multi-channel retrieval tasks over realistic synthetic knowledge-worker environments, generated with Tonic Fabricate. Each task drops an agent into one persona's work world — mail (Outlook or Gmail), Slack, Google Docs, calendar, attachments — and asks a question a real chief-of-staff-style assistant would get: "brief me for tomorrow's sync", "where did we land on the renewal, and what forced the timeline?". Answering requires finding and… See the full description on the dataset page: https://huggingface.co/datasets/TonicAI/knowledge-worker-search-bench.textn<1K1 likes169 downloads2mo agoHugging Face15Severian /Internal-Knowledge-Map Internal Knowledge Map: Experiments in Deeper Understanding and Novel Thinking for LLMs Designed for Cross-Discipline/Interconnected Critical Thinking, Nuanced Understanding, Diverse Role Playing and Innovative Problem Solving By integrating a cohesively structured dataset emphasizing the interconnectedness of knowledge across a myriad of domains, exploring characters/role playing/community discourse, solving impossible problems and developing inner dialogues; this project aspires… See the full description on the dataset page: https://huggingface.co/datasets/Severian/Internal-Knowledge-Map.text1K<n<10K49 likes153 downloads3y agoHugging Face16Yxanul /Mephisto-Knowledge_538k Mephisto-Knowledge_538k 538,861 English knowledge SFT examples generated by Qwen/Qwen3.5-4B in non-thinking (Instruct) mode on the Knowledge prompts of openbmb/UltraData-SFT-2605. Responses contain no chain-of-thought — thinking was disabled at generation time, so each assistant turn is a direct answer, usually with a short justification. Companion dataset: Mephisto-IF_172k (instruction-following, same teacher and pipeline). Read this before training: ref_agrees… See the full description on the dataset page: https://huggingface.co/datasets/Yxanul/Mephisto-Knowledge_538k.textquestion-answering100K<n<1M2 likes142 downloads2mo agoHugging Face17shimo4228 /agent-knowledge-cycle Agent Knowledge Cycle (AKC) — Knowledge Graph JSON-LD knowledge graph encoding the concept layer of the Agent Knowledge Cycle (AKC) — a six-phase bidirectional growth loop in which agent behavior and the operator's judgment co-develop over time, sustaining intent alignment that tests cannot check on their own. What this dataset is This dataset is a mirror of the graph.jsonld file at the root of the AKC GitHub repository. It is provided here for LLM training… See the full description on the dataset page: https://huggingface.co/datasets/shimo4228/agent-knowledge-cycle.tabularn<1K1 likes139 downloads2d agoHugging Face18ego0op /earth-love-united-climate-knowledge 🌍 Earth Love United Climate Knowledge Dataset The most comprehensive open climate science knowledge dataset. 10,128 text chunks + 124 structured facts + 4.54B year geological memory + 10 tipping points. Built to power GAIA — an AI that embodies the living consciousness of Earth. Dataset Overview This dataset gives an AI system authoritative, sourced knowledge about climate change, carbon, Earth science, and solutions. It has four layers: Layer 1: Text Knowledge… See the full description on the dataset page: https://huggingface.co/datasets/ego0op/earth-love-united-climate-knowledge.tabulartext-retrieval10K<n<100K2 likes134 downloads5mo agoHugging Face19fatcat55 /delvantic-stock-knowledge-layer Delvantic Stock Knowledge Layer A 872k-word, source-cited textbook of stock analysis and trading, organized as a tree — the reference layer behind a live AI research engine, published in full. Every finance dataset on the Hub is numbers: prices, filings, labelled headlines. This is the missing other half — the explanations. 771 documents on how the machinery of markets actually works, from reading a cash-flow statement to why volatility regimes break strategies, each one written… See the full description on the dataset page: https://huggingface.co/datasets/fatcat55/delvantic-stock-knowledge-layer.tabulartext-retrieval1K<n<10K0 likes131 downloads2mo agoHugging Face20sosa123454321 /greenpars-knowledge GreenPars knowledge base (RAG) Trilingual reports (FA/EN/TR) of the GreenPars plastic-recycling business plan: business plan, financial model, funding, scopes, Turkish partners, sanctions critique, WtE & sponsor assessment. chunks.json = 130 pre-chunked passages used by the GreenPars AI assistant. textquestion-answeringn<1K0 likes130 downloads12d agoHugging Face21dataformer /self-knowledgetextn<1K1 likes128 downloads2y agoHugging Face22HmyHxy /finance-Knowledge-Credit-Chinesetextn<1K4 likes113 downloads2y agoHugging Face23xihao1 /Traditional-Chinese-Medicine-Knowledgetext10K<n<100K1 likes108 downloads1y agoHugging Face24nemiling-official /nemiling-knowledge-base Nemiling Knowledge Base Nemiling Knowledge Base is the official structured knowledge dataset about Nemiling. Nemiling is a Russian platform for automating the monetization of Telegram projects through paid subscriptions, paid messages, paid consultations, and donations. The platform can be used for projects with Russian and international audiences. The dataset is maintained by the official Nemiling organization and provides structured, machine-readable information about the… See the full description on the dataset page: https://huggingface.co/datasets/nemiling-official/nemiling-knowledge-base.tabularquestion-answeringn<1K0 likes88 downloads2mo agoHugging Face25trumancai /talkie-1930-knowledge-bench Talkie-1930 Agentic Knowledge Injection Benchmark Benchmark for measuring whether an autonomous agent can durably write "verifiable post-1930 knowledge" into the parameters of a base language model (talkie-1930), evaluated standalone (no retrieval, no in-context). Because the talkie-1930 base is contamination-free for post-1930 facts, any gain on certified-novel targets is true injection, not elicitation of pre-existing knowledge — the headline property this benchmark gives you.… See the full description on the dataset page: https://huggingface.co/datasets/trumancai/talkie-1930-knowledge-bench.textquestion-answering10K<n<100K0 likes87 downloads4mo agoHugging Face26Query-of-CC /Knowledge_PileKnowledge Pile is a knowledge-related data leveraging Query of CC. This dataset is a partial of Knowledge Pile(about 40GB disk size), full datasets have been released in [🤗 knowledge_pile_full], a total of 735GB disk size and 188B tokens (using Llama2 tokenizer). Query of CC Just like the figure below, we initially collected seed information in some specific domains, such as keywords, frequently asked questions, and textbooks, to serve as inputs for the Query Bootstrapping stage.… See the full description on the dataset page: https://huggingface.co/datasets/Query-of-CC/Knowledge_Pile.text1M<n<10M22 likes86 downloads3y agoHugging Face27astroBench /knowledge_application!!!当前数据集仅为了方便测试使用,不保证题目答案正确!!! !!!如想用于科学研究,请留意后续正式发布!!! textn<1K0 likes80 downloads2y agoHugging Face28Uunan /turkish-knowledge-sft Turkish Knowledge SFT Turkish Knowledge SFT is a large-scale synthetic instruction-following dataset designed to improve the factual knowledge, explanation quality, and instructional capabilities of Turkish Large Language Models (LLMs). The dataset is designed for Supervised Fine-Tuning (SFT) and follows a conversation-oriented format compatible with modern chat models. Features 🇹🇷 Entirely in Turkish 🤖 Synthetic instruction-following dataset 📚… See the full description on the dataset page: https://huggingface.co/datasets/Uunan/turkish-knowledge-sft.texttext-generation100K<n<1M0 likes80 downloads3mo agoHugging Face29snuh /essential-level_medical_knowledge_dataset_sft essential-level_medical_knowledge_dataset_sft Dataset Summary essential-level_medical_knowledge_dataset_sft is an integrated collection of augmented SFT data across 4 distinct medical domains, developed by the Healthcare AI Research Institute (HARI) at SNUH. This dataset is derived and augmented from the Essential Medical Knowledge Data (필수의료 의학지식 데이터) provided by AI-Hub. It focuses exclusively on complex clinical scenarios generated using the "Add Constraints"… See the full description on the dataset page: https://huggingface.co/datasets/snuh/essential-level_medical_knowledge_dataset_sft.textquestion-answering10K<n<100K0 likes80 downloads20d agoHugging Face30intertwine-expel /knowledge-base Expel Knowledge Base Articles textn<1K0 likes76 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.