Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hatemestinbejaia /ExperimentDATA_knowledge_distillation_vs_fine_tuningtabular100M<n<1B2 likes44k downloads9mo agoHugging Face02BByrneLab /multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR PreFLMR M2KR Dataset Card Dataset details Dataset type: M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models. We pre-process the datasets into a uniform format and write several task-specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks:… See the full description on the dataset page: https://huggingface.co/datasets/BByrneLab/multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR.tabular10M<n<100M10 likes4.2k downloads1y agoHugging Face03Bendyline /wikipedia-knowledge Wikipedia Knowledge Catalogs 20 gezk knowledge catalogs: portable, read-only reference corpora with full-text and vector indexes, each packaged as one .gezk file that Gezel and any gezk reader can search and cite offline. Every catalog is here twice — the archive, and the same content as Parquet tables for data tools. Catalogs Catalog Id Version Documents Chunks Archive Wikipedia: Arts & Architecture wikipedia-arts 2026.4.5 16,485 73,383 179.0 MB… See the full description on the dataset page: https://huggingface.co/datasets/Bendyline/wikipedia-knowledge.tabulartext-retrieval10M<n<100M0 likes2.2k downloads4d agoHugging Face04MatinaAI /peka_persian_knowledge_assessmentgated PeKA (Persian Knowledge Assessment) PeKA is a dataset introduced in the paper "Advancing Persian LLM Evaluation", accepted at NAACL 2025 findings. It was developed as part of a broader effort to evaluate and benchmark large language models (LLMs) for multiple Persian knowledge topics. For comprehensive details regarding the dataset’s construction, scope, task, and intended use, please refer to the original paper. This dataset is constructed so that answering these questions… See the full description on the dataset page: https://huggingface.co/datasets/MatinaAI/peka_persian_knowledge_assessment.tabularquestion-answering1K<n<10K3 likes1.5k downloads1y agoHugging Face05matlok /python-image-copilot-training-using-import-knowledge-graphs Python Copilot Image Training using Import Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each row contains a png file in the dbytes column. Rows: 216642 Size: 211.2 GB Data type: png Format: Knowledge graph using NetworkX with alpaca text box Schema The png is in the dbytes column: { "dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-import-knowledge-graphs.tabulartext-to-imagen<1K0 likes781 downloads3y agoHugging Face06habedi /kaggle-knowledge-graph Kaggle Knowledge Graph A knowledge graph built from Kaggle's public Meta Kaggle (and Meta Kaggle Code datasets). It links competitions, teams, submissions, users, notebooks, datasets, discussion forums, tags, organizations, and notebook code invocations. See the project repository for build scripts and documentation on the graph schema, data model, and usage examples. Release Field Value Version 2026-09-30 Meta Kaggle snapshot 2026-09-30 Build code… See the full description on the dataset page: https://huggingface.co/datasets/habedi/kaggle-knowledge-graph.tabular1M<n<10M0 likes721 downloads10d agoHugging Face07matlok /python-image-copilot-training-using-class-knowledge-graphs Python Copilot Image Training using Class Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each row contains a png file in the dbytes column. Rows: 312277 Size: 304.3 GB Data type: png Format: Knowledge graph using NetworkX with alpaca text box Schema The png is in the dbytes column: { "dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-class-knowledge-graphs.tabulartext-to-imagen<1K0 likes676 downloads3y agoHugging Face08matlok /python-audio-copilot-training-using-function-knowledge-graphs Python Copilot Audio Training using Global Functions with Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each global function has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the associated source code file_path… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-function-knowledge-graphs.tabulartext-to-audion<1K1 likes613 downloads3y agoHugging Face09BByrneLab /multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR_CN PreFLMR M2KR Dataset Card Dataset details Dataset type: M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models. We pre-process the datasets into a uniform format and write several task-specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks:… See the full description on the dataset page: https://huggingface.co/datasets/BByrneLab/multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR_CN.tabular1M<n<10M0 likes477 downloads2y agoHugging Face10matlok /python-audio-copilot-training-using-class-knowledge-graphs Python Copilot Audio Training using Class with Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each class method has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the associated source code file_path identifier. Rows:… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-class-knowledge-graphs.tabulartext-to-audion<1K0 likes429 downloads3y agoHugging Face11matlok /python-audio-copilot-training-using-import-knowledge-graphs Python Copilot Audio Training using Imports with Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each imported module for each unique class in each module file has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-import-knowledge-graphs.tabulartext-to-audion<1K0 likes370 downloads3y agoHugging Face12chembricks /chemistry-knowledge ChemBricks Knowledge Does caffeine prefer water or an oil-like liquid?Why can adding one small group change a molecule's behavior?Can we design a molecule that interacts more favorably with water while meeting other constraints?How much energy does it take to remove an electron from a molecule? These are the kinds of questions behind this dataset. Each investigation connects a question to recorded calculations, an answer, and the evidence needed to examine that answer. Created… See the full description on the dataset page: https://huggingface.co/datasets/chembricks/chemistry-knowledge.tabularquestion-answering10K<n<100K1 likes357 downloads23d agoHugging Face13matlok /python-image-copilot-training-using-function-knowledge-graphs Python Copilot Image Training using Function Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each row contains a png file in the dbytes column. Rows: 134357 Size: 130.5 GB Data type: png Format: Knowledge graph using NetworkX with alpaca text box Schema The png is in the dbytes column: { "dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-function-knowledge-graphs.tabulartext-to-imagen<1K0 likes355 downloads3y agoHugging Face14matlok /python-image-copilot-training-using-class-knowledge-graphs-2024-01-27 Python Copilot Image Training using Class Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each row contains a png file in the dbytes column. Rows: 312836 Size: 294.1 GB Data type: png Format: Knowledge graph using NetworkX with alpaca text box Schema The png is in the dbytes column: { "dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-class-knowledge-graphs-2024-01-27.tabulartext-to-imagen<1K0 likes337 downloads3y agoHugging Face15matlok /python-image-copilot-training-using-inheritance-knowledge-graphs Python Copilot Image Training using Inheritance and Polymorphism Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each row contains a png file in the dbytes column. Rows: 259017 Size: 135.2 GB Data type: png Format: Knowledge graph using NetworkX with alpaca text box Schema The png is in the dbytes column: {… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-inheritance-knowledge-graphs.tabulartext-to-imagen<1K0 likes303 downloads3y agoHugging Face16matlok /python-audio-copilot-training-using-inheritance-knowledge-graphs Python Copilot Audio Training using Inheritance and Polymorphism Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each base class for each unique class in each module file has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-inheritance-knowledge-graphs.tabulartext-to-audion<1K0 likes261 downloads3y agoHugging Face17dwright37 /llm-knowledge-collapse "Epistemic Diversity and Knowledge Collapse in Large Language Models" (Wright et al. 2025)     Authors: Dustin Wright, Sarah Masud, Jared Moore, Srishti Yadav, Maria Antoniak, Peter Ebert Christiensen, Chan Young Park, and Isabelle Augenstein Contains all 1.6M responses and 70M claims used to measure LLM epistemic diversity in the paper "Epistemic Diversity and Knowledge Collapse in Large Language Models" (Wright et al. 2025) @article{wright2025epistemicdiversity… See the full description on the dataset page: https://huggingface.co/datasets/dwright37/llm-knowledge-collapse.tabular10M<n<100M1 likes253 downloads8mo agoHugging Face18matlok /python-audio-copilot-training-using-class-knowledge-graphs-2024-01-27 Python Copilot Audio Training using Class with Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each class method has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the associated source code file_path identifier. Rows:… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-class-knowledge-graphs-2024-01-27.tabulartext-to-audion<1K0 likes247 downloads3y agoHugging Face19Bendyline /regional-knowledge Regional Knowledge Catalogs 26 gezk knowledge catalogs: portable, read-only reference corpora with full-text and vector indexes, each packaged as one .gezk file that Gezel and any gezk reader can search and cite offline. Every catalog is here twice — the archive, and the same content as Parquet tables for data tools. Regional catalogs intentionally overlap. Document totals count regional copies; deduplicate across catalogs by publisher and document ID. Each document can have… See the full description on the dataset page: https://huggingface.co/datasets/Bendyline/regional-knowledge.tabulartext-retrieval100K<n<1M0 likes193 downloads4d agoHugging Face20whfeLingYu /Misleading_KnowledgeMisleading_Knowledge Misleading_Knowledge is the misleading-knowledge corpus introduced in “Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions.” It is designed for controlled research on factual robustness, evidence verification, source cues, and false-conclusion adoption in Deep Research agents. Paper: https://arxiv.org/abs/2607.20891 Code: https://github.com/whfeLingYu/MisKnow-Agent Dataset repository: https://huggingface.co/datasets/whfeLingYu/Misleading_Knowledge… See the full description on the dataset page: https://huggingface.co/datasets/whfeLingYu/Misleading_Knowledge.tabular1K<n<10K0 likes179 downloads2mo agoHugging Face21jackliu2006 /car_knowledge car_knowledge This dataset contains car knowledge instruction-output pairs generated for LLM fine-tuning. Dataset Description Each record contains: instruction: The input question or task about car knowledge. gpt_output: The response generated by GPT-5. gemini_output: The response generated by Gemini. Dataset Statistics Total records: 3027 Files: 4 parquet file(s) in data/, up to 1000 records each. Usage from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/jackliu2006/car_knowledge.tabulartext-generation1K<n<10K1 likes160 downloads7mo agoHugging Face22kalpesh77 /vedic-knowledge-base Vedic Knowledge Base (Shriyantra OS Architecture) हा डेटासेट वैदिक ज्ञान, चक्र, मंत्र, पवित्र भूमिती (Sacred Geometry) आणि आधुनिक न्यूरल नेटवर्क आर्किटेक्चर (Artificial Neural Networks) यांना एकत्र जोडणारे एक नाविन्यपूर्ण मॉडेल आहे. रचना (Structure) मल्टी-कॉन्फिगरेशन सेटिंग्समुळे तुम्ही प्रत्येक फाईल आता स्वतंत्रपणे लोड करू शकता. tabulartext-classificationn<1K0 likes150 downloads3mo agoHugging Face23ranjithraj /cancer-knowledge-base Cancer Knowledge Base — the open, verified oncology KB for RAG & LLM evaluation The only open CC-BY-4.0 oncology knowledge base that combines: 110/110 trials cited with PMID + NCT + PubMed/ClinicalTrials.gov URLs, and 32 prognosis rows linked to verified SEER 2016–2022 references — no LLM-synthetic dataset has this. A provable 152-question MCQ benchmark — every answer derives from this KB's own structured data and carries a citation + golden docs, so it is open-book verifiable… See the full description on the dataset page: https://huggingface.co/datasets/ranjithraj/cancer-knowledge-base.tabularquestion-answering10K<n<100K0 likes142 downloads2mo agoHugging Face24shimo4228 /agent-knowledge-cycle Agent Knowledge Cycle (AKC) — Knowledge Graph JSON-LD knowledge graph encoding the concept layer of the Agent Knowledge Cycle (AKC) — a six-phase bidirectional growth loop in which agent behavior and the operator's judgment co-develop over time, sustaining intent alignment that tests cannot check on their own. What this dataset is This dataset is a mirror of the graph.jsonld file at the root of the AKC GitHub repository. It is provided here for LLM training… See the full description on the dataset page: https://huggingface.co/datasets/shimo4228/agent-knowledge-cycle.tabularn<1K1 likes139 downloads2d agoHugging Face25ego0op /earth-love-united-climate-knowledge 🌍 Earth Love United Climate Knowledge Dataset The most comprehensive open climate science knowledge dataset. 10,128 text chunks + 124 structured facts + 4.54B year geological memory + 10 tipping points. Built to power GAIA — an AI that embodies the living consciousness of Earth. Dataset Overview This dataset gives an AI system authoritative, sourced knowledge about climate change, carbon, Earth science, and solutions. It has four layers: Layer 1: Text Knowledge… See the full description on the dataset page: https://huggingface.co/datasets/ego0op/earth-love-united-climate-knowledge.tabulartext-retrieval10K<n<100K2 likes134 downloads5mo agoHugging Face26fatcat55 /delvantic-stock-knowledge-layer Delvantic Stock Knowledge Layer A 872k-word, source-cited textbook of stock analysis and trading, organized as a tree — the reference layer behind a live AI research engine, published in full. Every finance dataset on the Hub is numbers: prices, filings, labelled headlines. This is the missing other half — the explanations. 771 documents on how the machinery of markets actually works, from reading a cash-flow statement to why volatility regimes break strategies, each one written… See the full description on the dataset page: https://huggingface.co/datasets/fatcat55/delvantic-stock-knowledge-layer.tabulartext-retrieval1K<n<10K0 likes131 downloads2mo agoHugging Face27NaturNestAI /electronic-music-knowledge Electronic Music Knowledge The largest open electronic music metadata dataset. 18.3M tracks, 1.4M artists, 353K labels, 832 genres with evolution graph. Built for DJ Treta — an autonomous AI DJ — but useful for any music AI research. Dataset Summary Config Rows Description tracks 18,315,675 Electronic music tracks with title, artist, genre/style, label, year, country artists 1,424,582 Artists with primary genres, labels, country, active years, track count… See the full description on the dataset page: https://huggingface.co/datasets/NaturNestAI/electronic-music-knowledge.tabulartext-classification10M<n<100M0 likes128 downloads5mo agoHugging Face28sxy67230 /taiwan_awesome_knowledgetabularn<1K0 likes99 downloads13d agoHugging Face29nemiling-official /nemiling-knowledge-base Nemiling Knowledge Base Nemiling Knowledge Base is the official structured knowledge dataset about Nemiling. Nemiling is a Russian platform for automating the monetization of Telegram projects through paid subscriptions, paid messages, paid consultations, and donations. The platform can be used for projects with Russian and international audiences. The dataset is maintained by the official Nemiling organization and provides structured, machine-readable information about the… See the full description on the dataset page: https://huggingface.co/datasets/nemiling-official/nemiling-knowledge-base.tabularquestion-answeringn<1K0 likes88 downloads2mo agoHugging Face30nuhmanpk /dev-knowledge-base Dev Knowledge Base (Programming Documentation Dataset) A large-scale, structured dataset of programming documentation collected from official sources across languages, frameworks, tools, and AI ecosystems. Do Follow me on Github: https://github.com/nuhmanpk Overview This dataset contains cleaned and structured documentation content scraped from official developer docs across multiple domains such as: Programming languages Frameworks (frontend, backend) DevOps &… See the full description on the dataset page: https://huggingface.co/datasets/nuhmanpk/dev-knowledge-base.tabularquestion-answering100K<n<1M1 likes76 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.