Team Ai
26 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01aereecho /xpo-explainer-data0 likes366 downloads5mo agoHugging Face02build-small-hackathon /blood-test-explainer-traces Blood Test Explainer - agent traces Agent traces from the Blood Test Explainer app (Build Small hackathon). Each row is one publicly-available sample lab report (fake patients, no PHI) run through the full agent pipeline: a small vision model reads the document and extracts the markers, then a curated medical knowledge base turns the values into a grounded, per-marker explanation plus cross-marker patterns. Model: build-small-hackathon/blood-test-minicpmv-4_6-medreason, a… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/blood-test-explainer-traces.textimage-text-to-textn<1K1 likes82 downloads4mo agoHugging Face03sagard21 /autotrain-data-code-explainergated AutoTrain Dataset for project: code-explainer Dataset Description This dataset has been automatically processed by AutoTrain for project code-explainer. Languages The BCP-47 code for the dataset's language is unk. Dataset Structure Data Instances A sample from this dataset looks as follows: [ { "text": "def upload_to_s3(local_file, bucket, s3_file):\n ## This function is responsible for uploading the file into the S3 bucket using… See the full description on the dataset page: https://huggingface.co/datasets/sagard21/autotrain-data-code-explainer.summarization2 likes66 downloads4y agoHugging Face04osbm /unet-explainer-dataimage1K<n<10K1 likes26 downloads3y agoHugging Face05Ramitha /alignllm-explainer-metricstabular1K<n<10K0 likes26 downloads8mo agoHugging Face06klimczakjakubdev /pharmaco-explainer Pharmaco-Explainer Datasets This repository contains datasets used in the Pharmaco-Explainer project. They are shared separately on Hugging Face and are used by the training and experimentation code hosted on GitHub: 👉 Training code and scripts:https://github.com/AdamSulek/pharmaco-explainer/ Available Datasets The following datasets are available: k3: 3-element pharmacophore k4_2ar: 4-element pharmacophore with two aromatic features k4: 4-element pharmacophore… See the full description on the dataset page: https://huggingface.co/datasets/klimczakjakubdev/pharmaco-explainer.tabular10M<n<100M0 likes15 downloads4mo agoHugging Face07fireblaster234 /smolified-offline-legal-explainer 🤏 smolified-offline-legal-explainer Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model fireblaster234/smolified-offline-legal-explainer. 📦 Asset Details Origin: Smolify Foundry (Job ID: 0b4fe722) Records: 1363 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by fireblaster234. Generated via Smolify.ai. texttext-generation1K<n<10K0 likes14 downloads8mo agoHugging Face08Betha /fen-position-explainer-validationtext1K<n<10K0 likes12 downloads2y agoHugging Face09shabul /feynman-explainer-dataset Feynman Explainer Synthetic Dataset A compact synthetic instruction dataset for training models to explain concepts in a Feynman-style voice: analogy first, intuition before jargon, and flowing prose instead of bullets. This dataset was created for the qwen2.5-3b-feynman-explainer fine-tune and includes both the raw synthetic examples and the chat-formatted train/validation splits used for MLX LoRA training. What is in the repo data/raw_feynman.jsonl: 575 raw… See the full description on the dataset page: https://huggingface.co/datasets/shabul/feynman-explainer-dataset.text-generationn<1K0 likes12 downloads6mo agoHugging Face10NguyenAn05 /exact_2026_model2_explainertext1K<n<10K0 likes10 downloads4mo agoHugging Face11Ramitha /rq3-all-records-combination-explainer-metricstabular1K<n<10K0 likes9 downloads9mo agoHugging Face12JakeClark /soliaudit-dasp-sequence-gcn-explainertabular10K<n<100K0 likes9 downloads6mo agoHugging Face13carseng /titleix-explainergated Title IX Respondent Explainer (Atomizer-ready) Purpose. An instruction-tuning dataset designed to train an information-only explainer bot for Title IX respondents. The bot helps users understand fields on a Title IX form, timelines, rights, and process basics. It does not give legal advice and does not make determinations about responsibility. Audience: Respondents (the party accused) using a Title IX website or form.Scope: Descriptive/educational answers only — no adjudication, no… See the full description on the dataset page: https://huggingface.co/datasets/carseng/titleix-explainer.texttext-generation1K<n<10K0 likes8 downloads1y agoHugging Face14JakeClark /soliaudit-dasp-sequence-gnn-no-explainertabular10K<n<100K0 likes7 downloads6mo agoHugging Face15kh4dien /explainer-gemma-2_simulator-qwen2.5text10K<n<100K0 likes6 downloads2y agoHugging Face16JakeClark /soliaudit-dasp-sequence-gnn-explainertabular10K<n<100K0 likes6 downloads6mo agoHugging Face17raniero /sn96g-science-explainers-10chunk1-20250919_133203 Subnet 96 — Clean Q/A Dataset Format: one JSONL per line: {"system": null, "conversations":[{"role":"user","content":"..."}, {"role":"assistant","content":"..."}]} Total pairs: 26 Avg answer length (tokens): 125.7 (median 124.0, min 96, max 176) Schema errors: 0 (should be 0) File size: 0.02 MB SHA256 (data.jsonl): bededf0d5e8154fd243766c0faf1e1e9bbb3071c97b9c4664d273e3b9180a7d3 Language: English Intended for: Bittensor Subnet 96 validators Generation: local LLaMA (GPU) +… See the full description on the dataset page: https://huggingface.co/datasets/raniero/sn96g-science-explainers-10chunk1-20250919_133203.textn<1K0 likes4 downloads1y agoHugging Face18raniero /sn96g-science-explainers-2chunk1-20250919_193451 Subnet 96 — Clean Q/A Dataset Format: one JSONL per line: {"system": null, "conversations":[{"role":"user","content":"..."}, {"role":"assistant","content":"..."}]} Total pairs: 2 Avg answer length (tokens): 35 (median 35.0, min 26, max 44) Schema errors: 0 (should be 0) File size: 0.00 MB SHA256 (data.jsonl): 3d134bc47c74ba1f958620f97a9ed58f30fa62c14311e565ac699bde4aa5f089 Language: English Intended for: Bittensor Subnet 96 validators Generation: local LLaMA (GPU) +… See the full description on the dataset page: https://huggingface.co/datasets/raniero/sn96g-science-explainers-2chunk1-20250919_193451.textn<1K0 likes4 downloads1y agoHugging Face19carseng /titleix_explainer_reformatgated Dataset Card for Title IX Respondent Explainers (Structured, 2020 Regs) Dataset Summary A curated instruction-tuning dataset for plain-language, respondent-focused Title IX explanations aligned with the 2020 federal regulations.Each example follows a six-section template: What this means Who it applies to What to expect Your options now Important cautions Where to confirm The dataset avoids legal advice and school-specific promises; timelines are framed as… See the full description on the dataset page: https://huggingface.co/datasets/carseng/titleix_explainer_reformat.tabulartext-generation1K<n<10K0 likes4 downloads1y agoHugging Face20VarunKVK /code-explainer-datasettextn<1K0 likes4 downloads9mo agoHugging Face21Advids /explainer-videos0 likes3 downloads2y agoHugging Face22raniero /sn96g-science-explainers-2chunk1-20250919_175433 Subnet 96 — Clean Q/A Dataset Format: one JSONL per line: {"system": null, "conversations":[{"role":"user","content":"..."}, {"role":"assistant","content":"..."}]} Total pairs: 2 Avg answer length (tokens): 36.5 (median 36.5, min 30, max 43) Schema errors: 0 (should be 0) File size: 0.00 MB SHA256 (data.jsonl): 9cef190fed1c2111e72c6871b8192450eb97ba7037bf614467acb849825cf6f8 Language: English Intended for: Bittensor Subnet 96 validators Generation: local LLaMA (GPU) +… See the full description on the dataset page: https://huggingface.co/datasets/raniero/sn96g-science-explainers-2chunk1-20250919_175433.textn<1K0 likes3 downloads1y agoHugging Face23raniero /sn96g-science-explainers-2chunk1-20250919_184511 Subnet 96 — Clean Q/A Dataset Format: one JSONL per line: {"system": null, "conversations":[{"role":"user","content":"..."}, {"role":"assistant","content":"..."}]} Total pairs: 2 Avg answer length (tokens): 32.5 (median 32.5, min 13, max 52) Schema errors: 0 (should be 0) File size: 0.00 MB SHA256 (data.jsonl): 217328e167aa722bfe23ac052014390d4aed1b453113248639790ab73e0a600e Language: English Intended for: Bittensor Subnet 96 validators Generation: local LLaMA (GPU) +… See the full description on the dataset page: https://huggingface.co/datasets/raniero/sn96g-science-explainers-2chunk1-20250919_184511.textn<1K0 likes3 downloads1y agoHugging Face24protolyze /explainer_sys_prompt0 likes2 downloads1y agoHugging Face25protolyze /explainer_model0 likes2 downloads1y agoHugging Face26carseng /titleix_explainer_casualgated pretty_name: "Title IX Respondent Explainers (Casual, 2020 Regs)" license: "cc-by-sa-4.0" language: - en tags: - law - education - safety - assistant - instruction-tuning - compliance task_categories: - text2text-generation size_categories: - 1K<n<10K Dataset Card for Title IX Respondent Explainers (Casual, 2020 Regs) Dataset Summary A companion dataset that paraphrases the structured six-section explanations into 2–3 short paragraphs with a… See the full description on the dataset page: https://huggingface.co/datasets/carseng/titleix_explainer_casual.text1K<n<10K0 likes2 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.