datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
xpo-explainer-datablood-test-explainer-traces
Blood Test Explainer - agent traces
Agent traces from the Blood Test Explainer app (Build Small hackathon). Each row is one publicly-available sample lab report (fake patients, no PHI) run through the full agent pipeline: a small vision model reads the document and extracts the markers, then a curated medical knowledge base turns the values into a grounded, per-marker explanation plus cross-marker patterns.
Model: build-small-hackathon/blood-test-minicpmv-4_6-medreason, a… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/blood-test-explainer-traces.autotrain-data-code-explainer
AutoTrain Dataset for project: code-explainer
Dataset Description
This dataset has been automatically processed by AutoTrain for project code-explainer.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"text": "def upload_to_s3(local_file, bucket, s3_file):\n ## This function is responsible for uploading the file into the S3 bucket using… See the full description on the dataset page: https://huggingface.co/datasets/sagard21/autotrain-data-code-explainer.unet-explainer-dataalignllm-explainer-metricspharmaco-explainer
Pharmaco-Explainer Datasets
This repository contains datasets used in the Pharmaco-Explainer project.
They are shared separately on Hugging Face and are used by the training and
experimentation code hosted on GitHub:
👉 Training code and scripts:https://github.com/AdamSulek/pharmaco-explainer/
Available Datasets
The following datasets are available:
k3: 3-element pharmacophore
k4_2ar: 4-element pharmacophore with two aromatic features
k4: 4-element pharmacophore… See the full description on the dataset page: https://huggingface.co/datasets/klimczakjakubdev/pharmaco-explainer.smolified-offline-legal-explainer
🤏 smolified-offline-legal-explainer
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model fireblaster234/smolified-offline-legal-explainer.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 0b4fe722)
Records: 1363
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by fireblaster234.
Generated via Smolify.ai.
fen-position-explainer-validationfeynman-explainer-dataset
Feynman Explainer Synthetic Dataset
A compact synthetic instruction dataset for training models to explain concepts in a Feynman-style voice: analogy first, intuition before jargon, and flowing prose instead of bullets.
This dataset was created for the qwen2.5-3b-feynman-explainer fine-tune and includes both the raw synthetic examples and the chat-formatted train/validation splits used for MLX LoRA training.
What is in the repo
data/raw_feynman.jsonl: 575 raw… See the full description on the dataset page: https://huggingface.co/datasets/shabul/feynman-explainer-dataset.exact_2026_model2_explainerrq3-all-records-combination-explainer-metricssoliaudit-dasp-sequence-gcn-explainertitleix-explainer
Title IX Respondent Explainer (Atomizer-ready)
Purpose. An instruction-tuning dataset designed to train an information-only explainer bot for Title IX respondents. The bot helps users understand fields on a Title IX form, timelines, rights, and process basics. It does not give legal advice and does not make determinations about responsibility.
Audience: Respondents (the party accused) using a Title IX website or form.Scope: Descriptive/educational answers only — no adjudication, no… See the full description on the dataset page: https://huggingface.co/datasets/carseng/titleix-explainer.soliaudit-dasp-sequence-gnn-no-explainerexplainer-gemma-2_simulator-qwen2.5soliaudit-dasp-sequence-gnn-explainersn96g-science-explainers-10chunk1-20250919_133203
Subnet 96 — Clean Q/A Dataset
Format: one JSONL per line:
{"system": null, "conversations":[{"role":"user","content":"..."}, {"role":"assistant","content":"..."}]}
Total pairs: 26
Avg answer length (tokens): 125.7 (median 124.0, min 96, max 176)
Schema errors: 0 (should be 0)
File size: 0.02 MB
SHA256 (data.jsonl): bededf0d5e8154fd243766c0faf1e1e9bbb3071c97b9c4664d273e3b9180a7d3
Language: English
Intended for: Bittensor Subnet 96 validators
Generation: local LLaMA (GPU) +… See the full description on the dataset page: https://huggingface.co/datasets/raniero/sn96g-science-explainers-10chunk1-20250919_133203.sn96g-science-explainers-2chunk1-20250919_193451
Subnet 96 — Clean Q/A Dataset
Format: one JSONL per line:
{"system": null, "conversations":[{"role":"user","content":"..."}, {"role":"assistant","content":"..."}]}
Total pairs: 2
Avg answer length (tokens): 35 (median 35.0, min 26, max 44)
Schema errors: 0 (should be 0)
File size: 0.00 MB
SHA256 (data.jsonl): 3d134bc47c74ba1f958620f97a9ed58f30fa62c14311e565ac699bde4aa5f089
Language: English
Intended for: Bittensor Subnet 96 validators
Generation: local LLaMA (GPU) +… See the full description on the dataset page: https://huggingface.co/datasets/raniero/sn96g-science-explainers-2chunk1-20250919_193451.titleix_explainer_reformat
Dataset Card for Title IX Respondent Explainers (Structured, 2020 Regs)
Dataset Summary
A curated instruction-tuning dataset for plain-language, respondent-focused Title IX explanations aligned with the 2020 federal regulations.Each example follows a six-section template:
What this means
Who it applies to
What to expect
Your options now
Important cautions
Where to confirm
The dataset avoids legal advice and school-specific promises; timelines are framed as… See the full description on the dataset page: https://huggingface.co/datasets/carseng/titleix_explainer_reformat.code-explainer-datasetexplainer-videossn96g-science-explainers-2chunk1-20250919_175433
Subnet 96 — Clean Q/A Dataset
Format: one JSONL per line:
{"system": null, "conversations":[{"role":"user","content":"..."}, {"role":"assistant","content":"..."}]}
Total pairs: 2
Avg answer length (tokens): 36.5 (median 36.5, min 30, max 43)
Schema errors: 0 (should be 0)
File size: 0.00 MB
SHA256 (data.jsonl): 9cef190fed1c2111e72c6871b8192450eb97ba7037bf614467acb849825cf6f8
Language: English
Intended for: Bittensor Subnet 96 validators
Generation: local LLaMA (GPU) +… See the full description on the dataset page: https://huggingface.co/datasets/raniero/sn96g-science-explainers-2chunk1-20250919_175433.sn96g-science-explainers-2chunk1-20250919_184511
Subnet 96 — Clean Q/A Dataset
Format: one JSONL per line:
{"system": null, "conversations":[{"role":"user","content":"..."}, {"role":"assistant","content":"..."}]}
Total pairs: 2
Avg answer length (tokens): 32.5 (median 32.5, min 13, max 52)
Schema errors: 0 (should be 0)
File size: 0.00 MB
SHA256 (data.jsonl): 217328e167aa722bfe23ac052014390d4aed1b453113248639790ab73e0a600e
Language: English
Intended for: Bittensor Subnet 96 validators
Generation: local LLaMA (GPU) +… See the full description on the dataset page: https://huggingface.co/datasets/raniero/sn96g-science-explainers-2chunk1-20250919_184511.explainer_sys_promptexplainer_modeltitleix_explainer_casual
pretty_name: "Title IX Respondent Explainers (Casual, 2020 Regs)"
license: "cc-by-sa-4.0"
language:
- en
tags:
- law
- education
- safety
- assistant
- instruction-tuning
- compliance
task_categories:
- text2text-generation
size_categories:
- 1K<n<10K
Dataset Card for Title IX Respondent Explainers (Casual, 2020 Regs)
Dataset Summary
A companion dataset that paraphrases the structured six-section explanations into 2–3 short paragraphs with a… See the full description on the dataset page: https://huggingface.co/datasets/carseng/titleix_explainer_casual.
