Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01LocalLLaMA /typed-decisions Typed Decisions A benchmark for typed probabilistic decisions. A model gets one piece of unstructured state and answers five typed questions about it at once, and every answer is a probability distribution, not a single label. The schema follows the System One primitives (noul, choice, score) used by TypeSafe AI, so a row replays against any API with that shape. The benchmark is independent: it is not affiliated with TypeSafe and does not reproduce their Jev model.… See the full description on the dataset page: https://huggingface.co/datasets/LocalLLaMA/typed-decisions.tabulartext-classification1K<n<10K159 likes37k downloads15h agoHugging Face02Stage-jh-monitor /total-300-lambda00-s_signal_type6-jh-epoch4 total-300-lambda00-s_signal_type6-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.3875 Action score: 0.43125 Valid samples: 320/320 tabularn<1K0 likes8.4k downloads1mo agoHugging Face03Stage-jh-monitor /total-300-lambda10-s_signal_type6-jh-epoch4 total-300-lambda10-s_signal_type6-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.36640625 Action score: 0.41875 Valid samples: 320/320 tabularn<1K0 likes8.4k downloads1mo agoHugging Face04Stage-jh-monitor /total-300noapp-lambda02-s_signal_type6-jh-epoch4 total-300noapp-lambda02-s_signal_type6-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.36640625 Action score: 0.409375 Valid samples: 320/320 tabularn<1K0 likes8.4k downloads1mo agoHugging Face05Stage-jh-monitor /total-300-lambda02-s_signal_type6-jh-retry-epoch4 total-300-lambda02-s_signal_type6-jh-retry-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.36953125 Action score: 0.3984375 Valid samples: 320/320 tabularn<1K0 likes8.4k downloads1mo agoHugging Face06Stage-jh-monitor /total-300-lambda02-s_signal_type6-jh-epoch4-reeval1 total-300-lambda02-s_signal_type6-jh-epoch4-reeval1 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.38828125 Action score: 0.4234375 Valid samples: 320/320 tabularn<1K0 likes8.4k downloads1mo agoHugging Face07Stage-jh-monitor /total-300-lambda02-s_signal_type6-jh-epoch4-reeval2 total-300-lambda02-s_signal_type6-jh-epoch4-reeval2 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.4125 Action score: 0.4265625 Valid samples: 320/320 tabularn<1K0 likes8.4k downloads1mo agoHugging Face08Stage-jh-monitor /total-300-lambda05-s_signal_type6-jh-epoch4 total-300-lambda05-s_signal_type6-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.35703125 Action score: 0.4375 Valid samples: 320/320 tabularn<1K0 likes8.4k downloads1mo agoHugging Face09Stage-jh-monitor /total-300-lambda08-s_signal_type6-jh-epoch4 total-300-lambda08-s_signal_type6-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.38046875 Action score: 0.4078125 Valid samples: 320/320 tabularn<1K0 likes8.4k downloads1mo agoHugging Face10Stage-jh-monitor /total-300-lambda02-s_signal_type6-jh-epoch4 total-300-lambda02-s_signal_type6-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.4046875 Action score: 0.4140625 Valid samples: 320/320 tabularn<1K0 likes8.4k downloads1mo agoHugging Face11Stage-jh-monitor /total-300app-lambda02-s_signal_type6-jh-epoch4 total-300app-lambda02-s_signal_type6-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.3625 Action score: 0.4015625 Valid samples: 320/320 tabularn<1K0 likes8.4k downloads1mo agoHugging Face12Stage-jh-monitor /total-131-lambda02-residual-s_signal_type6-jh-epoch4 total-131-lambda02-residual-s_signal_type6-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.3765625 Action score: 0.4171875 Valid samples: 320/320 tabularn<1K0 likes8.4k downloads1mo agoHugging Face13tasksource /tasksource-jev-typed-decisions tasksource-jev-typed-decisions 2.5 million typed decisions (choices, ratings and probabilities) from 670 sources. Why use it Real supervision. Labels, ratings, and annotator votes come from established datasets, not a teacher model. Every row names its source. Breadth. Over 300 dataset families: NLI and reasoning, QA and commonsense, sentiment, intent and topic, toxicity and safety, preference pairs, fact checking, entity tagging, and dozens of languages. GLUE… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/tasksource-jev-typed-decisions.imagezero-shot-classification1M<n<10M20 likes6.4k downloads3h agoHugging Face14RealSR /batch_0602_typeIdocument0 likes3k downloads4mo agoHugging Face15tasksource /procedural-typed-decisions procedural-typed-decisions Procedurally generated decision problems. Each row is one structured state (JSON, or a table, CSV, key=value lines, or prose for the arithmetic, retrieval, and aggregation configs) with several typed questions over that same state, following the Jev / System One request shape: choice (pick one criterion), noul (a number in [0, 1]; a probability or a yes/no), and score (an ordered rubric). Every answer is computed exactly from the state by rules that… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/procedural-typed-decisions.tabulartext-classification100K<n<1M4 likes1.9k downloads11d agoHugging Face16typeof /ultrachat-sharegpt-5GBtext100K<n<1M0 likes1.6k downloads3y agoHugging Face17typesafe /evalsafe-invoice-processing Invoice processing Snapshot: 2026-09-28. 150 cases and 6,874 question instances. Default reference: consensus. Labels are model-generated references. Data Load configuration cases, questions, or run_results; all have a test split. cases: one row per case_id, with the complete input in input_json, descriptive metadata_json, and openai, anthropic, and consensus labelsets. Decisions are grouped by policy_id and contain status, actions, and primary_action. questions:… See the full description on the dataset page: https://huggingface.co/datasets/typesafe/evalsafe-invoice-processing.tabular1K<n<10K6 likes1.3k downloads11d agoHugging Face18RealSR /batch_0602_typeIIdocument0 likes1.2k downloads4mo agoHugging Face19typesafe /evalsafe-customer-service Customer service Snapshot: 2026-09-28. 204 cases and 3,287 question instances. Default reference: consensus. Labels are model-generated references. Data Load configuration cases, questions, or run_results; all have a test split. cases: one row per case_id, with the complete input in input_json, descriptive metadata_json, and openai, anthropic, and consensus labelsets. Decisions are grouped by policy_id and contain status, actions, and primary_action. questions:… See the full description on the dataset page: https://huggingface.co/datasets/typesafe/evalsafe-customer-service.tabular1K<n<10K2 likes1.2k downloads12d agoHugging Face20typesafe /evalsafe-security-incidents Security incidents Snapshot: 2026-09-28. 240 cases and 1,820 question instances. Default reference: consensus. Labels are model-generated references. Data Load configuration cases, questions, or run_results; all have a test split. cases: one row per case_id, with the complete input in input_json, descriptive metadata_json, and openai, anthropic, and consensus labelsets. Decisions are grouped by policy_id and contain status, actions, and primary_action. questions:… See the full description on the dataset page: https://huggingface.co/datasets/typesafe/evalsafe-security-incidents.tabular1K<n<10K2 likes1.1k downloads12d agoHugging Face21typesafe /evalsafe-agent-trace-observability Agent trace triage Snapshot: 2026-09-28. 111 cases and 1,124 question instances. Default reference: consensus. Labels are model-generated references. Data Load configuration cases, questions, or run_results; all have a test split. cases: one row per case_id, with the complete input in input_json, descriptive metadata_json, and openai, anthropic, and consensus labelsets. Decisions are grouped by policy_id and contain status, actions, and primary_action. questions:… See the full description on the dataset page: https://huggingface.co/datasets/typesafe/evalsafe-agent-trace-observability.tabular1K<n<10K2 likes1k downloads12d agoHugging Face22mathlib-initiative /mathlib-types Mathlib Types This dataset contains information about types defined in Mathlib, the mathematical library for the Lean 4 theorem prover, extracted with lean_scout. Extracted from the Mathlib commit with the following hash. d13f23b723b8a846827a245b89c10fc7d3f11612 The dataset follows this schema: fields: - type: datatype: string nullable: false name: name - type: datatype: string nullable: true name: module - type: datatype: string nullable: false name:… See the full description on the dataset page: https://huggingface.co/datasets/mathlib-initiative/mathlib-types.text100K<n<1M0 likes956 downloads14d agoHugging Face23typesafe /evalsafe-onet EvalSafe O*NET 150 documents · 7,500 consensus-labeled questions · 9 candidate models. Snapshot: 2026-09-29. Default reference: consensus. Only questions with an available consensus target and their corresponding documents and final model results are included. The documents are synthetic workplace examples. The reference targets are model-generated, using Astra (gpt-6-astra) and Fable (claude-fable-5-1). The default reference is their consensus. Load from datasets… See the full description on the dataset page: https://huggingface.co/datasets/typesafe/evalsafe-onet.tabular10K<n<100K7 likes920 downloads11d agoHugging Face24Eitanli /meal_type Dataset Card for "meal_type" More Information needed text10K<n<100K0 likes860 downloads3y agoHugging Face25typeof /algebraic-stack NOTE: Please see EleutherAI/proof-pile-2 This is a cherry-picked repackaging of the algebraic-stack segment from the proof-pile-2 dataset as parquet files License see EleutherAI/proof-pile-2 Citation see EleutherAI/proof-pile-2 texttext-generation1M<n<10M6 likes802 downloads3y agoHugging Face26n4ze3m /typed-decisions-synth Typed Decisions Synth This is the synthetic dataset I made for Hmm, a small open model that answers questions about your data with probabilities instead of text. It has 7,414 cases with 25,859 questions across 149 domains and workflows. Every question has an answer and a soft label (a probability for every option), so you can train a model to be unsure when it should be. Code and the model: github.com/n4ze3m/hmm Note: Everything here is written and labelled by an LLM. Nobody… See the full description on the dataset page: https://huggingface.co/datasets/n4ze3m/typed-decisions-synth.texttext-classification1K<n<10K2 likes657 downloads5h agoHugging Face27SamuelChien821 /typed-decision-bench Typed Decision Bench v0.3 Built by Blobfish AI. A benchmark for one-pass decision models: 5,387 items, 25 tasks, 5 use-case suites. Blobfish designed the tasks, wrote the typed questions, framed each one as a decision a business actually delegates (use case, vertical), drew stratified seeded panels, froze them, and built the scoring, the contamination tiers and the quality scorecard. The underlying records are drawn from 21 openly licensed public datasets plus one generator of… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/typed-decision-bench.tabulartext-classification10K<n<100K0 likes592 downloads20d agoHugging Face28Universal-NER /Pile-NER-type Intro Pile-NER-type is a set of GPT-generated data for named entity recognition using the type-based data construction prompt. It was collected by prompting gpt-3.5-turbo-0301 and augmented by negative sampling. Check our project page for more information. License Attribution-NonCommercial 4.0 International text10K<n<100K29 likes526 downloads3y agoHugging Face29kl08 /myers-briggs-type-indicatortext1K<n<10K2 likes413 downloads3y agoHugging Face30pngwn /typed-decisions-v2-system-onetext10K<n<100K1 likes399 downloads24d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.