Team Ai
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Stage-jh-monitor /total-300-lambda00-s_signal_type6-jh-epoch4 total-300-lambda00-s_signal_type6-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.3875 Action score: 0.43125 Valid samples: 320/320 tabularn<1K0 likes8.4k downloads1mo agoHugging Face02Stage-jh-monitor /total-300-lambda10-s_signal_type6-jh-epoch4 total-300-lambda10-s_signal_type6-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.36640625 Action score: 0.41875 Valid samples: 320/320 tabularn<1K0 likes8.4k downloads1mo agoHugging Face03Stage-jh-monitor /total-300noapp-lambda02-s_signal_type6-jh-epoch4 total-300noapp-lambda02-s_signal_type6-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.36640625 Action score: 0.409375 Valid samples: 320/320 tabularn<1K0 likes8.4k downloads1mo agoHugging Face04Stage-jh-monitor /total-300-lambda02-s_signal_type6-jh-retry-epoch4 total-300-lambda02-s_signal_type6-jh-retry-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.36953125 Action score: 0.3984375 Valid samples: 320/320 tabularn<1K0 likes8.4k downloads1mo agoHugging Face05Stage-jh-monitor /total-300-lambda02-s_signal_type6-jh-epoch4-reeval1 total-300-lambda02-s_signal_type6-jh-epoch4-reeval1 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.38828125 Action score: 0.4234375 Valid samples: 320/320 tabularn<1K0 likes8.4k downloads1mo agoHugging Face06Stage-jh-monitor /total-300-lambda02-s_signal_type6-jh-epoch4-reeval2 total-300-lambda02-s_signal_type6-jh-epoch4-reeval2 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.4125 Action score: 0.4265625 Valid samples: 320/320 tabularn<1K0 likes8.4k downloads1mo agoHugging Face07Stage-jh-monitor /total-300-lambda05-s_signal_type6-jh-epoch4 total-300-lambda05-s_signal_type6-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.35703125 Action score: 0.4375 Valid samples: 320/320 tabularn<1K0 likes8.4k downloads1mo agoHugging Face08Stage-jh-monitor /total-300-lambda08-s_signal_type6-jh-epoch4 total-300-lambda08-s_signal_type6-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.38046875 Action score: 0.4078125 Valid samples: 320/320 tabularn<1K0 likes8.4k downloads1mo agoHugging Face09Stage-jh-monitor /total-300-lambda02-s_signal_type6-jh-epoch4 total-300-lambda02-s_signal_type6-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.4046875 Action score: 0.4140625 Valid samples: 320/320 tabularn<1K0 likes8.4k downloads1mo agoHugging Face10Stage-jh-monitor /total-300app-lambda02-s_signal_type6-jh-epoch4 total-300app-lambda02-s_signal_type6-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.3625 Action score: 0.4015625 Valid samples: 320/320 tabularn<1K0 likes8.4k downloads1mo agoHugging Face11Stage-jh-monitor /total-131-lambda02-residual-s_signal_type6-jh-epoch4 total-131-lambda02-residual-s_signal_type6-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.3765625 Action score: 0.4171875 Valid samples: 320/320 tabularn<1K0 likes8.4k downloads1mo agoHugging Face12SamuelChien821 /typed-decision-bench Typed Decision Bench v0.3 Built by Blobfish AI. A benchmark for one-pass decision models: 5,387 items, 25 tasks, 5 use-case suites. Blobfish designed the tasks, wrote the typed questions, framed each one as a decision a business actually delegates (use case, vertical), drew stratified seeded panels, froze them, and built the scoring, the contamination tiers and the quality scorecard. The underlying records are drawn from 21 openly licensed public datasets plus one generator of… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/typed-decision-bench.tabulartext-classification10K<n<100K0 likes592 downloads20d agoHugging Face13open-llm-leaderboard /cluebbers__Llama-3.1-8B-paraphrase-type-generation-apty-sigmoid-detailsgated Dataset Card for Evaluation run of cluebbers/Llama-3.1-8B-paraphrase-type-generation-apty-sigmoid Dataset automatically created during the evaluation run of model cluebbers/Llama-3.1-8B-paraphrase-type-generation-apty-sigmoid The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/cluebbers__Llama-3.1-8B-paraphrase-type-generation-apty-sigmoid-details.tabular10K<n<100K0 likes61 downloads2y agoHugging Face14Santhiyarajan /typed-requirement-atoms Typed Requirement Atoms for Omission Detection The training set for the typed requirement layer of Rubric (Lacuna v3). Each example is a single source atom annotated with the four decisions the layer must make about it, jointly: is it required for this task? (requiredness), what ontology category is it? (multi-label), how severe is its omission? (criticality), and does it apply at all? (applicability). This is what lets omission detection move from a flat lexical salience… See the full description on the dataset page: https://huggingface.co/datasets/Santhiyarajan/typed-requirement-atoms.tabulartext-classification10K<n<100K0 likes48 downloads3mo agoHugging Face15open-llm-leaderboard /cluebbers__Llama-3.1-8B-paraphrase-type-generation-apty-ipo-detailsgated Dataset Card for Evaluation run of cluebbers/Llama-3.1-8B-paraphrase-type-generation-apty-ipo Dataset automatically created during the evaluation run of model cluebbers/Llama-3.1-8B-paraphrase-type-generation-apty-ipo The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/cluebbers__Llama-3.1-8B-paraphrase-type-generation-apty-ipo-details.tabular10K<n<100K0 likes27 downloads2y agoHugging Face16open-llm-leaderboard /cluebbers__Llama-3.1-8B-paraphrase-type-generation-etpc-detailsgated Dataset Card for Evaluation run of cluebbers/Llama-3.1-8B-paraphrase-type-generation-etpc Dataset automatically created during the evaluation run of model cluebbers/Llama-3.1-8B-paraphrase-type-generation-etpc The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/cluebbers__Llama-3.1-8B-paraphrase-type-generation-etpc-details.tabular10K<n<100K0 likes22 downloads2y agoHugging Face17Type-1-Civilisation /SLM-Math-Bench-1-mediumtabular10K<n<100K0 likes15 downloads3mo agoHugging Face18sumit-178 /dhrona-content-typestabularn<1K0 likes2 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.