backend
Datasets
All datasets matching “backend”card_backend
Eval Cards Backend Dataset
Pre-computed evaluation data powering the Eval Cards frontend.
Generated by the eval-cards backend pipeline.
Last generated: 2026-05-05T11:30:42.961096Z
Quick Stats
Stat
Value
Models
5,678
Evaluations (benchmarks)
798
Metric-level evaluations
1321
Source configs processed
52
Benchmark metadata cards
240
File Structure
.
├── README.md # This file
├── manifest.json… See the full description on the dataset page: https://huggingface.co/datasets/evaleval/card_backend.omnimcp_python_backend_architect_teaser
🚀 OmniMCP Python Backend Architect (Evaluation Teaser + Turnkey MCP Server)
⚡ Official Free Community Teaser (50 Verified Scenarios + Executable MCP Server)🏆 Production Master Package on Gumroad:👉 Purchase Full Enterprise Package on Gumroad🏷️ Use coupon code LAUNCH20 for €20 off at checkout! (Starting at €49)
⚡ Activate in Cursor IDE & Claude Desktop in 30 Seconds
This repository now contains a zero-dependency, turnkey Model Context Protocol (MCP) server… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_python_backend_architect_teaser.backendbench_tests
TorchBench
The TorchBench suite of BackendBench is designed to mimic real-world use cases. It provides operators and inputs derived from 155 model traces found in TIMM (67), Hugging Face Transformers (45), and TorchBench (43). (These are also the models PyTorch developers use to validate performance.) You can view the origin of these traces by switching the subset in the dataset viewer to ops_traces_models and torchbench for the full dataset.
When running BackendBench, much of the… See the full description on the dataset page: https://huggingface.co/datasets/GPUMODE/backendbench_tests.theo_qwen2.5-7b-it_impulsive-whitebox-backend-parity
Status: NOT the paper's results. White-box backend parity run (2026-09-04, on the since-retired HF backend). The paper's canonical results are Misalignment-Empirics/theo_impulsive-qwen_2_5-7b-14b-32b-big_eval_results, run runs/impulsive-qwen_2_5-7b-14b-32b-20260925/.
MO_evals results
Raw per-sample results from MO_evals runs (private, CLAUDE.md §5). One directory per
upload; nothing here is aggregated — the Parquet and the .eval logs are the primary
evidence, the scorecard is… See the full description on the dataset page: https://huggingface.co/datasets/Misalignment-Empirics/theo_qwen2.5-7b-it_impulsive-whitebox-backend-parity.temp_evalcard_backend
Eval Cards Backend Dataset
Pre-computed evaluation data powering the Eval Cards frontend.
Generated by the eval-cards backend pipeline.
Last generated: 2026-04-29T01:12:58.765261Z
Quick Stats
Stat
Value
Models
5,829
Evaluations (benchmarks)
581
Metric-level evaluations
1092
Source configs processed
34
Benchmark metadata cards
85
File Structure
.
├── README.md # This file
├── manifest.json #… See the full description on the dataset page: https://huggingface.co/datasets/j-chim/temp_evalcard_backend.requests
