Team Ai
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01witcheer /rtx-5090-benchmarks RTX 5090 LLM Benchmarks Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig. Quality Benchmarks Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency. Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/rtx-5090-benchmarks.imagetext-generationn<1K1 likes1.5k downloads3d agoHugging Face02omegaprime669 /rtx-5090-benchmarks RTX 5090 LLM Benchmarks Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig. Quality Benchmarks Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency. Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/omegaprime669/rtx-5090-benchmarks.imagetext-generationn<1K0 likes218 downloads3mo agoHugging Face03Multimedia-SMU /seeingculture-benchmarkPaper | Project Page | Leaderboard | Explorer | Code | CMB, the video successor Seeing Culture Benchmark (SCB) Evaluating Visual Reasoning and Grounding in Cultural Context The Seeing Culture Benchmark (SCB) evaluates cultural reasoning in vision-language models in two stages: i) selecting the correct visual option with multiple-choice visual question answering (VQA), and ii) segmenting the relevant cultural artifact as evidence of reasoning. Visual options in… See the full description on the dataset page: https://huggingface.co/datasets/Multimedia-SMU/seeingculture-benchmark.imageimage-text-to-text1K<n<10K3 likes128 downloads15d agoHugging Face04ivitskiy /google-ads-benchmark-2026 Note on checksums. This README.md carries the YAML dataset-card header required by the Hugging Face hub, so its SHA-256 differs from the entry in checksums.txt; that entry refers to the canonical README in the GitHub mirror. All data files are byte-identical across mirrors. Load any table with load_dataset("ivitskiy/google-ads-benchmark-2026", "<config_name>"). Ivitskiy Ads Lab: Google Ads Panel & Benchmark Compilation 2026 (Open Research Dataset) Two things in one package.… See the full description on the dataset page: https://huggingface.co/datasets/ivitskiy/google-ads-benchmark-2026.image1K<n<10K0 likes106 downloads19d agoHugging Face05boczkakaroly /hungarian-riddles-benchmark Hungarian Riddles Benchmark Overview This dataset is a cultural and reasoning benchmark based on 100 metaphorical, trivia-style Hungarian riddles. The riddles are intentionally tricky and culturally grounded. They are designed to test answer correctness and reasoning quality, not only surface-level language fluency. Dataset structure Each row contains one riddle with reference material for evaluation. Fields ID – unique identifier topic – general… See the full description on the dataset page: https://huggingface.co/datasets/boczkakaroly/hungarian-riddles-benchmark.imagequestion-answeringn<1K0 likes59 downloads9mo agoHugging Face06UCSC-VLAA /gpt-image-edit-benchmark-results GPT-Image-Edit — Benchmark Results This repository contains evaluation results of GPT-Image-Edit across four standard image-editing benchmarks. All scores were computed using the official evaluation scripts provided by each benchmark. 📊 Benchmarks Benchmark Metrics Folder GEdit-EN 12 editing categories + Avg gedit/ Complex-Edit IF, IP, PQ, Overall complex_edit/ ImgEdit-Full 10 editing operations + Overall imgedit/ OmniContext Contextual edit scores… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/gpt-image-edit-benchmark-results.image1K<n<10K1 likes52 downloads1y agoHugging Face07nnanwube /controlnet-brand-fidelity-benchmark ControlNet Preprocessor Brand Fidelity Benchmark Study 1A — NaviTask Marketing Flyer | Phase 1 Results Status: Phase 1 complete (6 runs). Phase 2 in progress. Last updated: May 2026 The Enterprise Problem The global content marketing market is valued at $524.73 billion in 2025, projected to reach $989.84 billion by 2030 at a 13.53% CAGR. Enterprise adoption of generative AI has accelerated significantly, with 65% of organisations reporting regular use of… See the full description on the dataset page: https://huggingface.co/datasets/nnanwube/controlnet-brand-fidelity-benchmark.documentimage-to-imagen<1K0 likes33 downloads5mo agoHugging Face08lujunxi57 /debias_benchmarkimage10K<n<100K0 likes7 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.