benchmarks
SynTTS-Commands-Media-Benchmarksedge-llm-benchmarks-gguftinyllama-1.1b-gguf-benchmarkshacking-fairness-benchmarks-qwen2.5-7b-z251hacking-fairness-benchmarks-qwen2.5-7b-z2hacking-fairness-benchmarks-qwen3-8b-base-z1BingoGuard-bert-base-portuguese-cased-benchmarkshacking-fairness-benchmarks-mistral-7b-v0.3-z1
transformersVideo-MME-v2
🔥 News
2026.06.11 Videos re-encoded to H265, maintaining consistent evaluation scores. Fixed 2 incorrect MP4s & 3 mismatched URLs. Original data preserved in the original branch.
2026.05.22 Task types are now available for Q1-Q3 in coherence (logic) groups.
🤗 About This Repo
This repository contains annotation data for "Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding". It mainly consists of three… See the full description on the dataset page: https://huggingface.co/datasets/MME-Benchmarks/Video-MME-v2.daytrader-benchmarkswhat-ai-benchmarks-actually-measure
What AI Benchmarks Actually Measure: Item-Level Model Outputs and Scores for 53 Models
Item-level model responses and scores for 53 language models across the
56 benchmarks analyzed in What AI Benchmarks Actually Measure: Adapting
Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks
(Desai et al., 2026,
arxiv.org/abs/2609.08812).
We do not release the prompts from the benchmark datasets, but instead refer to them by
item ids. To regenerate the prompts from… See the full description on the dataset page: https://huggingface.co/datasets/madesai/what-ai-benchmarks-actually-measure.MiroFlow-BenchmarksThese are the benchmarking datasets used for MiroFlow Framework. More information: https://github.com/MiroMindAI/MiroThinker
speculator_benchmarksThis dataset contains dataset splits for evaluating speculative decoding algorithms on different tasks.
File:
Coding: HumanEval.jsonl
Math: math_reasoning.jsonl
Question Answering: qa.jsonl
MT_bench: question.jsonl
Retrieval-Augmented Generation: rag.jsonl
Summarization: summarization.jsonl
Translation (German to English): translation.jsonl
Writing: writing.jsonl
The data comes from two sources:
https://github.com/openai/human-eval (1). (The MIT License)… See the full description on the dataset page: https://huggingface.co/datasets/RedHatAI/speculator_benchmarks.
