Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hallucinations-leaderboard /results22 likes851k downloads2y agoHugging Face02gorilla-llm /Berkeley-Function-Calling-Leaderboard Berkeley Function Calling Leaderboard The Berkeley function calling leaderboard is a live leaderboard to evaluate the ability of different LLMs to call functions (also referred to as tools). We built this dataset from our learnings to be representative of most users' function calling use-cases, for example, in agents, as a part of enterprise workflows, etc. To this end, our evaluation dataset spans diverse categories, and across multiple languages. Checkout the Leaderboard at… See the full description on the dataset page: https://huggingface.co/datasets/gorilla-llm/Berkeley-Function-Calling-Leaderboard.127 likes148k downloads5mo agoHugging Face03eduagarcia-temp /llm_pt_leaderboard_raw_results0 likes97k downloads1y agoHugging Face04open-llm-leaderboard-old /requests Open LLM Leaderboard Requests This repository contains the request files of models that have been submitted to the Open LLM Leaderboard. You can take a look at the current status of your model by finding its request file in this dataset. If your model failed, feel free to open an issue on the Open LLM Leaderboard! (We don't follow issues in this repository as often) Evaluation Methodology The evaluation process involves running your models against several benchmarks from… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/requests.textn<1K22 likes96k downloads2y agoHugging Face05open-llm-leaderboard /requests13 likes90k downloads15h agoHugging Face06lmarena-ai /leaderboard-dataset Arena Leaderboard Dataset Historical snapshots of the Arena leaderboard. Usage from datasets import load_dataset # Load all historical text style control data ds = load_dataset("lmarena-ai/leaderboard-dataset", "text_style_control", split="full") # Load the current text style control leaderboard ds = load_dataset("lmarena-ai/leaderboard-dataset", "text_style_control", split="latest") # Filter to overall category ds =… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset.tabular1M<n<10M28 likes72k downloads3d agoHugging Face07morteza20 /mteb_leaderboard [!NOTE]Previously it was possible to submit models results to MTEB by adding the results to the model metadata. This is no longer an option as we want to ensure high quality metadata. This repository contain the results of the embedding benchmark evaluated using the package mteb. Reference 🦾 Leaderboard An up to date leaderboard of embedding models 📚 mteb Guides and instructions on how to use mteb, including running, submitting scores, etc. 🙋 Questions Questions about the… See the full description on the dataset page: https://huggingface.co/datasets/morteza20/mteb_leaderboard.0 likes43k downloads2y agoHugging Face08qimma /leaderboard-detailstext1M<n<10M0 likes32k downloads1mo agoHugging Face09open-llm-leaderboard /results20 likes29k downloads2y agoHugging Face10hallucinations-leaderboard /requests0 likes28k downloads2y agoHugging Face11open-cn-llm-leaderboard /requests1 likes27k downloads2y agoHugging Face12huggingface-projects /drlc-leaderboard-datatabular10K<n<100K2 likes24k downloads3d agoHugging Face13open-ko-llm-leaderboard /requests0 likes23k downloads2y agoHugging Face14hf-audio /open-asr-leaderboard ESB Test Sets: Parquet & Sorted This dataset takes the open-asr-leaderboard/datasets-test-only data and sorts each split by audio length. The format is also changed, from custom loading script (un-safe remote code) to parquet (safe). Broadly speaking, this dataset was generated with the following code-snippet: from datasets import load_dataset, get_dataset_config_names DATASET = "open-asr-leaderboard/datasets-test-only" # dataset to load from HUB_DATASET_ID =… See the full description on the dataset page: https://huggingface.co/datasets/hf-audio/open-asr-leaderboard.audio100K<n<1M84 likes23k downloads3mo agoHugging Face15llm-jp /leaderboard-results1 likes19k downloads1y agoHugging Face16cot-leaderboard /cot-eval-traces-2.0text1M<n<10M9 likes14k downloads2y agoHugging Face17open-llm-leaderboard /contentstabular1K<n<10K25 likes14k downloads2y agoHugging Face18qimma /leaderboard-requests0 likes13k downloads28d agoHugging Face19benchflow /skillsbench-leaderboard SkillsBench Leaderboard and Evidence Archive This repository stores public SkillsBench submissions, raw BenchFlow trial artifacts, trajectory evidence, audit reports, and the release-aligned official leaderboard exports. Official benchmark definition: benchflow/skillsbenchLatest public benchmark release: SkillsBench v1.1Latest source commit: 27738384b1df694ea2ae466e416f476e94d8fab9 Current Official Release The latest public results are under:… See the full description on the dataset page: https://huggingface.co/datasets/benchflow/skillsbench-leaderboard.3 likes13k downloads4mo agoHugging Face20hf-audio /open-asr-leaderboard-resultstabularn<1K1 likes10k downloads6h agoHugging Face21open-ko-llm-leaderboard /requests-backup5 likes9.2k downloads2y agoHugging Face22harborframework /terminal-bench-2-leaderboard Terminal-Bench 2.0 Leaderboard Submissions This repository accepts leaderboard submissions for Terminal-Bench 2.0. How to Submit Fork this repository Create a new branch for your submission Add your submission (a job or folder of jobs) under submissions/terminal-bench/2.0/<agent>__<model(s)>/ Open a Pull Request Submission Structure submissions/ terminal-bench/ 2.0/ <agent>__<model>/ metadata.yaml # Required: agent and model info… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench-2-leaderboard.36 likes9.2k downloads5mo agoHugging Face23open-cn-llm-leaderboard /results0 likes7.6k downloads2y agoHugging Face24nebius /SWE-rebench-leaderboard Dataset Summary ❗❗❗ Please use Harbour Hub for the July 2026 evaluation split:https://hub.harborframework.com/datasets/ibragim-badertdinov/swe-rebench-07-2026/latest SWE-rebench-leaderboard is a continuously updated, curated subset of the full SWE-rebench corpus, tailored for benchmarking software engineering agents on real-world tasks. These tasks are used in the SWE-rebench leaderboard. For more details on the benchmark methodology and data collection process, please refer to… See the full description on the dataset page: https://huggingface.co/datasets/nebius/SWE-rebench-leaderboard.tabular1K<n<10K30 likes7.4k downloads2mo agoHugging Face25open-cn-llm-leaderboard /vlm_results0 likes7.3k downloads1y agoHugging Face26eduagarcia-temp /llm_pt_leaderboard_requests0 likes7.1k downloads7mo agoHugging Face27IntelligenceLab /LHTB-leaderboard LHTB Leaderboard — Long-Horizon Terminal-Bench This repository hosts submitted runs for Long-Horizon Terminal-Bench (LHTB), a 46-task benchmark measuring how well LLM agents sustain useful work in a containerized terminal over hundreds of steps. Every entry below ships its complete run artifacts — per-trial configs, results, verifier outputs and terminal recordings — so any score on this board can be audited without rerunning the suite. 📊 Benchmark dataset:… See the full description on the dataset page: https://huggingface.co/datasets/IntelligenceLab/LHTB-leaderboard.tabular1K<n<10K3 likes5.8k downloads14d agoHugging Face28open-cn-llm-leaderboard /vlm_requests0 likes5.6k downloads1y agoHugging Face29bezzam /tts_leaderboard_screenshotsimagen<1K0 likes4.9k downloads7d agoHugging Face30Weyaxi /huggingface-leaderboard Huggingface Leaderboard's History Dataset 🏆 This is the history dataset of Huggingface Leaderboard. 🗒️ This dataset contains full dataframes in a CSV file for each time lapse. ⌛ This dataset is automatically updated when space restarts. (Which is approximately every 6 hours) Leaderboard Link 🔗 Weyaxi/huggingface-leaderboard 12 likes4.9k downloads4d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.