leaderboard
zeta-2.1-autoround-W4A16KAT-Coder-V2.5-Dev-AutoRound-W4A16-Tuninggpt2-demoQwopus3.6-27B-Coder-AutoRound-W4A16-RTNENERGY-DRINK-LOVE_-_leaderboard_inst_v1.3_deup_LDCC-SOLAR-10.7B_SFT-ggufENERGY-DRINK-LOVE_-_leaderboard_inst_v1.3_Open-Hermes_LDCC-SOLAR-10.7B_SFT-ggufENERGY-DRINK-LOVE_-_leaderboard_inst_v1.5_LDCC-SOLAR-10.7B_SFT-ggufENERGY-DRINK-LOVE_-_eeve_leaderboard_inst_v1.5-gguf
resultsBerkeley-Function-Calling-Leaderboard
Berkeley Function Calling Leaderboard
The Berkeley function calling leaderboard is a live leaderboard to evaluate the ability of different LLMs to call functions (also referred to as tools).
We built this dataset from our learnings to be representative of most users' function calling use-cases, for example, in agents, as a part of enterprise workflows, etc.
To this end, our evaluation dataset spans diverse categories, and across multiple languages.
Checkout the Leaderboard at… See the full description on the dataset page: https://huggingface.co/datasets/gorilla-llm/Berkeley-Function-Calling-Leaderboard.llm_pt_leaderboard_raw_resultsrequests
Open LLM Leaderboard Requests
This repository contains the request files of models that have been submitted to the Open LLM Leaderboard.
You can take a look at the current status of your model by finding its request file in this dataset. If your model failed, feel free to open an issue on the Open LLM Leaderboard! (We don't follow issues in this repository as often)
Evaluation Methodology
The evaluation process involves running your models against several benchmarks from… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/requests.requestsleaderboard-dataset
Arena Leaderboard Dataset
Historical snapshots of the Arena leaderboard.
Usage
from datasets import load_dataset
# Load all historical text style control data
ds = load_dataset("lmarena-ai/leaderboard-dataset", "text_style_control", split="full")
# Load the current text style control leaderboard
ds = load_dataset("lmarena-ai/leaderboard-dataset", "text_style_control", split="latest")
# Filter to overall category
ds =… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset.
