datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lm-eval-results-AbacusResearch-jaLLAbi2-7b-private
Dataset Card for Evaluation run of AbacusResearch/jaLLAbi2-7b
Dataset automatically created during the evaluation run of model AbacusResearch/jaLLAbi2-7b
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-AbacusResearch-jaLLAbi2-7b-private.abacusai__Llama-3-Smaug-8B-details
Dataset Card for Evaluation run of abacusai/Llama-3-Smaug-8B
Dataset automatically created during the evaluation run of model abacusai/Llama-3-Smaug-8B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/abacusai__Llama-3-Smaug-8B-details.abacusai__Dracarys-72B-Instruct-details
Dataset Card for Evaluation run of abacusai/Dracarys-72B-Instruct
Dataset automatically created during the evaluation run of model abacusai/Dracarys-72B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/abacusai__Dracarys-72B-Instruct-details.aba07990a873258bba0d0b32325be11386712474abacusai__Smaug-72B-v0.1-details
Dataset Card for Evaluation run of abacusai/Smaug-72B-v0.1
Dataset automatically created during the evaluation run of model abacusai/Smaug-72B-v0.1
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/abacusai__Smaug-72B-v0.1-details.abacusai__Smaug-Mixtral-v0.1-details
Dataset Card for Evaluation run of abacusai/Smaug-Mixtral-v0.1
Dataset automatically created during the evaluation run of model abacusai/Smaug-Mixtral-v0.1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/abacusai__Smaug-Mixtral-v0.1-details.abacusai__Liberated-Qwen1.5-14B-details
Dataset Card for Evaluation run of abacusai/Liberated-Qwen1.5-14B
Dataset automatically created during the evaluation run of model abacusai/Liberated-Qwen1.5-14B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/abacusai__Liberated-Qwen1.5-14B-details.abacusai__Smaug-Qwen2-72B-Instruct-details
Dataset Card for Evaluation run of abacusai/Smaug-Qwen2-72B-Instruct
Dataset automatically created during the evaluation run of model abacusai/Smaug-Qwen2-72B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/abacusai__Smaug-Qwen2-72B-Instruct-details.abacusai__bigstral-12b-32k-details
Dataset Card for Evaluation run of abacusai/bigstral-12b-32k
Dataset automatically created during the evaluation run of model abacusai/bigstral-12b-32k
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/abacusai__bigstral-12b-32k-details.abacusai__Smaug-Llama-3-70B-Instruct-32K-details
Dataset Card for Evaluation run of abacusai/Smaug-Llama-3-70B-Instruct-32K
Dataset automatically created during the evaluation run of model abacusai/Smaug-Llama-3-70B-Instruct-32K
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/abacusai__Smaug-Llama-3-70B-Instruct-32K-details.abacusai__Smaug-34B-v0.1-details
Dataset Card for Evaluation run of abacusai/Smaug-34B-v0.1
Dataset automatically created during the evaluation run of model abacusai/Smaug-34B-v0.1
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/abacusai__Smaug-34B-v0.1-details.AbacusResearch__Jallabi-34B-details
Dataset Card for Evaluation run of AbacusResearch/Jallabi-34B
Dataset automatically created during the evaluation run of model AbacusResearch/Jallabi-34B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/AbacusResearch__Jallabi-34B-details.abacusai__bigyi-15b-details
Dataset Card for Evaluation run of abacusai/bigyi-15b
Dataset automatically created during the evaluation run of model abacusai/bigyi-15b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/abacusai__bigyi-15b-details.synthetic-abandoned-cart-email-examples
Synthetic Abandoned Cart Email Examples
An entirely synthetic, bilingual collection of abandoned-cart email drafts with transparent checklist annotations. It is intended for education, prototyping, and evaluation, and contains no real recipients, customer messages, orders, merchant data, or campaign results.
Dataset Description
The dataset mirrors the five visible checks in NeuroCheckout's public Abandoned Cart Email Checker:
message clarity;
primary call to… See the full description on the dataset page: https://huggingface.co/datasets/neurocheckout-ai/synthetic-abandoned-cart-email-examples.lm-eval-results-AbacusResearch-haLLawa4-7b-private
Dataset Card for Evaluation run of AbacusResearch/haLLawa4-7b
Dataset automatically created during the evaluation run of model AbacusResearch/haLLawa4-7b
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-AbacusResearch-haLLawa4-7b-private.fuzzeval-humaneval-mbpp
FuzzEval unit tests for HumanEval-f and MBPP-f
Automatically generated unit tests for a reproduction of the ICML 2026 paper
"Towards Functional Correctness of Large Code Models with Selective Generation"
(Jeong, Kim & Park — arXiv:2505.13553,
official repo trustml-lab/selective-code-generation).
The paper's FuzzEval paradigm replaces a benchmark's handful of hand-written
unit tests with hundreds of unit tests obtained by fuzzing the reference
solution. This dataset is our… See the full description on the dataset page: https://huggingface.co/datasets/ababa134/fuzzeval-humaneval-mbpp.
