datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Womens_Clothing_E-Commerce_Reviewsabag-xm
AbAg-XM
Computed on Tenstorrent hardware with TT-Bio.
335,360 antibody-antigen structure predictions from four independently trained models, every one
scored against the experimental structure with DockQ. 512 samples per target per model, no cell
shallower than 512.
The targets are 2026ARK-AB, the antibody-antigen benchmark
released with OpenDDE. 164 PDB targets, 404 interfaces, 159 clusters at 40% MMseqs2 entity
clustering. We did not assemble that set and take no credit for… See the full description on the dataset page: https://huggingface.co/datasets/Tenstorrent/abag-xm.WikiQA-Free_Form_QA
Dataset Card for "WikiQA-Free_Form_QA"
The WikiQA task is the task of answering a question based on the information given in a Wikipedia document. We have built upon the short answer format data in Google Natural Questions to construct our QA task. It is formatted as a document and a question. We ensure the answer to the question is a short answer which is either a single word or a small sentence directly cut pasted from the document. Having the task structured as such, we can… See the full description on the dataset page: https://huggingface.co/datasets/abacusai/WikiQA-Free_Form_QA.Frappe-mobile-app-usageDataset Description: Frappe Processed Dataset
The Frappe dataset has been processed to refine the quality of user-item interactions by removing entries where either users or items had fewer than 5 interactions. This pruning resulted in a significant reduction in the dataset size:
Number of Users: 651 (a reduction of 31.97% from the original dataset)
Number of Items: 1127 (a reduction of 72.39%)
Total Number of Interactions: 84,373 (a reduction of 12.30%)
Columns Overview:
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/abadesalex/Frappe-mobile-app-usage.r14-d136r14-d135r14-d137r14-d133data_abaw_auLongChat-Lines
Dataset Card for "LongChat-Lines"
This dataset is was used to evaluate the performance of model finetuned to operate on longer contexts. It is based on
a task template proposed by LMSys to evaluate attention to arbitrary points in the context. See the full details at
https;//github.com/abacusai/Long-Context.
WikiQA-Altered_Numeric_QA
Dataset Card for "WikiQA-Altered_Numeric_QA"
The WikiQA task is the task of answering a question based on the information given in a Wikipedia document. We have built upon the short answer format data in Google Natural Questions to construct our QA task. It is formatted as a document and a question. We ensure the answer to the question is a short answer which is either a single word or a small sentence directly cut pasted from the document. Having the task structured as such, we… See the full description on the dataset page: https://huggingface.co/datasets/abacusai/WikiQA-Altered_Numeric_QA.details_abacusai__Smaug-72B-v0.1
Dataset Card for Evaluation run of abacusai/Smaug-72B-v0.1
Dataset automatically created during the evaluation run of model abacusai/Smaug-72B-v0.1 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_abacusai__Smaug-72B-v0.1.Cognitive_Atrophy_Benchmark
Cognitive Atrophy Benchmark — LLM Responses Across Four Mental-Health Conversation Datasets
This dataset releases the LLM-response component of the Cognitive Atrophy Benchmark: five large language models prompted under identical conditions across four mental-health conversation datasets. It is a building block for a forthcoming evaluation framework that quantifies cognitive atrophy — the gradual erosion of users' own reasoning, recall, and decisional autonomy when an LLM… See the full description on the dataset page: https://huggingface.co/datasets/abadawi/Cognitive_Atrophy_Benchmark.leyu-amharic-addis-ababa-dialect
Leyu Amharic - Addis Ababa Dialect Speech Corpus
Dataset Description
This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Addis Ababa dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent… See the full description on the dataset page: https://huggingface.co/datasets/leyu-amharic/leyu-amharic-addis-ababa-dialect.HUVER
Dataset Card for HUVER
The dataset is comprised of a 6,051 unique UAV configurations, where each configuration is described by multiple data for-
mats, including a grammar string, an RGB image, and an GLB file.
Complementing these representation modalities, we also provide a configuration-based description, i.e., a text descriptor describing the features of each UAV using natural language
Curated by: Abhiram Karri, Gary Stump, Christopher McComb, Binyang Song
Language(s)… See the full description on the dataset page: https://huggingface.co/datasets/AbaloneVH/HUVER.sudokubench
Dataset Card for SudokuBench
Dataset Details
This dataset contains a list of sudoku puzzles and their solutions, all at
varying levels of difficulty.
The difficulties are based on the number of squares (also sometimes referred to
as cells) that are provided at the start of the puzzle.
The puzzles are guaranteed to have a single unique solution without any overlap.
Within a difficulty config, you will find 10,000 puzzles at every number of
available cells at the start of… See the full description on the dataset page: https://huggingface.co/datasets/abatilo/sudokubench.robocasa-30and100demos-6chosen-tasks-for-aBaoMetaMath_DPO_FewShot
Dataset Card for "MetaMath_DPO_FewShot"
GSM8K \citep{cobbe2021training} is a dataset of diverse grade school maths word problems, which has been commonly adopted as a measure of the math and reasoning skills of LLMs.
The MetaMath dataset is an extension of the training set of GSM8K using data augmentation.
It is partitioned into queries and responses, where the query is a question involving mathematical calculation or reasoning, and the response is a logical series of steps and… See the full description on the dataset page: https://huggingface.co/datasets/abacusai/MetaMath_DPO_FewShot.details_abacusai__Smaug-2-72B
Dataset Card for Evaluation run of abacusai/Smaug-2-72B
Dataset automatically created during the evaluation run of model abacusai/Smaug-2-72B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_abacusai__Smaug-2-72B.abalone
Abalone
The Abalone dataset from the UCI ML repository.
Predict the age of the given abalone.
Configurations and tasks
Configuration
Task
Description
abalone
Regression
Predict the age of the abalone.
binary
Binary classification
Does the abalone have more than 9 rings?
Usage
from datasets import load_dataset
dataset = load_dataset("mstz/abalone")["train"]
Features
Target feature in bold.
Feature
Type
sex
[string]… See the full description on the dataset page: https://huggingface.co/datasets/mstz/abalone.MentalBench-Align
MentalBench–100k & MentalAlign–70k: Dual Benchmark Suite for Mental Health LLM Evaluation
📄 Paper (arXiv): When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation📎 Paper Link: https://arxiv.org/pdf/2510.19032
📦 Code & Documentation: https://github.com/abeerbadawi/MentalBench-Align
📘 Overview
This repository introduces two complementary datasets that enable systematic evaluation of large language models (LLMs) in… See the full description on the dataset page: https://huggingface.co/datasets/abadawi/MentalBench-Align.MetaMathFewshot
A few-shot version of the MetaMath (https://huggingface.co/datasets/meta-math/MetaMathQA) dataset.
Each entry is formatted with 'question' and 'answer' keys. The 'question' key has a random number of query-answer pairs between 0 and 4 inclusive, before a final target query; the expected answer to this is stored in the content of 'answer'.
parameter-golf-sp4096AbAssayBench
AbAssayBench
This dataset repository contains the processed data package for AbAssayBench,
a multi-endpoint antibody developability benchmark. The release combines the
FLAb2.0-derived antibody measurements used for model development with the
PROPHET-Ab measurements used for external validation.
The repository is intended to be used together with the MAP-Ab source code:
https://github.com/gu-yaowen/MAP-Ab.
Package layout
Path
Contents
tables/
Release… See the full description on the dataset page: https://huggingface.co/datasets/yg3191/AbAssayBench.details_AbacusResearch__Jallabi-34B
Dataset Card for Evaluation run of AbacusResearch/Jallabi-34B
Dataset automatically created during the evaluation run of model AbacusResearch/Jallabi-34B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_AbacusResearch__Jallabi-34B.ham10000_bbox
HAM10000 with Spatial Annotations and Bounding Box Coordinates
Enhanced version of HAM10000 dataset with bounding box coordinates and spatial descriptions for skin lesion localization.
Dataset Description
This dataset extends the original HAM10000 dermatology dataset with:
Bounding box coordinates for lesion localization
Spatial descriptions (e.g., "located in center-center region")
Area coverage statistics
Mask availability flags
Features
image: RGB skin… See the full description on the dataset page: https://huggingface.co/datasets/abaryan/ham10000_bbox.details_abacusai__Llama-3-Giraffe-70B-Instructlm-eval-results-AbacusResearch-jaLLAbi2-7b-private
Dataset Card for Evaluation run of AbacusResearch/jaLLAbi2-7b
Dataset automatically created during the evaluation run of model AbacusResearch/jaLLAbi2-7b
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-AbacusResearch-jaLLAbi2-7b-private.details_abacusai__Llama-3-Giraffe-70Bleyu-amharic-addis-ababa-dialect
Leyu Amharic - Addis Ababa Dialect Speech Corpus
Dataset Description
This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Addis Ababa dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent… See the full description on the dataset page: https://huggingface.co/datasets/gheero-Leyu/leyu-amharic-addis-ababa-dialect.
