datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Womens_Clothing_E-Commerce_Reviewsabag-xm
AbAg-XM
Computed on Tenstorrent hardware with TT-Bio.
335,360 antibody-antigen structure predictions from four independently trained models, every one
scored against the experimental structure with DockQ. 512 samples per target per model, no cell
shallower than 512.
The targets are 2026ARK-AB, the antibody-antigen benchmark
released with OpenDDE. 164 PDB targets, 404 interfaces, 159 clusters at 40% MMseqs2 entity
clustering. We did not assemble that set and take no credit for… See the full description on the dataset page: https://huggingface.co/datasets/Tenstorrent/abag-xm.LongChat-Lines
Dataset Card for "LongChat-Lines"
This dataset is was used to evaluate the performance of model finetuned to operate on longer contexts. It is based on
a task template proposed by LMSys to evaluate attention to arbitrary points in the context. See the full details at
https;//github.com/abacusai/Long-Context.
Cognitive_Atrophy_Benchmark
Cognitive Atrophy Benchmark — LLM Responses Across Four Mental-Health Conversation Datasets
This dataset releases the LLM-response component of the Cognitive Atrophy Benchmark: five large language models prompted under identical conditions across four mental-health conversation datasets. It is a building block for a forthcoming evaluation framework that quantifies cognitive atrophy — the gradual erosion of users' own reasoning, recall, and decisional autonomy when an LLM… See the full description on the dataset page: https://huggingface.co/datasets/abadawi/Cognitive_Atrophy_Benchmark.abalone
Abalone
The Abalone dataset from the UCI ML repository.
Predict the age of the given abalone.
Configurations and tasks
Configuration
Task
Description
abalone
Regression
Predict the age of the abalone.
binary
Binary classification
Does the abalone have more than 9 rings?
Usage
from datasets import load_dataset
dataset = load_dataset("mstz/abalone")["train"]
Features
Target feature in bold.
Feature
Type
sex
[string]… See the full description on the dataset page: https://huggingface.co/datasets/mstz/abalone.AbAssayBench
AbAssayBench
This dataset repository contains the processed data package for AbAssayBench,
a multi-endpoint antibody developability benchmark. The release combines the
FLAb2.0-derived antibody measurements used for model development with the
PROPHET-Ab measurements used for external validation.
The repository is intended to be used together with the MAP-Ab source code:
https://github.com/gu-yaowen/MAP-Ab.
Package layout
Path
Contents
tables/
Release… See the full description on the dataset page: https://huggingface.co/datasets/yg3191/AbAssayBench.lm-eval-results-AbacusResearch-jaLLAbi2-7b-private
Dataset Card for Evaluation run of AbacusResearch/jaLLAbi2-7b
Dataset automatically created during the evaluation run of model AbacusResearch/jaLLAbi2-7b
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-AbacusResearch-jaLLAbi2-7b-private.abacusai__Llama-3-Smaug-8B-details
Dataset Card for Evaluation run of abacusai/Llama-3-Smaug-8B
Dataset automatically created during the evaluation run of model abacusai/Llama-3-Smaug-8B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/abacusai__Llama-3-Smaug-8B-details.abacusai__Dracarys-72B-Instruct-details
Dataset Card for Evaluation run of abacusai/Dracarys-72B-Instruct
Dataset automatically created during the evaluation run of model abacusai/Dracarys-72B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/abacusai__Dracarys-72B-Instruct-details.robocasa-100demos-6chosen-tasks-for-aBao_lerobot_v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 414,
"total_frames": 140509,
"total_tasks": 146,
"total_videos": 2484,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:414"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/binhng/robocasa-100demos-6chosen-tasks-for-aBao_lerobot_v1.abacus_4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 10,
"total_frames": 13486,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot-edinburgh-white-team/abacus_4.abacus_3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 10,
"total_frames": 11986,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot-edinburgh-white-team/abacus_3.AbAgym
AbAgym
AbAgym is a curated dataset of deep mutational scanning (DMS) measurements for antibody-antigen complexes. This Hugging Face version reorganizes the original AbAgym files into loadable dataset configurations using Apache Parquet, while preserving the original structure archive.
The original AbAgym repository describes the dataset as containing 68 DMS datasets on antibody-antigen complexes, approximately 324,000 non-redundant mutations, 36,541 non-redundant interface mutations… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/AbAgym.abacus_distractorsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 10,
"total_frames": 8991,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot-edinburgh-white-team/abacus_distractors.abacus_2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 10,
"total_frames": 8994,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot-edinburgh-white-team/abacus_2.abacus_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 10,
"total_frames": 8991,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot-edinburgh-white-team/abacus_test.robocasa-30-6chosen-tasks-for-aBao_lerobot_v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 123,
"total_frames": 41284,
"total_tasks": 84,
"total_videos": 738,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:123"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/binhng/robocasa-30-6chosen-tasks-for-aBao_lerobot_v1.aba07990a873258bba0d0b32325be11386712474abacus_1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 10,
"total_frames": 8991,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot-edinburgh-white-team/abacus_1.eval_abacus_1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 1,
"total_frames": 1995,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot-edinburgh-white-team/eval_abacus_1.abacusai__Smaug-72B-v0.1-details
Dataset Card for Evaluation run of abacusai/Smaug-72B-v0.1
Dataset automatically created during the evaluation run of model abacusai/Smaug-72B-v0.1
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/abacusai__Smaug-72B-v0.1-details.abacusai__Smaug-Mixtral-v0.1-details
Dataset Card for Evaluation run of abacusai/Smaug-Mixtral-v0.1
Dataset automatically created during the evaluation run of model abacusai/Smaug-Mixtral-v0.1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/abacusai__Smaug-Mixtral-v0.1-details.abacusai__Liberated-Qwen1.5-14B-details
Dataset Card for Evaluation run of abacusai/Liberated-Qwen1.5-14B
Dataset automatically created during the evaluation run of model abacusai/Liberated-Qwen1.5-14B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/abacusai__Liberated-Qwen1.5-14B-details.abacusai__Smaug-Qwen2-72B-Instruct-details
Dataset Card for Evaluation run of abacusai/Smaug-Qwen2-72B-Instruct
Dataset automatically created during the evaluation run of model abacusai/Smaug-Qwen2-72B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/abacusai__Smaug-Qwen2-72B-Instruct-details.abacusai__bigstral-12b-32k-details
Dataset Card for Evaluation run of abacusai/bigstral-12b-32k
Dataset automatically created during the evaluation run of model abacusai/bigstral-12b-32k
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/abacusai__bigstral-12b-32k-details.abacusai__Smaug-Llama-3-70B-Instruct-32K-details
Dataset Card for Evaluation run of abacusai/Smaug-Llama-3-70B-Instruct-32K
Dataset automatically created during the evaluation run of model abacusai/Smaug-Llama-3-70B-Instruct-32K
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/abacusai__Smaug-Llama-3-70B-Instruct-32K-details.abacusai__Smaug-34B-v0.1-details
Dataset Card for Evaluation run of abacusai/Smaug-34B-v0.1
Dataset automatically created during the evaluation run of model abacusai/Smaug-34B-v0.1
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/abacusai__Smaug-34B-v0.1-details.AbacusResearch__Jallabi-34B-details
Dataset Card for Evaluation run of AbacusResearch/Jallabi-34B
Dataset automatically created during the evaluation run of model AbacusResearch/Jallabi-34B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/AbacusResearch__Jallabi-34B-details.africa-synth-abattoir-inspection-surveillance-all
Abattoir Inspection Surveillance | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health datasets help… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-abattoir-inspection-surveillance-all.details_abacusai__bigyi-15b
Dataset Card for Evaluation run of abacusai/bigyi-15b
Dataset automatically created during the evaluation run of model abacusai/bigyi-15b.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_abacusai__bigyi-15b.
