Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01abayuu /Womens_Clothing_E-Commerce_Reviewstabular10K<n<100K0 likes2.9k downloads3y agoHugging Face02Tenstorrent /abag-xm AbAg-XM Computed on Tenstorrent hardware with TT-Bio. 335,360 antibody-antigen structure predictions from four independently trained models, every one scored against the experimental structure with DockQ. 512 samples per target per model, no cell shallower than 512. The targets are 2026ARK-AB, the antibody-antigen benchmark released with OpenDDE. 164 PDB targets, 404 interfaces, 159 clusters at 40% MMseqs2 entity clustering. We did not assemble that set and take no credit for… See the full description on the dataset page: https://huggingface.co/datasets/Tenstorrent/abag-xm.tabularother100K<n<1M1 likes1.3k downloads2mo agoHugging Face03abacusai /WikiQA-Free_Form_QA Dataset Card for "WikiQA-Free_Form_QA" The WikiQA task is the task of answering a question based on the information given in a Wikipedia document. We have built upon the short answer format data in Google Natural Questions to construct our QA task. It is formatted as a document and a question. We ensure the answer to the question is a short answer which is either a single word or a small sentence directly cut pasted from the document. Having the task structured as such, we can… See the full description on the dataset page: https://huggingface.co/datasets/abacusai/WikiQA-Free_Form_QA.text1K<n<10K17 likes950 downloads3y agoHugging Face04abadesalex /Frappe-mobile-app-usageDataset Description: Frappe Processed Dataset The Frappe dataset has been processed to refine the quality of user-item interactions by removing entries where either users or items had fewer than 5 interactions. This pruning resulted in a significant reduction in the dataset size: Number of Users: 651 (a reduction of 31.97% from the original dataset) Number of Items: 1127 (a reduction of 72.39%) Total Number of Interactions: 84,373 (a reduction of 12.30%) Columns Overview: The dataset… See the full description on the dataset page: https://huggingface.co/datasets/abadesalex/Frappe-mobile-app-usage.10K<n<100K2 likes666 downloads2y agoHugging Face05abacada /r14-d136textn<1K0 likes642 downloads5d agoHugging Face06abacada /r14-d135textn<1K0 likes636 downloads5d agoHugging Face07abacada /r14-d137textn<1K0 likes636 downloads5d agoHugging Face08abacada /r14-d133textn<1K0 likes631 downloads5d agoHugging Face09doanh25032004 /data_abaw_auimage10K<n<100K0 likes630 downloads2y agoHugging Face10abacusai /LongChat-Lines Dataset Card for "LongChat-Lines" This dataset is was used to evaluate the performance of model finetuned to operate on longer contexts. It is based on a task template proposed by LMSys to evaluate attention to arbitrary points in the context. See the full details at https;//github.com/abacusai/Long-Context. tabularn<1K23 likes508 downloads3y agoHugging Face11abacusai /WikiQA-Altered_Numeric_QA Dataset Card for "WikiQA-Altered_Numeric_QA" The WikiQA task is the task of answering a question based on the information given in a Wikipedia document. We have built upon the short answer format data in Google Natural Questions to construct our QA task. It is formatted as a document and a question. We ensure the answer to the question is a short answer which is either a single word or a small sentence directly cut pasted from the document. Having the task structured as such, we… See the full description on the dataset page: https://huggingface.co/datasets/abacusai/WikiQA-Altered_Numeric_QA.text1K<n<10K13 likes440 downloads3y agoHugging Face12open-llm-leaderboard-old /details_abacusai__Smaug-72B-v0.1 Dataset Card for Evaluation run of abacusai/Smaug-72B-v0.1 Dataset automatically created during the evaluation run of model abacusai/Smaug-72B-v0.1 on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_abacusai__Smaug-72B-v0.1.0 likes288 downloads3y agoHugging Face13abadawi /Cognitive_Atrophy_Benchmark Cognitive Atrophy Benchmark — LLM Responses Across Four Mental-Health Conversation Datasets This dataset releases the LLM-response component of the Cognitive Atrophy Benchmark: five large language models prompted under identical conditions across four mental-health conversation datasets. It is a building block for a forthcoming evaluation framework that quantifies cognitive atrophy — the gradual erosion of users' own reasoning, recall, and decisional autonomy when an LLM… See the full description on the dataset page: https://huggingface.co/datasets/abadawi/Cognitive_Atrophy_Benchmark.tabulartext-generation10K<n<100K2 likes253 downloads3mo agoHugging Face14leyu-amharic /leyu-amharic-addis-ababa-dialect Leyu Amharic - Addis Ababa Dialect Speech Corpus Dataset Description This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Addis Ababa dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent… See the full description on the dataset page: https://huggingface.co/datasets/leyu-amharic/leyu-amharic-addis-ababa-dialect.audioautomatic-speech-recognition1K<n<10K1 likes240 downloads2mo agoHugging Face15AbaloneVH /HUVER Dataset Card for HUVER The dataset is comprised of a 6,051 unique UAV configurations, where each configuration is described by multiple data for- mats, including a grammar string, an RGB image, and an GLB file. Complementing these representation modalities, we also provide a configuration-based description, i.e., a text descriptor describing the features of each UAV using natural language Curated by: Abhiram Karri, Gary Stump, Christopher McComb, Binyang Song Language(s)… See the full description on the dataset page: https://huggingface.co/datasets/AbaloneVH/HUVER.3dimage-to-text1K<n<10K0 likes223 downloads1mo agoHugging Face16abatilo /sudokubench Dataset Card for SudokuBench Dataset Details This dataset contains a list of sudoku puzzles and their solutions, all at varying levels of difficulty. The difficulties are based on the number of squares (also sometimes referred to as cells) that are provided at the start of the puzzle. The puzzles are guaranteed to have a single unique solution without any overlap. Within a difficulty config, you will find 10,000 puzzles at every number of available cells at the start of… See the full description on the dataset page: https://huggingface.co/datasets/abatilo/sudokubench.text1M<n<10M1 likes180 downloads1y agoHugging Face17binhng /robocasa-30and100demos-6chosen-tasks-for-aBao0 likes164 downloads9mo agoHugging Face18abacusai /MetaMath_DPO_FewShot Dataset Card for "MetaMath_DPO_FewShot" GSM8K \citep{cobbe2021training} is a dataset of diverse grade school maths word problems, which has been commonly adopted as a measure of the math and reasoning skills of LLMs. The MetaMath dataset is an extension of the training set of GSM8K using data augmentation. It is partitioned into queries and responses, where the query is a question involving mathematical calculation or reasoning, and the response is a logical series of steps and… See the full description on the dataset page: https://huggingface.co/datasets/abacusai/MetaMath_DPO_FewShot.text100K<n<1M28 likes160 downloads3y agoHugging Face19open-llm-leaderboard-old /details_abacusai__Smaug-2-72B Dataset Card for Evaluation run of abacusai/Smaug-2-72B Dataset automatically created during the evaluation run of model abacusai/Smaug-2-72B on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_abacusai__Smaug-2-72B.0 likes159 downloads3y agoHugging Face20mstz /abalone Abalone The Abalone dataset from the UCI ML repository. Predict the age of the given abalone. Configurations and tasks Configuration Task Description abalone Regression Predict the age of the abalone. binary Binary classification Does the abalone have more than 9 rings? Usage from datasets import load_dataset dataset = load_dataset("mstz/abalone")["train"] Features Target feature in bold. Feature Type sex [string]… See the full description on the dataset page: https://huggingface.co/datasets/mstz/abalone.tabulartabular-classification1K<n<10K1 likes156 downloads1y agoHugging Face21abadawi /MentalBench-Align MentalBench–100k & MentalAlign–70k: Dual Benchmark Suite for Mental Health LLM Evaluation 📄 Paper (arXiv): When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation📎 Paper Link: https://arxiv.org/pdf/2510.19032 📦 Code & Documentation: https://github.com/abeerbadawi/MentalBench-Align 📘 Overview This repository introduces two complementary datasets that enable systematic evaluation of large language models (LLMs) in… See the full description on the dataset page: https://huggingface.co/datasets/abadawi/MentalBench-Align.text10K<n<100K5 likes156 downloads1y agoHugging Face22abacusai /MetaMathFewshot A few-shot version of the MetaMath (https://huggingface.co/datasets/meta-math/MetaMathQA) dataset. Each entry is formatted with 'question' and 'answer' keys. The 'question' key has a random number of query-answer pairs between 0 and 4 inclusive, before a final target query; the expected answer to this is stored in the content of 'answer'. text100K<n<1M28 likes148 downloads3y agoHugging Face23abayb /parameter-golf-sp40960 likes143 downloads6mo agoHugging Face24yg3191 /AbAssayBench AbAssayBench This dataset repository contains the processed data package for AbAssayBench, a multi-endpoint antibody developability benchmark. The release combines the FLAb2.0-derived antibody measurements used for model development with the PROPHET-Ab measurements used for external validation. The repository is intended to be used together with the MAP-Ab source code: https://github.com/gu-yaowen/MAP-Ab. Package layout Path Contents tables/ Release… See the full description on the dataset page: https://huggingface.co/datasets/yg3191/AbAssayBench.imagetabular-regression100K<n<1M1 likes136 downloads2mo agoHugging Face25open-llm-leaderboard-old /details_AbacusResearch__Jallabi-34B Dataset Card for Evaluation run of AbacusResearch/Jallabi-34B Dataset automatically created during the evaluation run of model AbacusResearch/Jallabi-34B on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_AbacusResearch__Jallabi-34B.0 likes135 downloads3y agoHugging Face26abaryan /ham10000_bbox HAM10000 with Spatial Annotations and Bounding Box Coordinates Enhanced version of HAM10000 dataset with bounding box coordinates and spatial descriptions for skin lesion localization. Dataset Description This dataset extends the original HAM10000 dermatology dataset with: Bounding box coordinates for lesion localization Spatial descriptions (e.g., "located in center-center region") Area coverage statistics Mask availability flags Features image: RGB skin… See the full description on the dataset page: https://huggingface.co/datasets/abaryan/ham10000_bbox.imageimage-classification10K<n<100K0 likes133 downloads1y agoHugging Face27open-llm-leaderboard-old /details_abacusai__Llama-3-Giraffe-70B-Instruct0 likes125 downloads2y agoHugging Face28nyu-dice-lab /lm-eval-results-AbacusResearch-jaLLAbi2-7b-private Dataset Card for Evaluation run of AbacusResearch/jaLLAbi2-7b Dataset automatically created during the evaluation run of model AbacusResearch/jaLLAbi2-7b The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-AbacusResearch-jaLLAbi2-7b-private.tabular100K<n<1M0 likes123 downloads2y agoHugging Face29open-llm-leaderboard-old /details_abacusai__Llama-3-Giraffe-70B0 likes122 downloads2y agoHugging Face30gheero-Leyu /leyu-amharic-addis-ababa-dialect Leyu Amharic - Addis Ababa Dialect Speech Corpus Dataset Description This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Addis Ababa dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent… See the full description on the dataset page: https://huggingface.co/datasets/gheero-Leyu/leyu-amharic-addis-ababa-dialect.audioautomatic-speech-recognition1K<n<10K1 likes122 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.