Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01NoeFlandre /benchmark-llms-landuse-relevance Land-use relevance benchmark v3-multilingual · 85 languages x 300 items/language · 25,500 items · binary yes/no labels. Code Package version recorded in run metadata: 0.2.0 (some runs lack version metadata). Task and prompt Does a sentence describe a place's land or environment in ways visible to satellites? English prompt · greedy decoding · seed 0 · max_new_tokens=4096 · bfloat16 · batch varies by model. unsloth/Qwen3.8-27B-GGUF@UD-IQ2_XXS runs the UD-IQ2_XXS… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/benchmark-llms-landuse-relevance.tabulartext-classification10K<n<100K0 likes3.2k downloads13d agoHugging Face02bench-llms /or-bench OR-Bench: An Over-Refusal Benchmark for Large Language Models Please see our demo at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench.imagetext-generation10K<n<100K1 likes925 downloads2y agoHugging Face03bench-llms /or-bench-toxic-all OR-Bench: An Over-Refusal Benchmark for Large Language Models This dataset constains highly toxic prompts, use with caution!!! Please see our demo at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench-toxic-all.imagetext-generation10K<n<100K1 likes475 downloads2y agoHugging Face04neemiasbsilva /multimodal-LLMs-See-Sentiment MLLMsent — datasets and experiment results Every input and every output of "Multimodal LLMs See Sentiment" (arXiv:2508.16873): the image descriptions generated by six multimodal LLMs, the sentiment labels derived from the PerceptSent annotations, and the complete per-fold results of all 141 experiments. Paper: arXiv:2508.16873 Code, training and inference: https://github.com/neemiasbsilva/multimodal-LLMs-see-sentiment Model checkpoints:… See the full description on the dataset page: https://huggingface.co/datasets/neemiasbsilva/multimodal-LLMs-See-Sentiment.texttext-classification10K<n<100K1 likes305 downloads2mo agoHugging Face05davron04 /llm_stockstabular10M<n<100M1 likes192 downloads10mo agoHugging Face06tum-nlp /cognitive-biases-in-llms A Comprehensive Evaluation of Cognitive Biases in LLMs: Dataset Dataset for evaluating cognitive biases in large language models 1. Dataset Card Overview A tabular dataset for measuring the presence and strength of cognitive biases in large language models (LLMs), introduced in the paper “A Comprehensive Evaluation of Cognitive Biases in LLMs” by Malberg et al. Paper | Code This dataset is intended only for the evaluation of LLMs and not to be used for… See the full description on the dataset page: https://huggingface.co/datasets/tum-nlp/cognitive-biases-in-llms.text10K<n<100K1 likes161 downloads1y agoHugging Face07oxford-llms /ai-respondents-challenge AI Respondents Challenge — Oxford LLMs 2026 Predict a survey respondent's answer to a held-out question from their other answers (World Values Survey wave 7). Any method allowed; you must disclose the features and prompts you used. Ranked on normalized skill + distributional alignment, on in-domain and out-of-domain (held-out countries) boards. Configs train — 5,000 labeled respondents (100 per seen country): respondent_id, country + all WVS variables… See the full description on the dataset page: https://huggingface.co/datasets/oxford-llms/ai-respondents-challenge.tabular1K<n<10K1 likes150 downloads3mo agoHugging Face08sempite /llmstxt-corpus The llms.txt corpus Measurement data on the llms.txt convention, collected in one run on 5 August 2026. llms.txt is a plain-text file at a site's root, proposed as a curated map telling AI systems what the site contains. This is a measurement of what is actually being published under that name. Canonical release: https://doi.org/10.5281/zenodo.22859104 This repository mirrors that deposit. Cite the DOI, which always resolves to the newest version. Two observations… See the full description on the dataset page: https://huggingface.co/datasets/sempite/llmstxt-corpus.tabular10K<n<100K0 likes132 downloads20d agoHugging Face09mznaser /moral-tracing-in-LLMs LLM Moral Evolution Study A longitudinal dataset tracking moral reasoning patterns across 14 large language models from OpenAI and Anthropic, spanning multiple generations (2023–2025). The dataset measures how moral stances, ethical judgments, and value priorities shift across model updates using a 107-item probe instrument grounded in Moral Foundations Theory. Models OpenAI Model Release GPT-3.5 Turbo 2023-11 GPT-4 2023-03 GPT-4o… See the full description on the dataset page: https://huggingface.co/datasets/mznaser/moral-tracing-in-LLMs.documenttext-classification10K<n<100K0 likes89 downloads4mo agoHugging Face10danilocorsi /LLMs-Sentiment-Augmented-Bitcoin-Dataset Leveraging LLMs for Informed Bitcoin Trading Decisions: Prompting with Social and News Data Reveals Promising Predictive Abilities The work was carried out by: Danilo Corsi Cesare Campagnano Description This project investigates the potential of leveraging Large Language Models (LLMs) to support Bitcoin traders. Specifically, we analyze the correlation between Bitcoin price movements and sentiment expressed in news headlines, posts, and comments on social media. We… See the full description on the dataset page: https://huggingface.co/datasets/danilocorsi/LLMs-Sentiment-Augmented-Bitcoin-Dataset.tabulartext-classification10K<n<100K8 likes87 downloads2y agoHugging Face111-800-LLMs /piqatext10K<n<100K0 likes78 downloads1y agoHugging Face121-800-LLMs /go_emotionstext10K<n<100K0 likes57 downloads1y agoHugging Face131-800-LLMs /Boardgame-QAtext10K<n<100K0 likes51 downloads1y agoHugging Face14llmspeed /llm-speed-benchmarks llm-speed: signed LLM inference-speed benchmarks Crowdsourced, cryptographically signed measurements of how fast large language models actually run: decode tokens per second, time to first token, and latency, across consumer GPUs, Apple Silicon, and hosted APIs, under one reproducible workload suite (suite-v1). Live data and bulk downloads: https://llm-speed.com/data Per-run permalink: https://llm-speed.com/r/<id> Methodology: https://llm-speed.com/methodology DOI:… See the full description on the dataset page: https://huggingface.co/datasets/llmspeed/llm-speed-benchmarks.textn<1K0 likes51 downloads3mo agoHugging Face15jhu-clsp /astro-llms-benchmark-dataset AstroLLMs Gold Benchmark Dataset This dataset is a collection of queries that astronomers asked to an astronomy research Slack chatbot. Along with the questions, there are open coding labels determined by a team of researchers and expert astronomer answers to these queries. Astronomers were asked to respond using citations and without the help of Large Language Models. This dataset of answers and responses is called the "Gold Benchmark Dataset". Dataset Structure The… See the full description on the dataset page: https://huggingface.co/datasets/jhu-clsp/astro-llms-benchmark-dataset.textn<1K1 likes48 downloads1y agoHugging Face16tarekmasryo /llm-system-ops-production-telemetry-sft-data 🤖📈 LLM System Ops Telemetry (Synthetic) A synthetic, production-style, multi-table LLM telemetry dataset designed for LLMOps analytics and decision-grade experiments. It supports monitoring cost, latency, tokens, failures, safety flags, tool usage, and user feedback at the interaction level, with rollups at the session and user levels — plus an SFT table aligned 1:1 with interactions and a prompt/config dimension. Synthetic data (safe for teaching, prototyping, and portfolio… See the full description on the dataset page: https://huggingface.co/datasets/tarekmasryo/llm-system-ops-production-telemetry-sft-data.tabulartabular-classification10K<n<100K1 likes47 downloads8mo agoHugging Face17ucberkeley-dlab /normative_evaluation_llms_everyday_dilemmastabular10K<n<100K2 likes45 downloads1y agoHugging Face18drozado /llms_epistemic_consistency LLMs Epistemic Consistency Dataset This dataset artifact contains the stimuli and prompt templates used for experiments on epistemic consistency and political-cue sensitivity in LLM evaluations. Dataset URL: https://huggingface.co/datasets/drozado/llms_epistemic_consistency Contents croissant.json: root-level copy of the completed Croissant metadata for NeurIPS 2026 Evaluations and Datasets submission. metadata/croissant.json: same Croissant metadata, kept with the… See the full description on the dataset page: https://huggingface.co/datasets/drozado/llms_epistemic_consistency.imagen<1K0 likes36 downloads5mo agoHugging Face191-800-LLMs /chemistrytext10K<n<100K0 likes34 downloads1y agoHugging Face201-800-LLMs /physicstext10K<n<100K2 likes32 downloads1y agoHugging Face21joylarkin /cleverhack-llms-txtDescription: cleverhack.com llms.txt file as a dataset Last Update: 17 May 2026 textn<1K2 likes32 downloads5mo agoHugging Face22aliarda /LLMs-Turkish-TEOG-Leaderboard TEOG Scores Leaderboard Welcome to the TEOG Scores Leaderboard! This repository contains the results of evaluating various large language models (LLMs) on the TEOG (Temel Eğitimden Ortaöğretime Geçiş) exam dataset. The TEOG exam is a standardized test in Turkey used for high school admissions, and this dataset provides a benchmark for assessing the performance of LLMs in Turkish educational tasks. Please remember that full score for TEOG is 500 points. More Models Are… See the full description on the dataset page: https://huggingface.co/datasets/aliarda/LLMs-Turkish-TEOG-Leaderboard.textn<1K2 likes31 downloads2y agoHugging Face231-800-LLMs /okapi_mmlutext100K<n<1M0 likes31 downloads1y agoHugging Face241-800-LLMs /okapi_arc_challengetext100K<n<1M0 likes30 downloads1y agoHugging Face25pixeloffice /llm-smartrouter-benchmark LLM SmartRouter & Agent Highway Latency & Cost Benchmark (v1.4.0) Empirical performance benchmark dataset comparing direct model endpoints (OpenAI, Anthropic Claude, Google Gemini) against the PixelRouter / BLUN SmartRouter proxy layer and Autonomous Agent Web Highway (https://api.pixeloffice.eu/v1). v1.4.0 Benchmark Highlights Anthropic Claude Messages API: Sub-35ms proxy routing for native /v1/messages payloads with 94%+ cost savings. Machine Web Highway… See the full description on the dataset page: https://huggingface.co/datasets/pixeloffice/llm-smartrouter-benchmark.tabulartext-generationn<1K0 likes30 downloads1mo agoHugging Face26somosnlp /LLM_SQL_BaseDatosEspanol Usos Usos directos El objetivo principal de este dataset es proporcionar ejemplos simples para el fine-tuning de modelos de procesamiento de lenguaje natural (NLP) en el contexto de consultas SQL. Usos fuera de mira Podria usarse para el entrenamiento de una IA que sirva como creadora de base de datos artificiales Estructura del conjunto de datos Question: Es la pegunta que el usuario le dara al chatbot Answer: La respuesta el que chatbot le… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp/LLM_SQL_BaseDatosEspanol.textquestion-answeringn<1K10 likes29 downloads2y agoHugging Face27dhgottesman /keen_estimating_knowledge_in_llmstext1K<n<10K0 likes29 downloads11mo agoHugging Face281-800-LLMs /indian-medicinesimage10K<n<100K1 likes26 downloads1y agoHugging Face29jhu-clsp /astro-llms-full-query-data AstroLLMs Full Query Dataset This dataset includes all of the data collected in a four-week deployment of a Large Language Model-powered Slack chatbot trained on astrophysics papers. Astronomers were invited to interact with the chatbot, ask questions, and leave feedback. This data includes 368 question-answer pairs, including feedback, reactions, and labeling. Dataset Structure The columns of this dataset are thread_ts (unique time stamp of the query), channel_id… See the full description on the dataset page: https://huggingface.co/datasets/jhu-clsp/astro-llms-full-query-data.tabularn<1K1 likes22 downloads1y agoHugging Face30Sanjay1905 /ascii_art_dataset_for_llmstext100K<n<1M1 likes22 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.