Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Anthropic /values-in-the-wild Summary This dataset presents a comprehensive taxonomy of 3307 values expressed by Claude (an AI assistant) across hundreds of thousands of real-world conversations. Using a novel privacy-preserving methodology, these values were extracted and classified without human reviewers accessing any conversation content. The dataset reveals patterns in how AI systems express values "in the wild" when interacting with diverse users and tasks. We're releasing this resource to advance research… See the full description on the dataset page: https://huggingface.co/datasets/Anthropic/values-in-the-wild.tabular1K<n<10K156 likes445 downloads1y agoHugging Face02shivank21 /mmconflict-editable-values-1k MMConflict Editable Values 2K This dataset contains 2,000 source images with visible atomic values for multimodal conflict research. It has 100 images in each of 20 categories. Every image comes from a photograph, scan, captured website, software screenshot, or page of a source document. The dataset does not contain generated images or project-rendered examples. Each row records the source, source URL, license, attribution, visible value, question, and a candidate box around the… See the full description on the dataset page: https://huggingface.co/datasets/shivank21/mmconflict-editable-values-1k.imageimage-to-text1K<n<10K0 likes376 downloads1mo agoHugging Face03anicola /value-systems-in-llms-paraphrasing-and-profile-elicitation Value Systems in LLMs: Effects of Paraphrasing and Profile Elicitation on Decision-Making Consistency and Robustness (Versión en español más abajo.) Do large language models give stable answers to the same forced-choice question when the prompt is perturbed in ways that do not change its meaning — and does assigning them a personality or value profile change those answers? This dataset contains the full material of that experiment: the 9,350 prompts, the 561,000 model responses… See the full description on the dataset page: https://huggingface.co/datasets/anicola/value-systems-in-llms-paraphrasing-and-profile-elicitation.tabularmultiple-choice100K<n<1M0 likes142 downloads2mo agoHugging Face04dougalldeepmind /2026-09-17-da-lowstakes-values-in-advice-synth-smoke 18-row constitution-only low-stakes smoke; FAIL scaling gate; diagnostic candidates only field value experiment 18-row constitution-only low-stakes smoke; FAIL scaling gate; diagnostic candidates only date_generated 20260917_171453 constitution constitutions/claude_distilled_09_principles/constitution.md sha256 8e273b472d945aa23efa6236886da5e1171bff2193ee31ff73489ca54c4f0edc source_repo https://github.com/Matthew-Bozoukov/teaching_claude_why_replication.git @… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-17-da-lowstakes-values-in-advice-synth-smoke.textn<1K0 likes137 downloads23d agoHugging Face05DCDCgrupoG /spotify-moral-values-charts Spotify Top 100 Lyrics 2018-2024 (Bag of Words) Dataset Summary Corpus de 13.538 canciones únicas que estuvieron en el Top 100 diario de Spotify entre el 1-1-2018 y el 31-12-2024 en cinco listas: Global, Estados Unidos, Australia, España y Argentina. Cada fila es una canción, con la fecha de su primera aparición en el chart, métricas de popularidad y su letra representada como bolsa de palabras con stemming (stem:count). Las letras originales no se publican… See the full description on the dataset page: https://huggingface.co/datasets/DCDCgrupoG/spotify-moral-values-charts.tabulartext-classification1M<n<10M0 likes127 downloads6d agoHugging Face06rohinm /zip-training-hallucination-data-qwen06b-thinking-train-with-valuestabular10K<n<100K0 likes90 downloads1y agoHugging Face07fabianrausch /financial-entities-values-augmentedThis dataset is contains 200 sentences taken from German financial statements. In each sentence financial entities and financial values are annotated. Additionally there is an augmented version of this dataset where the financial entities in each sentence have been replaced by several other financial entities which are hardly/not covered in the original dataset. The augmented version consists of 7287 sentences. textn<1K2 likes67 downloads4y agoHugging Face08oxford-llms /world_values_survey_2017_2022_sfttabular100K<n<1M1 likes65 downloads2y agoHugging Face09luxury-lakehouse /spadl-vaep-action-values SPADL/VAEP Action Values Every on-ball action from ~9.5 million professional soccer events, converted to the SPADL unified format and scored with offensive, defensive, and net VAEP values. Built with the silly-kicks library — enabling player ranking by total contribution beyond goals and assists. Part of the (Right! Luxury!) Lakehouse soccer analytics platform. ⚠️ Schema change (cut-over 2026-07-22) This dataset now emits both legacy and canonical Kimball key… See the full description on the dataset page: https://huggingface.co/datasets/luxury-lakehouse/spadl-vaep-action-values.tabulartabular-classification1M<n<10M0 likes65 downloads2mo agoHugging Face10valuesimplex-ai-lab /FIR-Bench-Multi-Docs-FinQAtext10K<n<100K0 likes62 downloads1y agoHugging Face11values-md /when-agents-act Dataset Card for "When Agents Act" Dataset Summary This dataset contains 702 ethical decision judgements from 9 frontier LLMs (Claude Opus 4.5, GPT-5, GPT-5 Nano, Claude Sonnet 4.5, Claude Haiku 4.5, Gemini 3 Pro, Gemini 2.5 Flash, Grok-4, Grok-4 Fast) across 10 rigorously curated AI-relevant ethical dilemmas. Models were tested in both theory mode (hypothetical reasoning) and action mode (tool-enabled agents believing actions would execute). Key Finding: Models reverse… See the full description on the dataset page: https://huggingface.co/datasets/values-md/when-agents-act.tabulartext-classificationn<1K1 likes50 downloads11mo agoHugging Face12llm-lab /MENA_VALUES_Benchmarktabularn<1K0 likes41 downloads1y agoHugging Face13windfromthenorth /crafter-valuestext100K<n<1M0 likes40 downloads1y agoHugging Face14HC-85 /food-nutritional-valuesOriginally webscrapped by Aleksandr Antonov from Nutrition Value and posted on Kaggle as "Nutritional values for common foods and products". tabular10K<n<100K3 likes38 downloads2y agoHugging Face15richardcieplechowicz /tampa-bay-owner-home-values-zcta-2020-2024 Tampa Bay owner-occupied home values by ZCTA, 2020-2024 Author: Richard (Ryszard) Cieplechowicz. Study page: https://richardcieplechowicz.com/tampa-bay-owner-home-values-by-zcta-2020-2024/ Tampa Bay owner-occupied home values by ZCTA, 2020-2024 A second data cut by Richard Cieplechowicz: the distribution of self-reported values for owner-occupied homes across 132 selected ZCTAs assigned to four Tampa Bay counties. This is not a sale-price series or a listing-price survey. One… See the full description on the dataset page: https://huggingface.co/datasets/richardcieplechowicz/tampa-bay-owner-home-values-zcta-2020-2024.tabularn<1K0 likes38 downloads7d agoHugging Face16valuesimplex-ai-lab /FinCPRG FinCPRG Dataset 近年来,大语言模型(LLMs)在构建段落检索数据集方面展现出了巨大的潜力。然而,现有方法在表达跨文档查询需求和控制标注质量方面仍存在局限性。为了解决这些问题,本文提出了一个双向生成管道,旨在为文档内和跨文档场景生成3级层次化查询,并在直接映射标注的基础上挖掘额外的相关性标签。使用这个管道,我们从近1.3k份中文金融研究报告中构建了金融段落检索生成数据集(FinCPRG),该数据集包含层次化查询和丰富的相关性标签。 数据集结构 (Dataset Structure) 数据集遵循标准的段落检索格式,包含以下文件: corpus.jsonl: 包含段落(文档)语料。每行是一个JSON对象,包含 _id (段落ID) 和 text (段落文本) 字段。 queries.jsonl: 包含所有生成的查询。每行是一个JSON对象,包含 _id (查询ID) 和 text (查询文本) 字段。 qrels/: 包含查询-段落相关性标注文件 (qrels),格式为TSV (query-id \t corpus-id \t… See the full description on the dataset page: https://huggingface.co/datasets/valuesimplex-ai-lab/FinCPRG.text100K<n<1M0 likes36 downloads1y agoHugging Face17Baidicoot /simulators-political-valuestext10K<n<100K1 likes35 downloads1y agoHugging Face18valuesimplex-ai-lab /FIR-Bench-Sin-Doc-FinQAtext1K<n<10K0 likes35 downloads1y agoHugging Face19kali-ai /ai-values ai-values This dataset is made by Kali AI for you to train your AI models, Issues? Simply submit a PR to the Community section Dataset { "friendliness": 5.6, "emoji_usage": { "casual": 0.4, "non_casual": 0.05 }, "auto": "Act only after an explicit user request and when decisiveness is true.", "casual_messages": "auto", "clean_messages": "auto", "roleplay_mode": "auto", "decisions": { "true": 1, "false": 0 }, "decisive": "Is… See the full description on the dataset page: https://huggingface.co/datasets/kali-ai/ai-values.texttext-generationn<1K0 likes35 downloads8mo agoHugging Face20llm-values /self_response_2_deepseek_ai__deepseek_llm_67b_base_answerstabularn<1K0 likes32 downloads2y agoHugging Face21utk6 /anthropic-rlhf-human-values-dpotabular10K<n<100K0 likes32 downloads2y agoHugging Face22llm-values /self_response_deepseek_ai__deepseek_llm_67b_base_answerstabularn<1K0 likes31 downloads2y agoHugging Face23valuesimplex-ai-lab /FIR-Bench-Research-Reports-FinQAtext1M<n<10M0 likes30 downloads1y agoHugging Face24luxury-lakehouse /space-creation-values Space Creation Values — ELASTIC/OBSO Per-player per-frame space creation quantification — measuring each player's contribution to off-ball scoring opportunities via differential OBSO. For every sampled frame, the model computes OBSO with and without each player, yielding the area of scoring opportunity that player creates (or destroys) by their positioning. Part of the (Right! Luxury!) Lakehouse soccer analytics platform. Quick Start from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/luxury-lakehouse/space-creation-values.tabulartabular-regression100K<n<1M0 likes30 downloads5mo agoHugging Face25windfromthenorth /crafter-values-splittabular100K<n<1M0 likes28 downloads1y agoHugging Face26Baidicoot /simulators-political-values-centertext1K<n<10K0 likes27 downloads1y agoHugging Face27llm-values /self_response_Qwen_cont_Qwen1.5_32B_answerstabularn<1K0 likes24 downloads2y agoHugging Face28Baidicoot /simulators-political-values-center-righttext1K<n<10K0 likes22 downloads1y agoHugging Face29snap-stanford /reddit_baseline_gpt-5-mini-2025-08-07_persona_valuestext1K<n<10K0 likes22 downloads10mo agoHugging Face30xiaoqingsun004 /values_evals Dataset Details Eval for LLM values. Each query, framed as a user query, is an implicit tradeoff between value1 and value2 (with an optional bias towards one value). For each value1 and value2, there is a spectrum of responses biased away from (0) or towards (6) that value. original_eval is taken from jifanz/stress_testing_model_spec and reproduced here. We use the value mapping from Anthropic/values-in-the-wild to map fine-grained values to value1 and value2. new_eval is our… See the full description on the dataset page: https://huggingface.co/datasets/xiaoqingsun004/values_evals.text10K<n<100K0 likes22 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.