Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01infinite-dataset-hub /FINDER_API_KEY_AI_SEARCH_2023 FINDER_API_KEY_AI_SEARCH_2023 tags: data collection, machine learning, API performance Note: This is an AI-generated dataset so its content may be inaccurate or false Dataset Description: The 'FINDER_API_KEY_AI_SEARCH_2023' dataset is designed to collect and analyze data from various AI search engines and their associated API performance metrics. The dataset focuses on the effectiveness of API key-based access in enhancing the search capabilities of AI systems and includes a… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/FINDER_API_KEY_AI_SEARCH_2023.tabularn<1K0 likes963 downloads2y agoHugging Face02lmarena-ai /search-arena-24k Overview This dataset contains ALL in-the-wild conversation crowdsourced from Search Arena between March 18, 2025 and May 8, 2025. It includes 24,069 multi-turn conversations with search-LLMs across diverse intents, languages, and topics—alongside 12,652 human preference votes. The dataset spans approximately 11,000 users across 136 countries, 13 publicly released models, around 90 languages (including 11% multilingual prompts), and over 5,000 multi-turn sessions. While user… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/search-arena-24k.text10K<n<100K42 likes766 downloads7mo agoHugging Face03spectralbranding /r15-ai-search-metamerism R15: AI Search Metamerism — Cross-Cultural Brand Perception Dataset Citation: Zharnikov, D. (2026v) | DOI: 10.5281/zenodo.19422427 | Version: v3.2.0 Dataset Summary This dataset contains the full session logs, aggregated results, and analysis outputs from the R15 large-scale experiment testing whether Large Language Models systematically collapse multi-dimensional brand perception into Economic and Experiential dimensions ("spectral metamerism"). It comprises 21… See the full description on the dataset page: https://huggingface.co/datasets/spectralbranding/r15-ai-search-metamerism.tabulartext-generationn<1K0 likes355 downloads3mo agoHugging Face04uw-math-ai /theorem-search-dataset Theorem Search Dataset The largest open corpus of informal mathematical theorems: 1,341,083 theorem statements with natural-language slogans from 209,777 papers, designed for semantic theorem retrieval. Paper: Semantic Search over 9 Million Mathematical Theorems Demo: huggingface.co/spaces/uw-math-ai/theorem-search Benchmark results On 110 test queries written by research mathematicians, our best pipeline (Qwen3-Embedding-8B on DeepSeek-V3.1 slogans) outperforms… See the full description on the dataset page: https://huggingface.co/datasets/uw-math-ai/theorem-search-dataset.tabularquestion-answering1M<n<10M26 likes262 downloads9d agoHugging Face05NM268 /multimodal-ai-search-data Multimodal AI Search: precomputed index Artifacts loaded by the Multimodal AI Search Space (source), built with scripts/build_index.py. file contents index.json model name, image filenames, captions, caption → image index image_embeddings.npy float16, L2-normalised image embeddings, one row per photo caption_embeddings.npy float16, L2-normalised caption embeddings, one row per caption projection_2d.npy 2-D UMAP coordinates (photos, then one caption per photo)… See the full description on the dataset page: https://huggingface.co/datasets/NM268/multimodal-ai-search-data.image1K<n<10K0 likes200 downloads7d agoHugging Face06lmarena-ai /search-arena-v1-7k Overview This dataset contains 7k leaderboard conversation votes collected from Search Arena between March 18, 2025 and April 13, 2025. All entries have been redacted for PII and sensitive user information to ensure privacy. Each data point includes: Two model responses (messages_a and messages_b) The human vote result A timestamp Full system metadata, LLM + web search trace, and post-processed metadata for controlled experiments (conv_meta) To reproduce the leaderboard results… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/search-arena-v1-7k.tabular1K<n<10K28 likes163 downloads1y agoHugging Face07uw-math-ai /theorem-search-dataset-permissive Theorem Search Dataset The largest open corpus of informal mathematical theorems: 1,239,720 theorem statements with natural-language slogans from 197,889 papers, designed for semantic theorem retrieval. Paper: Semantic Search over 9 Million Mathematical Theorems Demo: huggingface.co/spaces/uw-math-ai/theorem-search Benchmark results On 110 test queries written by research mathematicians, our best pipeline (Qwen3-Embedding-8B on DeepSeek-V3.1 slogans) outperforms… See the full description on the dataset page: https://huggingface.co/datasets/uw-math-ai/theorem-search-dataset-permissive.textquestion-answering1M<n<10M0 likes93 downloads9d agoHugging Face08gentian-hajdaraj /ai-search-visibility-italia AI Search Visibility in Italia 2026 A public research dataset from Telescop Research examining how brands and sources appear across ChatGPT, Gemini, Google AI Mode and Google AI Overview for a frozen panel of Italian queries. Dataset summary 500 frozen Italian queries 480 main non-branded queries 20 brand-seeded controls 10 sectors 4 AI surfaces 3 complete and separate runs 6,000 final observations 529 commercial brand families identified Collection window: 3–5… See the full description on the dataset page: https://huggingface.co/datasets/gentian-hajdaraj/ai-search-visibility-italia.tabularn<1K1 likes93 downloads17d agoHugging Face09Ionio-ai /ecommerce-search-extraction Ionio E-commerce Search Query Extraction Built with: simula — schema-driven synthetic data generation with auditable taxonomy lineage. An English synthetic dataset for training and evaluating systems that translate natural-language shopping requests into narrow, atomic, database-queryable JSON. It contains 10,985 accepted examples from a 13,000-attempt generation run. No accepted rows were trimmed from this release. Each example pairs a realistic typed or spoken shopper query… See the full description on the dataset page: https://huggingface.co/datasets/Ionio-ai/ecommerce-search-extraction.texttext-generation10K<n<100K2 likes63 downloads2mo agoHugging Face10Broadcastwell /ai-search-visibility-prices-2026-09 What AI search visibility costs in 2026: prices, guarantees and checkout paths of 123 sellers Broadcastwell Research, Study 2, September 2026. Version 1.0. DOI 10.5281/zenodo.22982430. License CC BY 4.0. What this is Four CSV files behind the report "What AI search visibility costs in 2026". They are cut from Broadcastwell Research's September 2026 seller-price survey: 137 seller profiles read live on 25 Sep 2026. Every price, guarantee, checkout path and free… See the full description on the dataset page: https://huggingface.co/datasets/Broadcastwell/ai-search-visibility-prices-2026-09.text0 likes60 downloads14d agoHugging Face11MTS-AI-SearchSkill /MTSBerquadMTSBerquad is a cleaned and enriched dataset SberQuAD transferred to the Generative QA task. All entities were truecased, refactored by hand to improve readability and consistency. Answers have been expanded and rearranged from MLM QA task to Generative/Long Form QA task. MTSBerquad presented in PyCon 2024 by MTS AI Search Group. Developed by MTS AI Search Group (Krayko Nikita, Laputin Fedor, Sidorov Ivan) textquestion-answering10K<n<100K8 likes50 downloads2y agoHugging Face12deeplumen /deeplumen-ai-search-visibility-benchmarks DeepLumen AI Search Visibility Benchmarks Clean, source-linked benchmark data derived from public DeepLumen case study pages. The dataset is intended as a publish-ready starter package for GitHub and Hugging Face, with each metric tied back to an editorially relevant canonical source. Dataset Files data/ai_search_visibility_benchmarks.json and .csv: one row per public case study. data/metric_observations.json and .csv: one row per measured benchmark observation.… See the full description on the dataset page: https://huggingface.co/datasets/deeplumen/deeplumen-ai-search-visibility-benchmarks.tabular-classificationn<1K0 likes49 downloads4mo agoHugging Face13Ecomrules /ai-search-visibility-audit-of-100-swiss-smbs-ti AI Search Visibility Audit of 100 Swiss SMBs (Ticino, 2026) A reproducible dataset measuring the technical readiness of Swiss SMBs in the Ticino canton for AI-powered search engines (ChatGPT, Perplexity, Google AI Overviews). Background In May 2026, Uberall published a study showing 83% of Italian restaurants are completely invisible to AI-generated search recommendations. This dataset replicates the methodology on 100 Swiss SMBs in the Ticino canton… See the full description on the dataset page: https://huggingface.co/datasets/Ecomrules/ai-search-visibility-audit-of-100-swiss-smbs-ti.0 likes47 downloads4mo agoHugging Face14AI-BRIDGES /wikidata-search-tracesWikidata-Search-Traces is an open corpus of 10,235 reasoning trajectories over Wikidata, developed by Pleias as part of AI-BRIDGES, with funding from Wikimedia Switzerland and in close collaboration with Wikimedia Deutschland.In each trajectory, an agent answers a question by writing Python in a REPL: it searches entities, follows relations, reads qualifiers and references, keeps intermediate results in variables, and returns an answer checked exactly against the graph. Wikidata describes… See the full description on the dataset page: https://huggingface.co/datasets/AI-BRIDGES/wikidata-search-traces.tabularquestion-answering10K<n<100K0 likes47 downloads1d agoHugging Face15arcee-ai /bfcl_v4_web_searchtextn<1K6 likes43 downloads1y agoHugging Face16quotientai /natural-qa-random-67-with-AI-search-answers Dataset Details Dataset Description This dataset is a refined subset of the "Natural Questions" dataset, filtered to include only high-quality answers as labeled manually. The dataset includes ground truth examples of "good" answers, defined as responses that are correct, clear, and sufficient for the given questions. Additionally, answers generated by three AI search engines (Perplexity, Gemini, Exa AI) have been incorporated to provide both raw and parsed outputs for… See the full description on the dataset page: https://huggingface.co/datasets/quotientai/natural-qa-random-67-with-AI-search-answers.textn<1K0 likes39 downloads2y agoHugging Face17ai-shift /ameba_faq_search AMEBA Blog FAQ Search Dataset This data was obtained by crawling this website. The FAQ Data was processed to remove HTML tags and other formatting after crawling, and entries containing excessively long content were excluded. The Query Data was generated using a Large Language Model (LLM). Please refer to the following blog for information about the generation process. https://www.ai-shift.co.jp/techblog/3710 https://www.ai-shift.co.jp/techblog/3761 Column description… See the full description on the dataset page: https://huggingface.co/datasets/ai-shift/ameba_faq_search.textquestion-answering1K<n<10K6 likes37 downloads3y agoHugging Face18WebSEM-ai /ai-search-visibility-romania-electronics-market AI Search Visibility — Romania's Electronics & IT Market (August 2026) 18 brand-free purchase questions × 5 AI engines = 87 answers. 86 of them name a major retailer. Position, not presence, decides the market. Raw data CC BY 4.0. Canonical study (analysis, charts, interpretation): Romanian · English What this is Eighteen real purchase questions were put to ChatGPT, Google Gemini, Perplexity, Google AI Mode and Google AI Overviews, in Romanian, from Romania, in… See the full description on the dataset page: https://huggingface.co/datasets/WebSEM-ai/ai-search-visibility-romania-electronics-market.tabularn<1K0 likes35 downloads2mo agoHugging Face19ai-colombia /job-searcher-data AI Job Searcher Training Data (V3) Fine-tuning dataset for a career advisor AI specializing in Nordic and European job markets. V3 (March 2026): Added 150 freeform/narrative-style Analyze examples modeled after real job postings from finn.no and arbeidsplassen.nav.no. Now includes startup-style, agency, generalist, and narrative formats alongside V2 rigid-template examples. Dataset Description This dataset contains 1,190 training examples across 9 languages and 5… See the full description on the dataset page: https://huggingface.co/datasets/ai-colombia/job-searcher-data.texttext-generation1K<n<10K0 likes34 downloads7mo agoHugging Face20313Automations /geo-checklist-ai-search GEO Checklist for AI Search A practical 20-point checklist for GEO (generative engine optimization): the work that makes a business easy for ChatGPT, Perplexity, Gemini and Google AI Overviews to find, understand and cite. Each check explains why it matters, how to verify it, its priority and the effort involved. Curated by 313 Automations, an AI automation, web and SEO/GEO agency in Lahore, Pakistan (founded 2018). Use it interactively in the free GEO Readiness Checklist Space.… See the full description on the dataset page: https://huggingface.co/datasets/313Automations/geo-checklist-ai-search.texttext-classificationn<1K0 likes34 downloads7d agoHugging Face21Inabia-AI /image_searchingimagen<1K0 likes33 downloads2y agoHugging Face22rubiks-ai /SearchBench-Evaltextn<1K1 likes28 downloads2y agoHugging Face23DeepNLP /ai-search-agent AI Search Agent Agent Meta and Traffic Dataset in AI Agent Marketplace | AI Agent Directory | AI Agent Index from DeepNLP This dataset is collected from AI Agent Marketplace Index and Directory at http://www.deepnlp.org, which contains AI Agents's meta information such as agent's name, website, description, as well as the monthly updated Web performance metrics, including Google,Bing average search ranking positions, Github Stars, Arxiv References, etc. The dataset is helpful for AI… See the full description on the dataset page: https://huggingface.co/datasets/DeepNLP/ai-search-agent.texttext-generationn<1K1 likes25 downloads2y agoHugging Face24ai-colombia /ai-job-searcher-training-data AI Job Searcher Training Data Fine-tuning dataset for a career advisor AI specializing in Nordic and European job markets. Dataset Description This dataset contains 1,040 training examples across 9 languages and 5 task categories, formatted as chat conversations (system/user/assistant) suitable for fine-tuning LLMs. Task Categories Category Examples Description Cover Letter Generation 208 Professional cover letters from job description + user profile… See the full description on the dataset page: https://huggingface.co/datasets/ai-colombia/ai-job-searcher-training-data.texttext-generation1K<n<10K0 likes25 downloads7mo agoHugging Face25DeepNLP /search-recommendation-ai-agent Search Recommendation Agent Meta and Traffic Dataset in AI Agent Marketplace | AI Agent Directory | AI Agent Index from DeepNLP This dataset is collected from AI Agent Marketplace Index and Directory at http://www.deepnlp.org, which contains AI Agents's meta information such as agent's name, website, description, as well as the monthly updated Web performance metrics, including Google,Bing average search ranking positions, Github Stars, Arxiv References, etc. The dataset is helpful… See the full description on the dataset page: https://huggingface.co/datasets/DeepNLP/search-recommendation-ai-agent.textn<1K0 likes23 downloads2y agoHugging Face26Oni1xyz /ai-that-works-search-embeddings AI That Works Search Embeddings This private review dataset contains 2,224 normalized text embeddings for the AI That Works search project. Source attribution AI That Works is hosted by Vaibhav Gupta (@hellovai) and Dex Horthy (@dexhorthy). This independent dataset is derived from their show and its companion materials; all underlying show content remains credited to its creators and contributors. Official show repository:… See the full description on the dataset page: https://huggingface.co/datasets/Oni1xyz/ai-that-works-search-embeddings.textfeature-extraction1K<n<10K0 likes23 downloads1d agoHugging Face27roshbeed /ai-residency-vector-search-retrieval-datatextn<1K0 likes21 downloads2mo agoHugging Face28OSAS-AI /fuzzy_search_datasettext10K<n<100K0 likes18 downloads3mo agoHugging Face29WebSEM-ai /ai-search-visibility-romania-book-market AI Search Visibility — Romania's Book Market (July 2026) When someone asks ChatGPT "which online bookstore should I use for children's books?", they get one answer, not ten blue links. This dataset measures who is inside that answer — and who merely feeds it. Canonical study (analysis, charts, interpretation): Romanian · English What this is Eighteen real purchase questions were put to five AI engines — ChatGPT, Google Gemini, Perplexity, Google AI Mode, Google AI… See the full description on the dataset page: https://huggingface.co/datasets/WebSEM-ai/ai-search-visibility-romania-book-market.tabularn<1K0 likes18 downloads2mo agoHugging Face30ClarusC64 /cascade-ai-adversarial-search-simulator-v0.2 Clarus Adversarial Cascade Simulator v0.2 Adversarial boundary discovery for cascade-prone system configurations. You provide a configuration.The simulator maps how close it is to systemic collapse. Interactive Demo Live Gradio interface available in Hugging Face Spaces. Workflow: Input baseline configuration (6 sliders) Score configuration → View risk assessment Run adversarial search → Discover worst-case boundary states View scenario pack → Executable sandbox… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/cascade-ai-adversarial-search-simulator-v0.2.tabulartext-classificationn<1K1 likes16 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.