datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
polymarket-search-indexus-stocks-and-search-trends
US stocks and search trends
11636 symbols, one plain CSV per symbol and resolution.
import pandas as pd
df = pd.read_csv("hf://datasets/yadavsaurabh/us-stocks-and-search-trends/prices/1d/A.csv")
from datasets import load_dataset # every symbol of one resolution as one table
ds = load_dataset("yadavsaurabh/us-stocks-and-search-trends", "1d")
Prices (prices/<interval>/<SYMBOL>.csv)
Interval
Bars
History
1d
daily
full history (back to 1962 for… See the full description on the dataset page: https://huggingface.co/datasets/yadavsaurabh/us-stocks-and-search-trends.search-v3-embeddings
Hub Card Search Embeddings (v3)
One-sentence summaries and 1024-d embeddings for 1,173,030 dataset and model cards on the
Hugging Face Hub — 536,870 datasets and 636,160 models. It is the search index behind the revived
librarian-bots/huggingface-semantic-search
backend: you search over a short model-written summary of each card rather than the raw card, and
retrieve against the embedding of that summary.
The cards come from librarian-bots/dataset_cards_with_metadata
and… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/search-v3-embeddings.FINDER_API_KEY_AI_SEARCH_2023
FINDER_API_KEY_AI_SEARCH_2023
tags: data collection, machine learning, API performance
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'FINDER_API_KEY_AI_SEARCH_2023' dataset is designed to collect and analyze data from various AI search engines and their associated API performance metrics. The dataset focuses on the effectiveness of API key-based access in enhancing the search capabilities of AI systems and includes a… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/FINDER_API_KEY_AI_SEARCH_2023.rollouts-olmo7b-cue-search
rollouts-olmo7b-cue-search
Model: allenai/Olmo-3-1025-7B (snapshot a81bae42).
Tokenizer: allenai/Olmo-3-1025-7B (snapshot a81bae42).
Protocol: RL-Zero prompt, MATH-500 x 4 rollouts, budget 31,744, T 0.6, top-p 0.95, seed 20260819 (depth-2 exhaustive and n-gram chain: seed 20260821); the top-20 beam nominee screen, ten random-opener arms, every depth-2 opener (84 shards, arm names unique across shards) and the n-gram chain arms.
Rollouts generated on the CSAIL cluster for the… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-cues/rollouts-olmo7b-cue-search.hnm-search-data
HnM Search Dataset Created from Recommendations Dataset
This synthetic data-set is created using the recommendations dataset:
https://huggingface.co/datasets/einrafh/hnm-fashion-recommendations-data (Use of this dataset is subject to the terms and conditions set forth on the original distribution page. This dataset is intended for non-commercial and research use.)
https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/data (DATA ACCESS AND USE:… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/hnm-search-data.peak-search-300krest-graph-searchjob-search-replay-trial
Temporary job-search replay trial
Public job-document sample with controlled test updates. Not the production index.
split_search_qa
preprocessed_SearchQA
The SearchQA question-answer pairs originate from J! Archive2, which comprehensively archives all question-answer pairs
from the renowned television show Jeopardy! The passages, sourced from Google search web page snippets.
We offer passage metadata, encompassing details like 'air_date,' 'category,' 'value,' 'round,' and 'show_number,'
enabling you to enhance retrieval performance at your discretion.
Should you require further details about SearchQA, please… See the full description on the dataset page: https://huggingface.co/datasets/NomaDamas/split_search_qa.swellmeter-google-trending-searches
Swellmeter: what is trending on Google right now, in 30 countries
Today's trending searches plus 30 days of history, by API: $5 for 1,500 calls. Buy now → · API key on screen the moment checkout ends · no subscription · 14-day refund if it doesn't work as described · help: cybermax.tools@gmail.com
Need it fresh, filtered or via API? This free file is a snapshot (Google trending searches at the last refresh), last updated 2026-10-09.
Swellmeter Trending API (50 free calls a… See the full description on the dataset page: https://huggingface.co/datasets/CyberMax-tools/swellmeter-google-trending-searches.r15-ai-search-metamerism
R15: AI Search Metamerism — Cross-Cultural Brand Perception Dataset
Citation: Zharnikov, D. (2026v) | DOI: 10.5281/zenodo.19422427 | Version: v3.2.0
Dataset Summary
This dataset contains the full session logs, aggregated results, and analysis outputs from the R15 large-scale experiment testing whether Large Language Models systematically collapse multi-dimensional brand perception into Economic and Experiential dimensions ("spectral metamerism"). It comprises 21… See the full description on the dataset page: https://huggingface.co/datasets/spectralbranding/r15-ai-search-metamerism.code_x_glue_tc_nl_code_search_adv
Dataset Card for "code_x_glue_tc_nl_code_search_adv"
Dataset Summary
CodeXGLUE NL-code-search-Adv dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Text-Code/NL-code-search-Adv
The dataset we use comes from CodeSearchNet and we filter the dataset as the following:
Remove examples that codes cannot be parsed into an abstract syntax tree.
Remove examples that #tokens of documents is < 3 or >256
Remove examples that documents contain special tokens… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_tc_nl_code_search_adv.peak-search-content-70kMood_Based_Music_Search
Dataset Card for Acoustic_Features_From_Spotify
This dataset is created by merging multiple Kaggle datasets of Spotify audio features and track metadata into a unified, clean, and deduplicated collection of 2,577,667 unique tracks. Each record is indexed by Spotify track_id (with planned support for ISRC identifiers in future iterations).
Dataset Details
Dataset Description
Acoustic_Features_From_Spotify consolidates acoustic properties and… See the full description on the dataset page: https://huggingface.co/datasets/P-Arpan/Mood_Based_Music_Search.theorem-search-dataset
Theorem Search Dataset
The largest open corpus of informal mathematical theorems: 1,341,083 theorem statements with natural-language slogans from 209,777 papers, designed for semantic theorem retrieval.
Paper: Semantic Search over 9 Million Mathematical Theorems
Demo: huggingface.co/spaces/uw-math-ai/theorem-search
Benchmark results
On 110 test queries written by research mathematicians, our best pipeline (Qwen3-Embedding-8B on DeepSeek-V3.1 slogans) outperforms… See the full description on the dataset page: https://huggingface.co/datasets/uw-math-ai/theorem-search-dataset.superintelligence-search-questions
Keyfern: what people search about superintelligence (Sept 2026)
Google's trending searches for 30 countries with 30 days of history, by API: $5 for 1,500 calls. Buy now → · API key on screen the moment checkout ends · no subscription · 14-day refund if it doesn't work as described · help: cybermax.tools@gmail.com
Need it fresh, filtered or via API? This free file is a snapshot (search questions at the last refresh), last updated 2026-09-24.
Keyfern on Apify ($0.001 per… See the full description on the dataset page: https://huggingface.co/datasets/CyberMax-tools/superintelligence-search-questions.cabot-search
CaBot Clinical Literature Embedding Index
Exact embedding-search index over 3,474,244 works from 204 high-impact clinical
journals (2023 JIF >= 10), built from an OpenAlex snapshot (~June 2025) and
embedded with OpenAI text-embedding-3-small at 1536 dimensions (float32).
This is used for CaBot. The build and search code is in the CaBot-Search/
folder of the source repository.
Columns
column
type
notes
id
string
OpenAlex work id
doi
string
title… See the full description on the dataset page: https://huggingface.co/datasets/tbuckley/cabot-search.deepresearchgym-agentic-search-logs
DeepResearchGym Agentic Search Logs
This repository hosts the dataset accompanying the paper “Agentic Search in the Wild” (arXiv: https://arxiv.org/abs/2601.17617).
The dataset contains 14M+ search queries collected via DeepResearchGym (DRGym), an open-source search API designed for DeepResearch-style agentic search. For more background on DRGym, see: https://arxiv.org/abs/2505.19253.
All records have been anonymized and shuffled to prevent re-identification, and we additionally… See the full description on the dataset page: https://huggingface.co/datasets/cx-cmu/deepresearchgym-agentic-search-logs.amazon-product-searchPlease refer the original repository https://github.com/amazon-science/esci-data.
train-searchqa
SearchQA — Training, unified schema
A normalised copy of the dataset behind the mteb task SearchQA, a retrieval training set built from sentence-transformers/embedding-training-data. Same queries, documents
and relevance judgements as the benchmark evaluates — reshaped into one strict schema shared by every dataset
in this collection.
Source
sentence-transformers/embedding-training-data @ 03987af07011 (the revision pinned in mteb)
Domain · languages
Jeopardy QA /… See the full description on the dataset page: https://huggingface.co/datasets/Hyukkyu/train-searchqa.Llama-3.2-1B-Instruct-beam-search-completionsindia-stocks-and-search-trends
India stocks and search trends
2587 symbols, one plain CSV per symbol and resolution.
import pandas as pd
df = pd.read_csv("hf://datasets/yadavsaurabh/india-stocks-and-search-trends/prices/1d/20MICRONS.NS.csv")
from datasets import load_dataset # every symbol of one resolution as one table
ds = load_dataset("yadavsaurabh/india-stocks-and-search-trends", "1d")
Prices (prices/<interval>/<SYMBOL>.csv)
Interval
Bars
History
1d
daily
full history… See the full description on the dataset page: https://huggingface.co/datasets/yadavsaurabh/india-stocks-and-search-trends.weibo-hot-searchsearch-arena-v1-7k
Overview
This dataset contains 7k leaderboard conversation votes collected from Search Arena between March 18, 2025 and April 13, 2025. All entries have been redacted for PII and sensitive user information to ensure privacy.
Each data point includes:
Two model responses (messages_a and messages_b)
The human vote result
A timestamp
Full system metadata, LLM + web search trace, and post-processed metadata for controlled experiments (conv_meta)
To reproduce the leaderboard results… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/search-arena-v1-7k.superscout-sft-search
SuperScout search corpus (SFT)
The supervised fine-tuning corpus behind SuperScout-7B, a 7B searcher that
explores a repository, localizes the fault, writes a failing reproduction, and
emits a structured handoff. The dataset contains 19,911 examples as built and
frozen; six rows carrying malformed tool-call wrappers are dropped at load time,
giving the 19,905 examples actually trained on. A further 9,478 examples (the
third-best trace per issue) were held back as a shelf and… See the full description on the dataset page: https://huggingface.co/datasets/SuperAGI/superscout-sft-search.semantic-history-search
Semantic History (Synthetic)
Semantic History is a synthetic dataset for semantic history search research.It ships as three normalized Parquet tables:
Split
Description
docs
Search history records with url, title, description, frecency, last_visit_date, and tags.
queries
One row per query, tagged by profile and temporal/multi-label flags.
qrels
Relevance pairs linking queries ↔ docs with rank and relevance.
All content is synthetic (no real browsing logs).… See the full description on the dataset page: https://huggingface.co/datasets/frankjc2022/semantic-history-search.Test_Semantic_SearchLlama-3.2-3B-Instruct-beam-search-completionssearch-source-audit
Sources of Truth — AI Search Citations for Mental Health Queries
Which external sources do consumer AI search products actually cite when people ask about mental
health? This dataset is the annotated citation corpus behind "Sources of Truth: A Multi-Platform,
Multilingual Audit of Citations in AI Mental Health Information Queries."
Twenty English mental health questions were put to three free consumer AI search products (ChatGPT, Perplexity, and Google AI Overview) under two… See the full description on the dataset page: https://huggingface.co/datasets/MindBench/search-source-audit.
