Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Becky7777777 /polymarket-search-indextabular10M<n<100M0 likes4.6k downloads10h agoHugging Face02yadavsaurabh /us-stocks-and-search-trends US stocks and search trends 11636 symbols, one plain CSV per symbol and resolution. import pandas as pd df = pd.read_csv("hf://datasets/yadavsaurabh/us-stocks-and-search-trends/prices/1d/A.csv") from datasets import load_dataset # every symbol of one resolution as one table ds = load_dataset("yadavsaurabh/us-stocks-and-search-trends", "1d") Prices (prices/<interval>/<SYMBOL>.csv) Interval Bars History 1d daily full history (back to 1962 for… See the full description on the dataset page: https://huggingface.co/datasets/yadavsaurabh/us-stocks-and-search-trends.tabular10M<n<100M1 likes1.8k downloads9d agoHugging Face03davanstrien /search-v3-embeddings Hub Card Search Embeddings (v3) One-sentence summaries and 1024-d embeddings for 1,173,030 dataset and model cards on the Hugging Face Hub — 536,870 datasets and 636,160 models. It is the search index behind the revived librarian-bots/huggingface-semantic-search backend: you search over a short model-written summary of each card rather than the raw card, and retrieve against the embedding of that summary. The cards come from librarian-bots/dataset_cards_with_metadata and… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/search-v3-embeddings.tabularfeature-extraction1M<n<10M4 likes1.2k downloads2mo agoHugging Face04infinite-dataset-hub /FINDER_API_KEY_AI_SEARCH_2023 FINDER_API_KEY_AI_SEARCH_2023 tags: data collection, machine learning, API performance Note: This is an AI-generated dataset so its content may be inaccurate or false Dataset Description: The 'FINDER_API_KEY_AI_SEARCH_2023' dataset is designed to collect and analyze data from various AI search engines and their associated API performance metrics. The dataset focuses on the effectiveness of API key-based access in enhancing the search capabilities of AI systems and includes a… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/FINDER_API_KEY_AI_SEARCH_2023.tabularn<1K0 likes963 downloads2y agoHugging Face05reasoning-cues /rollouts-olmo7b-cue-search rollouts-olmo7b-cue-search Model: allenai/Olmo-3-1025-7B (snapshot a81bae42). Tokenizer: allenai/Olmo-3-1025-7B (snapshot a81bae42). Protocol: RL-Zero prompt, MATH-500 x 4 rollouts, budget 31,744, T 0.6, top-p 0.95, seed 20260819 (depth-2 exhaustive and n-gram chain: seed 20260821); the top-20 beam nominee screen, ten random-opener arms, every depth-2 opener (84 shards, arm names unique across shards) and the n-gram chain arms. Rollouts generated on the CSAIL cluster for the… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-cues/rollouts-olmo7b-cue-search.tabular100K<n<1M0 likes939 downloads27d agoHugging Face06microsoft /hnm-search-data HnM Search Dataset Created from Recommendations Dataset This synthetic data-set is created using the recommendations dataset: https://huggingface.co/datasets/einrafh/hnm-fashion-recommendations-data (Use of this dataset is subject to the terms and conditions set forth on the original distribution page. This dataset is intended for non-commercial and research use.) https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/data (DATA ACCESS AND USE:… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/hnm-search-data.imagetext-ranking10M<n<100M2 likes593 downloads8mo agoHugging Face07peakji /peak-search-300ktabular100K<n<1M0 likes525 downloads2y agoHugging Face08philgabr /rest-graph-searchtabular10M<n<100M0 likes492 downloads19h agoHugging Face09dtunkelang /job-search-replay-trial Temporary job-search replay trial Public job-document sample with controlled test updates. Not the production index. tabularn<1K0 likes380 downloads3d agoHugging Face10NomaDamas /split_search_qa preprocessed_SearchQA The SearchQA question-answer pairs originate from J! Archive2, which comprehensively archives all question-answer pairs from the renowned television show Jeopardy! The passages, sourced from Google search web page snippets. We offer passage metadata, encompassing details like 'air_date,' 'category,' 'value,' 'round,' and 'show_number,' enabling you to enhance retrieval performance at your discretion. Should you require further details about SearchQA, please… See the full description on the dataset page: https://huggingface.co/datasets/NomaDamas/split_search_qa.tabular10M<n<100M0 likes374 downloads3y agoHugging Face11CyberMax-tools /swellmeter-google-trending-searches Swellmeter: what is trending on Google right now, in 30 countries Today's trending searches plus 30 days of history, by API: $5 for 1,500 calls. Buy now → · API key on screen the moment checkout ends · no subscription · 14-day refund if it doesn't work as described · help: cybermax.tools@gmail.com Need it fresh, filtered or via API? This free file is a snapshot (Google trending searches at the last refresh), last updated 2026-10-09. Swellmeter Trending API (50 free calls a… See the full description on the dataset page: https://huggingface.co/datasets/CyberMax-tools/swellmeter-google-trending-searches.tabulartime-series-forecasting1K<n<10K0 likes374 downloads6h agoHugging Face12spectralbranding /r15-ai-search-metamerism R15: AI Search Metamerism — Cross-Cultural Brand Perception Dataset Citation: Zharnikov, D. (2026v) | DOI: 10.5281/zenodo.19422427 | Version: v3.2.0 Dataset Summary This dataset contains the full session logs, aggregated results, and analysis outputs from the R15 large-scale experiment testing whether Large Language Models systematically collapse multi-dimensional brand perception into Economic and Experiential dimensions ("spectral metamerism"). It comprises 21… See the full description on the dataset page: https://huggingface.co/datasets/spectralbranding/r15-ai-search-metamerism.tabulartext-generationn<1K0 likes355 downloads3mo agoHugging Face13google /code_x_glue_tc_nl_code_search_adv Dataset Card for "code_x_glue_tc_nl_code_search_adv" Dataset Summary CodeXGLUE NL-code-search-Adv dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Text-Code/NL-code-search-Adv The dataset we use comes from CodeSearchNet and we filter the dataset as the following: Remove examples that codes cannot be parsed into an abstract syntax tree. Remove examples that #tokens of documents is < 3 or >256 Remove examples that documents contain special tokens… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_tc_nl_code_search_adv.tabulartext-retrieval100K<n<1M11 likes346 downloads3y agoHugging Face14peakji /peak-search-content-70ktabular10K<n<100K0 likes293 downloads2y agoHugging Face15P-Arpan /Mood_Based_Music_Search Dataset Card for Acoustic_Features_From_Spotify This dataset is created by merging multiple Kaggle datasets of Spotify audio features and track metadata into a unified, clean, and deduplicated collection of 2,577,667 unique tracks. Each record is indexed by Spotify track_id (with planned support for ISRC identifiers in future iterations). Dataset Details Dataset Description Acoustic_Features_From_Spotify consolidates acoustic properties and… See the full description on the dataset page: https://huggingface.co/datasets/P-Arpan/Mood_Based_Music_Search.tabular1M<n<10M1 likes270 downloads3d agoHugging Face16uw-math-ai /theorem-search-dataset Theorem Search Dataset The largest open corpus of informal mathematical theorems: 1,341,083 theorem statements with natural-language slogans from 209,777 papers, designed for semantic theorem retrieval. Paper: Semantic Search over 9 Million Mathematical Theorems Demo: huggingface.co/spaces/uw-math-ai/theorem-search Benchmark results On 110 test queries written by research mathematicians, our best pipeline (Qwen3-Embedding-8B on DeepSeek-V3.1 slogans) outperforms… See the full description on the dataset page: https://huggingface.co/datasets/uw-math-ai/theorem-search-dataset.tabularquestion-answering1M<n<10M26 likes262 downloads9d agoHugging Face17CyberMax-tools /superintelligence-search-questions Keyfern: what people search about superintelligence (Sept 2026) Google's trending searches for 30 countries with 30 days of history, by API: $5 for 1,500 calls. Buy now → · API key on screen the moment checkout ends · no subscription · 14-day refund if it doesn't work as described · help: cybermax.tools@gmail.com Need it fresh, filtered or via API? This free file is a snapshot (search questions at the last refresh), last updated 2026-09-24. Keyfern on Apify ($0.001 per… See the full description on the dataset page: https://huggingface.co/datasets/CyberMax-tools/superintelligence-search-questions.tabulartext-classification1K<n<10K0 likes259 downloads6h agoHugging Face18tbuckley /cabot-search CaBot Clinical Literature Embedding Index Exact embedding-search index over 3,474,244 works from 204 high-impact clinical journals (2023 JIF >= 10), built from an OpenAlex snapshot (~June 2025) and embedded with OpenAI text-embedding-3-small at 1536 dimensions (float32). This is used for CaBot. The build and search code is in the CaBot-Search/ folder of the source repository. Columns column type notes id string OpenAlex work id doi string title… See the full description on the dataset page: https://huggingface.co/datasets/tbuckley/cabot-search.tabularfeature-extraction1M<n<10M0 likes224 downloads4mo agoHugging Face19cx-cmu /deepresearchgym-agentic-search-logs DeepResearchGym Agentic Search Logs This repository hosts the dataset accompanying the paper “Agentic Search in the Wild” (arXiv: https://arxiv.org/abs/2601.17617). The dataset contains 14M+ search queries collected via DeepResearchGym (DRGym), an open-source search API designed for DeepResearch-style agentic search. For more background on DRGym, see: https://arxiv.org/abs/2505.19253. All records have been anonymized and shuffled to prevent re-identification, and we additionally… See the full description on the dataset page: https://huggingface.co/datasets/cx-cmu/deepresearchgym-agentic-search-logs.tabulartext-retrieval10M<n<100M16 likes214 downloads8mo agoHugging Face20asahi417 /amazon-product-searchPlease refer the original repository https://github.com/amazon-science/esci-data. tabular1M<n<10M4 likes206 downloads2y agoHugging Face21Hyukkyu /train-searchqa SearchQA — Training, unified schema A normalised copy of the dataset behind the mteb task SearchQA, a retrieval training set built from sentence-transformers/embedding-training-data. Same queries, documents and relevance judgements as the benchmark evaluates — reshaped into one strict schema shared by every dataset in this collection. Source sentence-transformers/embedding-training-data @ 03987af07011 (the revision pinned in mteb) Domain · languages Jeopardy QA /… See the full description on the dataset page: https://huggingface.co/datasets/Hyukkyu/train-searchqa.tabulartext-retrieval10M<n<100M0 likes198 downloads8d agoHugging Face22HuggingFaceH4 /Llama-3.2-1B-Instruct-beam-search-completionstabular10K<n<100K1 likes196 downloads2y agoHugging Face23yadavsaurabh /india-stocks-and-search-trends India stocks and search trends 2587 symbols, one plain CSV per symbol and resolution. import pandas as pd df = pd.read_csv("hf://datasets/yadavsaurabh/india-stocks-and-search-trends/prices/1d/20MICRONS.NS.csv") from datasets import load_dataset # every symbol of one resolution as one table ds = load_dataset("yadavsaurabh/india-stocks-and-search-trends", "1d") Prices (prices/<interval>/<SYMBOL>.csv) Interval Bars History 1d daily full history… See the full description on the dataset page: https://huggingface.co/datasets/yadavsaurabh/india-stocks-and-search-trends.tabular10M<n<100M1 likes184 downloads9d agoHugging Face24LZYFirecn /weibo-hot-searchtabular10K<n<100K3 likes165 downloads1y agoHugging Face25lmarena-ai /search-arena-v1-7k Overview This dataset contains 7k leaderboard conversation votes collected from Search Arena between March 18, 2025 and April 13, 2025. All entries have been redacted for PII and sensitive user information to ensure privacy. Each data point includes: Two model responses (messages_a and messages_b) The human vote result A timestamp Full system metadata, LLM + web search trace, and post-processed metadata for controlled experiments (conv_meta) To reproduce the leaderboard results… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/search-arena-v1-7k.tabular1K<n<10K28 likes163 downloads1y agoHugging Face26SuperAGI /superscout-sft-search SuperScout search corpus (SFT) The supervised fine-tuning corpus behind SuperScout-7B, a 7B searcher that explores a repository, localizes the fault, writes a failing reproduction, and emits a structured handoff. The dataset contains 19,911 examples as built and frozen; six rows carrying malformed tool-call wrappers are dropped at load time, giving the 19,905 examples actually trained on. A further 9,478 examples (the third-best trace per issue) were held back as a shelf and… See the full description on the dataset page: https://huggingface.co/datasets/SuperAGI/superscout-sft-search.tabulartext-generation10K<n<100K0 likes155 downloads11d agoHugging Face27frankjc2022 /semantic-history-search Semantic History (Synthetic) Semantic History is a synthetic dataset for semantic history search research.It ships as three normalized Parquet tables: Split Description docs Search history records with url, title, description, frecency, last_visit_date, and tags. queries One row per query, tagged by profile and temporal/multi-label flags. qrels Relevance pairs linking queries ↔ docs with rank and relevance. All content is synthetic (no real browsing logs).… See the full description on the dataset page: https://huggingface.co/datasets/frankjc2022/semantic-history-search.tabular100K<n<1M5 likes143 downloads11mo agoHugging Face28Lexington120 /Test_Semantic_Searchtabularn<1K6 likes139 downloads3y agoHugging Face29HuggingFaceH4 /Llama-3.2-3B-Instruct-beam-search-completionstabular10K<n<100K1 likes123 downloads2y agoHugging Face30MindBench /search-source-audit Sources of Truth — AI Search Citations for Mental Health Queries Which external sources do consumer AI search products actually cite when people ask about mental health? This dataset is the annotated citation corpus behind "Sources of Truth: A Multi-Platform, Multilingual Audit of Citations in AI Mental Health Information Queries." Twenty English mental health questions were put to three free consumer AI search products (ChatGPT, Perplexity, and Google AI Overview) under two… See the full description on the dataset page: https://huggingface.co/datasets/MindBench/search-source-audit.tabular10K<n<100K0 likes117 downloads16d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.