Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mteb /stackexchange-clustering StackExchangeClustering.v2 An MTEB dataset Massive Text Embedding Benchmark Clustering of titles from 121 stackexchanges. Clustering of 25 sets, each with 10-50 classes, and each class with 100 - 1000 sentences. Task category t2c Domains Web, Written Reference https://arxiv.org/abs/2104.07081 How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/stackexchange-clustering.texttext-classificationn<1K1 likes7.3k downloads8mo agoHugging Face02common-pile /stackexchange_filtered Stack Exchange Description StackExchange is a collection of Q&A communities spanning a wide variety of topics.While StackExchange formerly provided structured dumps of all of their content, since July of 2024, StackExchange has stopped publishing XML dumps to the Internet Archive.Instead, each site can provide a logged-in user with a custom URL to download the dump for that site.This means that dumps for defunct sites like windowsphone.stackexchange.com are… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/stackexchange_filtered.texttext-generation10M<n<100M10 likes4.2k downloads1y agoHugging Face03common-pile /stackexchange Stack Exchange Description StackExchange is a collection of Q&A communities spanning a wide variety of topics. While StackExchange formerly provided structured dumps of all of their content, since July of 2024, StackExchange has stopped publishing XML dumps to the Internet Archive. Instead, each site can provide a logged in user with a custom url to download the dump for that site. This means that dumps for defunct sites like windowsphone.stackexchange.com are… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/stackexchange.texttext-generation10M<n<100M8 likes4k downloads1y agoHugging Face04mteb /stackexchange-clustering-p2p StackExchangeClusteringP2P.v2 An MTEB dataset Massive Text Embedding Benchmark Clustering of title+body from stackexchange. Clustering of 5 sets of 10k paragraphs and 5 sets of 5k paragraphs. Task category t2c Domains Web, Written Reference https://arxiv.org/abs/2104.07081 How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["StackExchangeClusteringP2P.v2"])… See the full description on the dataset page: https://huggingface.co/datasets/mteb/stackexchange-clustering-p2p.texttext-classificationn<1K1 likes3k downloads1y agoHugging Face05marianna13 /physics-stackexchangetext10K<n<100K2 likes1.2k downloads3y agoHugging Face06flax-sentence-embeddings /stackexchange_title_body_jsonljsonl.gz format from https://huggingface.co/datasets/flax-sentence-embeddings/stackexchange_xml Each line contains a dict in the format: {"text": ["title", "body"], "tags": ["tag1", "tag2"]} The following parameters have been used for filtering: min_title_len = 20 min_body_len = 20 max_body_len = 4096 min_score = 0 If a stackexchange contained less than 10k questions (after filtering), it is written to the small_stackexchanges.jsonl.gz file. This is a dump of the files from… See the full description on the dataset page: https://huggingface.co/datasets/flax-sentence-embeddings/stackexchange_title_body_jsonl.text1M<n<10M2 likes1k downloads5y agoHugging Face07marin-community /stackexchange-markdown Marin Markdownified StackExchange Markdownified Stack Exchange transforms the Stack Exchange's question-answer pairs into Markdown format consisting of 20.4B tokens. This dataset preserves the content contained in technical discussions while organizing it into a thread format for language model training. Value Tokens 20 413 785 853 Primary source https://archive.org/details/stackexchange File format JSONL License CC (mirrors upstream SE licenses)… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/stackexchange-markdown.texttext-generation10M<n<100M6 likes752 downloads1y agoHugging Face08raj2708 /stackexchange-all StackExchange All Communities — Preprocessed All 362 StackExchange communities (including Stack Overflow) processed into QA pairs and standalone questions, ready for LLM pretraining and instruction tuning. Stats Field Value Communities 362 Dump date March 2026 License CC-BY-SA 4.0 Record Types instruction — Question + Answer pair (qa_pair) text — Unanswered question (standalone) Format { "id": "uuid-v4"… See the full description on the dataset page: https://huggingface.co/datasets/raj2708/stackexchange-all.text10M<n<100M1 likes439 downloads4mo agoHugging Face09jonathanli /law-stack-exchange Dataset Card for Law Stack Exchange Dataset Dataset Summary Dataset from the Law Stack Exchange, as used in "Parameter-Efficient Legal Domain Adaptation". Citation Information @inproceedings{li-etal-2022-parameter, title = "Parameter-Efficient Legal Domain Adaptation", author = "Li, Jonathan and Bhambhoria, Rohan and Zhu, Xiaodan", booktitle = "Proceedings of the Natural Legal Language Processing Workshop 2022", month = dec… See the full description on the dataset page: https://huggingface.co/datasets/jonathanli/law-stack-exchange.tabulartext-classification1K<n<10K17 likes349 downloads4y agoHugging Face10ymoslem /Law-StackExchange Law-StackExchange Dataset Details All StackExchange legal questions and their answers from the Law site, up to 14 August 2023. The repository includes a notebook for the process using the official StackExchange API. Citation @misc{Moslem2023-LawStackExchangeDataset, author = {Moslem, Yasmin}, title = {Law-StackExchange Dataset}, year = 2023, url = {https://huggingface.co/datasets/ymoslem/Law-StackExchange}, doi =… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/Law-StackExchange.tabularquestion-answering10K<n<100K32 likes148 downloads1y agoHugging Face11ymoslem /MedicalSciences-StackExchangeAll StackExchange questions and their answers from the Medical Sciences site, up to 14 August 2023. The repository includes a notebook for the process using the official StackExchange API. tabularquestion-answering1K<n<10K11 likes111 downloads3y agoHugging Face12PsiPi /robotics_stackexchange_com_QLORAtext1K<n<10K0 likes79 downloads3y agoHugging Face13yyu /stackexchange-attrpromptThis is the data used in the paper Large Language Model as Attributed Training Data Generator: A Tale of Diversity and Bias. Checkout the paper: https://arxiv.org/abs/2306.15895 for details. label.txt: the label name for each class train.jsonl: The original training set. valid.jsonl: The original validation set. test.jsonl: The original test set. simprompt.jsonl: The training data generated by the simple prompt. attrprompt.jsonl: The training data generated by the attributed prompt. texttext-classification10K<n<100K0 likes78 downloads3y agoHugging Face14mteb-arena /arena_emb_stackexchangeEmbeddings used for the MTEB Arena. You can download this repo via git clone https://hf.co/datasets/mteb/arena_emb_stackexchange cd arena_emb_stackexchange git lfs pull or just download individual files, e.g. wget https://hf.co/datasets/mteb/arena_emb_stackexchange/resolve/main/emb_stackexchange_GritLM__GritLM-7B.json.aa As there is an upload limit of 50GB per file, we have split files using e.g. split --number=l/6 emb_stackexchange_GritLM__GritLM-7B.json… See the full description on the dataset page: https://huggingface.co/datasets/mteb-arena/arena_emb_stackexchange.text1M<n<10M0 likes70 downloads2y agoHugging Face15TigerResearch /tigerbot-stackexchange-qa-en-0.5mTigerbot 基于stackexchange问答站点dump数据生成sft数据集 原始来源:https://archive.org/details/stackexchange Usage import datasets ds_sft = datasets.load_dataset('TigerResearch/tigerbot-stackexchange-qa-en-0.5m') text100K<n<1M5 likes61 downloads3y agoHugging Face16bshada /reverseengineering.stackexchange.comtext1K<n<10K3 likes60 downloads2y agoHugging Face17bshada /3dprinting.stackexchange.comtext1K<n<10K4 likes57 downloads2y agoHugging Face18mteb /arena-stackexchange Dataset used for Stackexchange in MTEB/arena Overview The mteb/arena-stackexchange dataset is a curated collection of Stack Exchange questions and answers, designed for use in the MTEB (Massive Text Embedding Benchmark) Arena. This dataset allows various embedding models to compete and be ranked based on their performance on Stack Exchange content. What is Stack Exchange? Stack Exchange is a network of question-and-answer (Q&A) websites on topics in diverse… See the full description on the dataset page: https://huggingface.co/datasets/mteb/arena-stackexchange.text1M<n<10M3 likes57 downloads2y agoHugging Face19prhegde /preference-data-math-stack-exchangeThe preference dataset is derived from the stack exchange dataset which contains questions and answers from the Stack Overflow Data Dump. This contains questions and answers for various topics. For this work, we used only question and answers from math.stackexchange.com sub-folder. The questions are grouped with answers that are assigned a score corresponding to the Anthropic paper: score = log2 (1 + upvotes) rounded to the nearest integer, plus 1 if the answer was accepted by the questioner… See the full description on the dataset page: https://huggingface.co/datasets/prhegde/preference-data-math-stack-exchange.text10K<n<100K6 likes49 downloads3y agoHugging Face20ppak10 /3dprinting.stackexchange.comtext1K<n<10K1 likes48 downloads7mo agoHugging Face21suolyer /pile_stackexchangetext10K<n<100K2 likes40 downloads4y agoHugging Face22datajuicer /redpajama-pile-stackexchange-refined-by-data-juicer RedPajama & The Pile -- StackExchange (refined by Data-Juicer) A refined version of StackExchange dataset in RedPajama & The Pile by Data-Juicer. Removing some "bad" samples from the original merged dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 71GB). Dataset Information Number of samples: 26,309,203 (Keep ~57.89% from the… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-pile-stackexchange-refined-by-data-juicer.texttext-generationn<1K0 likes36 downloads3y agoHugging Face23memray /stackexchangetext100K<n<1M1 likes31 downloads4y agoHugging Face24bshada /electronics.stackexchange.comtext10K<n<100K3 likes29 downloads2y agoHugging Face25bshada /iot.stackexchange.comtextn<1K2 likes23 downloads2y agoHugging Face26NetherlandsForensicInstitute /stackexchange-duplicate-questions-translated-nlThis is a Dutch version of the Stackexchange duplicate questions dataset. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation. textsentence-similarity100K<n<1M0 likes21 downloads2y agoHugging Face27bshada /robotics.stackexchange.comtext10K<n<100K2 likes21 downloads2y agoHugging Face28orionweller /stackexchange-200-wordstext1M<n<10M2 likes19 downloads2y agoHugging Face29bshada /iot.meta.stackexchange.comtextn<1K1 likes18 downloads2y agoHugging Face30orionweller /stackexchange_dolmatext10M<n<100M2 likes18 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.