Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mikex86 /stackoverflow-posts StackOverflow Posts Markdown Dataset Summary This dataset contains all posts submitted to StackOverflow before the 14th of June 2023 formatted as Markdown text. The dataset contains ~60 Million posts, totaling ~35GB in size and ~65 billion characters of text. The data is sourced from Internet Archive StackExchange Data Dump. Dataset Structure Each record corresponds to one post of a particular type. Original ordering from the data dump is not exactly preserved… See the full description on the dataset page: https://huggingface.co/datasets/mikex86/stackoverflow-posts.tabularquestion-answering10M<n<100M63 likes14k downloads3y agoHugging Face02suriyagunasekar /stackoverflow-with-meta-data Dataset Card for "stackoverflow-with-meta-data" More Information needed text10M<n<100M13 likes5.2k downloads4y agoHugging Face03CoIR-Retrieval /stackoverflow-qaEmploying the MTEB evaluation framework's dataset version, utilize the code below for assessment: import mteb import logging from sentence_transformers import SentenceTransformer from mteb import MTEB logger = logging.getLogger(__name__) model_name = 'intfloat/e5-base-v2' model = SentenceTransformer(model_name) tasks = mteb.get_tasks( tasks=[ "AppsRetrieval", "CodeFeedbackMT", "CodeFeedbackST", "CodeTransOceanContest", "CodeTransOceanDL"… See the full description on the dataset page: https://huggingface.co/datasets/CoIR-Retrieval/stackoverflow-qa.text10K<n<100K0 likes2.4k downloads2y agoHugging Face04mteb /StackOverflowQA StackOverflowQA An MTEB dataset Massive Text Embedding Benchmark The dataset is a collection of natural language queries and their corresponding response which may include some text mixed with code snippets. The task is to retrieve the most relevant response for a given query. Task category t2t Domains Programming, Written Reference https://arxiv.org/abs/2407.02883 How to evaluate on this task You can evaluate an embedding model on this dataset using… See the full description on the dataset page: https://huggingface.co/datasets/mteb/StackOverflowQA.texttext-retrieval10K<n<100K0 likes2.1k downloads1y agoHugging Face05bigcode /stackoverflow-clean Dataset Card for "stackoverflow-clean" More Information needed tabular10M<n<100M8 likes1.6k downloads3y agoHugging Face06pacovaldez /stackoverflow-questions Dataset Card for [Stackoverflow Post Questions] Dataset Description Companies that sell Open-source software tools usually hire an army of Customer representatives to try to answer every question asked about their tool. The first step in this process is the prioritization of the question. The classification scale usually consists of 4 values, P0, P1, P2, and P3, with different meanings across every participant in the industry. On the other hand, every software developer… See the full description on the dataset page: https://huggingface.co/datasets/pacovaldez/stackoverflow-questions.texttext-classification1M<n<10M52 likes1.1k downloads4y agoHugging Face07suriyagunasekar /stackoverflow-python-with-meta-data Dataset Card for "stackoverflow-python-with-meta-data" More Information needed text1M<n<10M13 likes1.1k downloads4y agoHugging Face08MTEB-BR /stackoverflow-clustering StackoverflowPtClustering Cluster native Brazilian-Portuguese technical question titles from the Portuguese Stack Overflow (pt.stackoverflow.com) into 10 technology tags (python, java, php, javascript, android, mysql, c#, html, css, c). Programming domain. Part of MTEB-BR — the native Brazilian-Portuguese MTEB sub-benchmark. Task type: Clustering · Language: Brazilian Portuguese (mined from real-world sources) · Domains: Programming, Web, Written. Dataset structure… See the full description on the dataset page: https://huggingface.co/datasets/MTEB-BR/stackoverflow-clustering.texttext-classification1K<n<10K0 likes1.1k downloads2mo agoHugging Face09mteb /StackOverflowDupQuestions StackOverflowDupQuestions An MTEB dataset Massive Text Embedding Benchmark Stack Overflow Duplicate Questions Task for questions with the tags Java, JavaScript and Python Task category t2t Domains Written, Blog, Programming Reference https://www.microsoft.com/en-us/research/uploads/prod/2019/03/nl4se18LinkSO.pdf How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/StackOverflowDupQuestions.texttext-ranking1M<n<10M0 likes990 downloads1y agoHugging Face10easytpp /stackoverflowtabular1K<n<10K1 likes756 downloads3y agoHugging Face11koutch /stackoverflow_python Dataset Card for "stackoverflow_python" Dataset Summary This dataset comes originally from kaggle. It was originally split into three tables (CSV files) (Questions, Answers, and Tags) now merged into a single table. Each row corresponds to a pair (question-answer) and their associated tags. The dataset contains all questions asked between August 2, 2008 and Ocotober 19, 2016. Supported Tasks and Leaderboards This might be useful for open-domain… See the full description on the dataset page: https://huggingface.co/datasets/koutch/stackoverflow_python.tabularquestion-answering100K<n<1M33 likes643 downloads4y agoHugging Face12mteb /stackoverflowdupquestions-rerankingtext10K<n<100K3 likes632 downloads4y agoHugging Face13mirzaei2114 /stackoverflowVQA-filteredimagevisual-question-answering100K<n<1M3 likes403 downloads3y agoHugging Face14mlfoundations-dev /stackoverflowtext100M<n<1B2 likes331 downloads2y agoHugging Face15terryyz /stackoverflow_ngram_13text10M<n<100M1 likes302 downloads2y agoHugging Face16code-rag-bench /stackoverflow-postsThe StackOverflow posts retrieval source for code-rag-bench. text1M<n<10M2 likes293 downloads2y agoHugging Face17nvidia /Nemotron-RL-math-stack_overflow Dataset Description: The Nemotron-RL-math-stack_overflow dataset contains mathematical problems and solutions sourced from the Stack Overflow forums. The method of extracting problems and solutions from forum posts was similar to the one used to create the OpenMathReasoning dataset, which is described in this paper. Only problems for which an answer was extracted are included in the present dataset. This dataset is released as part of NVIDIA NeMo Gym, a framework for building… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-math-stack_overflow.4 likes236 downloads10d agoHugging Face18farida5gaber /stackoverflow-posts StackOverflow Posts Markdown Dataset Summary This dataset contains all posts submitted to StackOverflow before the 14th of June 2023 formatted as Markdown text. The dataset contains ~60 Million posts, totaling ~35GB in size and ~65 billion characters of text. The data is sourced from Internet Archive StackExchange Data Dump. Dataset Structure Each record corresponds to one post of a particular type. Original ordering from the data dump is not exactly preserved… See the full description on the dataset page: https://huggingface.co/datasets/farida5gaber/stackoverflow-posts.tabularquestion-answering10M<n<100M0 likes219 downloads6mo agoHugging Face19raymondzmc /stackoverflow_Llama-3.1-8B-Instruct_vocab_2000_lasttabular10K<n<100K0 likes195 downloads10mo agoHugging Face20reubenjohn /stackoverflow-open-status-classification-albert-tokenized Dataset Card for "stackoverflow-open-status-classification-albert-tokenized" More Information needed text1M<n<10M0 likes167 downloads4y agoHugging Face21BramVanroy /stackoverflow-chat-dutch Dataset Card for Stack Overflow Chat Dutch Dataset Summary This dataset contains 56,964 conversations between een AI assistant and a (fake) "Human" (generated) in Dutch, specifically in the domain of programming (Stack Overflow). They are translations of Baize's machine-generated answers to the Stack Overflow dataset. ☕ Want to help me out? Translating the data with the OpenAI API, and prompt testing, cost me 💸$133.60💸. If you like this dataset, please consider buying… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/stackoverflow-chat-dutch.textquestion-answering10K<n<100K2 likes138 downloads3y agoHugging Face22pacovaldez /stackoverflow-questions-2016 Dataset Card for [Stackoverflow Post Questions] Dataset Description Companies that sell Open-source software tools usually hire an army of Customer representatives to try to answer every question asked about their tool. The first step in this process is the prioritization of the question. The classification scale usually consists of 4 values, P0, P1, P2, and P3, with different meanings across every participant in the industry. On the other hand, every software developer… See the full description on the dataset page: https://huggingface.co/datasets/pacovaldez/stackoverflow-questions-2016.texttext-classification100K<n<1M1 likes137 downloads4y agoHugging Face23IlyaGusev /ru_stackoverflow Russian StackOverflow dataset Description Summary: Dataset of questions, answers, and comments from ru.stackoverflow.com. Script: create_stackoverflow.py Point of Contact: Ilya Gusev Languages: The dataset is in Russian with some programming code. Usage Prerequisites: pip install datasets zstandard jsonlines pysimdjson Loading: from datasets import load_dataset dataset = load_dataset('IlyaGusev/ru_stackoverflow', split="train") for example in dataset:… See the full description on the dataset page: https://huggingface.co/datasets/IlyaGusev/ru_stackoverflow.text-generation100K<n<1M12 likes136 downloads4y agoHugging Face24reubenjohn /stackoverflow-unified-text-open-status-classification Dataset Card for "stackoverflow-unified-text-open-status-classification" More Information needed tabular1M<n<10M0 likes135 downloads4y agoHugging Face25Harryxun /stackoverflow-parquettext100M<n<1B0 likes129 downloads3d agoHugging Face26KonradSzafer /stackoverflow_linux Dataset Card for "stackoverflow_linux" Dataset information: Source: Stack Overflow Category: Linux Number of samples: 300 Train/Test split: 270/30 Quality: Data come from the top 1k most upvoted questions Additional Information License All Stack Overflow user contributions are licensed under CC-BY-SA 3.0 with attribution required. More Information needed textquestion-answeringn<1K9 likes126 downloads4y agoHugging Face27xPXXX /stackoverflow_DL-related_questionstabular10K<n<100K0 likes122 downloads3y agoHugging Face28c17hawke /stackoverflow-datasettabular10K<n<100K6 likes119 downloads4y agoHugging Face29mcipriano /stackoverflow-kubernetes-questionsThe purpose of this dataset is to provide the opportunity to perform any training, fine-tuning, etc. for any Language Model. In the 'data' folder, you will find the dataset in Parquet format, which is one of the formats used for these processes. In case it may be useful for other purposes, I have also included the dataset in CSV format. All data in this dataset were retrieved from the Stack Exchange network using the Stack Exchange Data explorer tool… See the full description on the dataset page: https://huggingface.co/datasets/mcipriano/stackoverflow-kubernetes-questions.textquestion-answering10K<n<100K31 likes118 downloads3y agoHugging Face30CoIR-Retrieval /stackoverflow-qa-queries-corpusEmploying the CoIR evaluation framework's dataset version, utilize the code below for assessment: import coir from coir.data_loader import get_tasks from coir.evaluation import COIR from coir.models import YourCustomDEModel model_name = "intfloat/e5-base-v2" # Load the model model = YourCustomDEModel(model_name=model_name) # Get tasks #all task ["codetrans-dl","stackoverflow-qa","apps","codefeedback-mt","codefeedback-st","codetrans-contest","synthetic- # text2sql","cosqa","codesearchnet"… See the full description on the dataset page: https://huggingface.co/datasets/CoIR-Retrieval/stackoverflow-qa-queries-corpus.text10K<n<100K0 likes116 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.