stackoverflow
mlfoundations-dev_-_llama3-1_8b_mlfoundations-dev-stackoverflow_25000tasks__5p-ggufmlfoundations-dev_-_stackoverflow_5000tasks_0p-ggufmlfoundations-dev_-_stackoverflow_10000tasks_0p-ggufmlfoundations-dev_-_llama3-1_8b_mlfoundations-dev-stackoverflow_25000tasks_1p-ggufmlfoundations-dev_-_llama3-1_8b_mlfoundations-dev-stackoverflow_25000tasks__75p-ggufmlfoundations-dev_-_stackoverflow_5000tasks_.25p-ggufmlfoundations-dev_-_stackoverflow_5000tasks_1p-ggufmlfoundations-dev_-_stackoverflow_10000tasks_.25p-gguf
stackoverflow-posts
StackOverflow Posts Markdown
Dataset Summary
This dataset contains all posts submitted to StackOverflow before the 14th of June 2023 formatted as Markdown text.
The dataset contains ~60 Million posts, totaling ~35GB in size and ~65 billion characters of text.
The data is sourced from Internet Archive StackExchange Data Dump.
Dataset Structure
Each record corresponds to one post of a particular type.
Original ordering from the data dump is not exactly preserved… See the full description on the dataset page: https://huggingface.co/datasets/mikex86/stackoverflow-posts.stackoverflow-with-meta-data
Dataset Card for "stackoverflow-with-meta-data"
More Information needed
stackoverflow-qaEmploying the MTEB evaluation framework's dataset version, utilize the code below for assessment:
import mteb
import logging
from sentence_transformers import SentenceTransformer
from mteb import MTEB
logger = logging.getLogger(__name__)
model_name = 'intfloat/e5-base-v2'
model = SentenceTransformer(model_name)
tasks = mteb.get_tasks(
tasks=[
"AppsRetrieval",
"CodeFeedbackMT",
"CodeFeedbackST",
"CodeTransOceanContest",
"CodeTransOceanDL"… See the full description on the dataset page: https://huggingface.co/datasets/CoIR-Retrieval/stackoverflow-qa.StackOverflowQA
StackOverflowQA
An MTEB dataset
Massive Text Embedding Benchmark
The dataset is a collection of natural language queries and their corresponding response which may include some text mixed with code snippets. The task is to retrieve the most relevant response for a given query.
Task category
t2t
Domains
Programming, Written
Reference
https://arxiv.org/abs/2407.02883
How to evaluate on this task
You can evaluate an embedding model on this dataset using… See the full description on the dataset page: https://huggingface.co/datasets/mteb/StackOverflowQA.stackoverflow-clean
Dataset Card for "stackoverflow-clean"
More Information needed
stackoverflow-python-with-meta-data
Dataset Card for "stackoverflow-python-with-meta-data"
More Information needed
