datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cqadupstack-programmers
CQADupstackProgrammersRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
CQADupStack: A Benchmark Data Set for Community Question-Answering Research
Task category
t2t
Domains
Programming, Written, Non-fiction
Referencehttp://nlp.cis.unimelb.edu.au/resources/cqadupstack/
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/cqadupstack-programmers.sfia-9-scraped
SFIA-9-Scraped Dataset
This repository contains the SFIA-9-Scraped dataset, a JSON collection of the Skills Framework for the Information Age (SFIA) version 9 categories and levels, scraped for non-commercial research use.
🚀 Dataset Overview
Name: SFIA-9-Scraped
Hugging Face: Programmer-RD-AI/sfia-9-scraped
DOI: 10.57967/hf/5746
Author: Ranuga Disansa Gamage
Revision: 89feeb8
Publisher: Hugging Face
Year: 2025
Use this dataset to build RAG systems, taxonomy-driven… See the full description on the dataset page: https://huggingface.co/datasets/Programmer-RD-AI/sfia-9-scraped.cqadupstack-programmers-top-20-gen-queries
NFCorpus: 20 generated queries (BEIR Benchmark)
This HF dataset contains the top-20 synthetic queries generated for each passage in the above BEIR benchmark dataset.
DocT5query model used: BeIR/query-gen-msmarco-t5-base-v1
id (str): unique document id in NFCorpus in the BEIR benchmark (corpus.jsonl).
Questions generated: 20
Code used for generation: evaluate_anserini_docT5query_parallel.py
Below contains the old dataset card for the BEIR benchmark.
Dataset Card for BEIR… See the full description on the dataset page: https://huggingface.co/datasets/income/cqadupstack-programmers-top-20-gen-queries.cqadupstack-programmers-fa
Dataset Summary
CQADupstack-programmers-Fa is a Persian (Farsi) dataset developed for the Retrieval task, with a focus on duplicate question detection in community question-answering (CQA) platforms. This dataset is a translated version of the "Programmers" (Software Engineering) StackExchange subforum from the English CQADupstack collection and is part of the FaMTEB benchmark under the BEIR-Fa suite.
Language(s): Persian (Farsi)
Task(s): Retrieval (Duplicate Question Retrieval)… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/cqadupstack-programmers-fa.reddit-ProgrammerHumor-testaurora_programmer_data
My Awesome Dataset
A comprehensive description of my awesome dataset.
Dataset Description
This dataset contains images of cats and dogs. The images were collected from [mention data source(s), e.g., a specific website, scraped from the internet]. It is intended for use in image classification tasks. The dataset consists of [number] images, with approximately [percentage]% allocated to the training set and [percentage]% to the test set. [Add more details about the… See the full description on the dataset page: https://huggingface.co/datasets/naimulislam/aurora_programmer_data.sfia-9-chunks
sfia-9-chunks Dataset
Overview
The sfia-9-chunks dataset is a derived dataset from sfia-9-scraped. It uses sentence embeddings and hierarchical clustering to split each SFIA-9 document into coherent semantic chunks. This chunking facilitates more efficient downstream tasks like semantic search, question answering, and topic modeling.
Chunking Methodology
We employ the following procedure to generate chunks:
from sentence_transformers import… See the full description on the dataset page: https://huggingface.co/datasets/Programmer-RD-AI/sfia-9-chunks.MNLP_M3_mcqa_datasetMNLP_M2_mcqa_dataset
