datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cqadupstack-programmers-vn
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["CQADupstackProgrammers-VN"])
evaluator = mteb.MTEB(task)
model = mteb.get_model(YOUR_MODEL)
evaluator.run(model)
To learn more about how to run models on mteb task check out the GitHub repitory.
Citation
If you use this dataset, please cite the dataset as well as mteb, as this dataset likely includes additional processing… See the full description on the dataset page: https://huggingface.co/datasets/GreenNode/cqadupstack-programmers-vn.CQADupstack-Programmers-PL
CQADupstack-Programmers-PL
An MTEB dataset
Massive Text Embedding Benchmark
CQADupStack: A Stack Exchange Question Duplicate Pairs Dataset
Task category
t2t
Domains
Programming, Written, Non-fiction
Reference
https://huggingface.co/datasets/clarin-knext/cqadupstack-programmers-pl
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["CQADupstack-Programmers-PL"])… See the full description on the dataset page: https://huggingface.co/datasets/mteb/CQADupstack-Programmers-PL.cqadupstack-programmers-vn-rawCQADupstackProgrammersRetrieval-Fa
CQADupstackProgrammersRetrieval-Fa
An MTEB dataset
Massive Text Embedding Benchmark
CQADupstackProgrammersRetrieval-Fa
Task category
t2t
Domains
Web
Reference
https://huggingface.co/datasets/MCINext/cqadupstack-programmers-fa
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["CQADupstackProgrammersRetrieval-Fa"])
evaluator = mteb.MTEB(task)
model =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/CQADupstackProgrammersRetrieval-Fa.beir_cqadupstack_programmers
CQADupStack / Programmers (BEIR) — programming Q&A retrieval
Dataset description
CQADupStack is a benchmark for community question answering (cQA) built from Stack Exchange data. It was introduced by Hoogeveen, Verspoor, and Baldwin at ADCS 2015 to support research on duplicate questions: finding earlier posts that match or subsume a newly asked question, so users can reuse existing answers instead of opening redundant threads.
The full CQADupStack release aggregates… See the full description on the dataset page: https://huggingface.co/datasets/orgrctera/beir_cqadupstack_programmers.cqadupstack-programmers-qrels
Dataset Card for "cqadupstack-programmers-qrels"
More Information needed
beir-cqadupstack-programmers
CQADupstackProgrammersRetrieval — BEIR, unified schema
A normalised copy of the dataset behind the mteb task CQADupstackProgrammersRetrieval, one of the tasks of the BEIR benchmark as mteb defines it (a member of the aggregate task CQADupstackRetrieval). Same queries, documents
and relevance judgements as the benchmark evaluates — reshaped into one strict schema shared by every dataset
in this collection.
Source
mteb/cqadupstack-programmers @ 6184bc1440d2 (the… See the full description on the dataset page: https://huggingface.co/datasets/Hyukkyu/beir-cqadupstack-programmers.beir_cqadupstack_programmers_test
beir_cqadupstack_programmers_test
BEIR CQADupStack/programmers test split
Field
Value
Benchmark
beir
Sub-benchmark
cqadupstack_programmers
Type
retrieval
Items
876
Exported from Langfuse.
cqadupstack-programmers
Dataset Card for "cqadupstack-programmers"
More Information needed
CQADupstackProgrammers-NL
