Team Ai
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mteb /cqadupstack-programmers CQADupstackProgrammersRetrieval An MTEB dataset Massive Text Embedding Benchmark CQADupStack: A Benchmark Data Set for Community Question-Answering Research Task category t2t Domains Programming, Written, Non-fiction Referencehttp://nlp.cis.unimelb.edu.au/resources/cqadupstack/ How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/cqadupstack-programmers.texttext-retrieval10K<n<100K0 likes1.2k downloads1y agoHugging Face02GreenNode /cqadupstack-programmers-vn How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["CQADupstackProgrammers-VN"]) evaluator = mteb.MTEB(task) model = mteb.get_model(YOUR_MODEL) evaluator.run(model) To learn more about how to run models on mteb task check out the GitHub repitory. Citation If you use this dataset, please cite the dataset as well as mteb, as this dataset likely includes additional processing… See the full description on the dataset page: https://huggingface.co/datasets/GreenNode/cqadupstack-programmers-vn.texttext-retrieval10K<n<100K0 likes62 downloads1y agoHugging Face03mteb /CQADupstack-Programmers-PL CQADupstack-Programmers-PL An MTEB dataset Massive Text Embedding Benchmark CQADupStack: A Stack Exchange Question Duplicate Pairs Dataset Task category t2t Domains Programming, Written, Non-fiction Reference https://huggingface.co/datasets/clarin-knext/cqadupstack-programmers-pl How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["CQADupstack-Programmers-PL"])… See the full description on the dataset page: https://huggingface.co/datasets/mteb/CQADupstack-Programmers-PL.texttext-retrieval10K<n<100K0 likes57 downloads1y agoHugging Face04clarin-knext /cqadupstack-programmers-plPart of BEIR-PL: Zero Shot Information Retrieval Benchmark for the Polish Language. Link to arxiv: https://arxiv.org/pdf/2305.19840.pdf Contact: konrad.wojtasik@pwr.edu.pl 0 likes43 downloads2y agoHugging Face05Hyukkyu /beir-cqadupstack-programmers CQADupstackProgrammersRetrieval — BEIR, unified schema A normalised copy of the dataset behind the mteb task CQADupstackProgrammersRetrieval, one of the tasks of the BEIR benchmark as mteb defines it (a member of the aggregate task CQADupstackRetrieval). Same queries, documents and relevance judgements as the benchmark evaluates — reshaped into one strict schema shared by every dataset in this collection. Source mteb/cqadupstack-programmers @ 6184bc1440d2 (the… See the full description on the dataset page: https://huggingface.co/datasets/Hyukkyu/beir-cqadupstack-programmers.texttext-retrieval10K<n<100K0 likes43 downloads1mo agoHugging Face06income /cqadupstack-programmers-top-20-gen-queries NFCorpus: 20 generated queries (BEIR Benchmark) This HF dataset contains the top-20 synthetic queries generated for each passage in the above BEIR benchmark dataset. DocT5query model used: BeIR/query-gen-msmarco-t5-base-v1 id (str): unique document id in NFCorpus in the BEIR benchmark (corpus.jsonl). Questions generated: 20 Code used for generation: evaluate_anserini_docT5query_parallel.py Below contains the old dataset card for the BEIR benchmark. Dataset Card for BEIR… See the full description on the dataset page: https://huggingface.co/datasets/income/cqadupstack-programmers-top-20-gen-queries.texttext-retrieval10K<n<100K0 likes39 downloads4y agoHugging Face07BaoLocTown /cqadupstack-programmers-vn-rawtext10K<n<100K0 likes37 downloads2y agoHugging Face08MCINext /cqadupstack-programmers-fa Dataset Summary CQADupstack-programmers-Fa is a Persian (Farsi) dataset developed for the Retrieval task, with a focus on duplicate question detection in community question-answering (CQA) platforms. This dataset is a translated version of the "Programmers" (Software Engineering) StackExchange subforum from the English CQADupstack collection and is part of the FaMTEB benchmark under the BEIR-Fa suite. Language(s): Persian (Farsi) Task(s): Retrieval (Duplicate Question Retrieval)… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/cqadupstack-programmers-fa.text10K<n<100K0 likes26 downloads1y agoHugging Face09mteb /CQADupstackProgrammersRetrieval-Fa CQADupstackProgrammersRetrieval-Fa An MTEB dataset Massive Text Embedding Benchmark CQADupstackProgrammersRetrieval-Fa Task category t2t Domains Web Reference https://huggingface.co/datasets/MCINext/cqadupstack-programmers-fa How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["CQADupstackProgrammersRetrieval-Fa"]) evaluator = mteb.MTEB(task) model =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/CQADupstackProgrammersRetrieval-Fa.texttext-retrieval10K<n<100K0 likes22 downloads1y agoHugging Face10dmrau /cqadupstack-programmers-qrels Dataset Card for "cqadupstack-programmers-qrels" More Information needed text1K<n<10K0 likes20 downloads3y agoHugging Face11orgrctera /beir_cqadupstack_programmers CQADupStack / Programmers (BEIR) — programming Q&A retrieval Dataset description CQADupStack is a benchmark for community question answering (cQA) built from Stack Exchange data. It was introduced by Hoogeveen, Verspoor, and Baldwin at ADCS 2015 to support research on duplicate questions: finding earlier posts that match or subsume a newly asked question, so users can reuse existing answers instead of opening redundant threads. The full CQADupStack release aggregates… See the full description on the dataset page: https://huggingface.co/datasets/orgrctera/beir_cqadupstack_programmers.texttext-retrievaln<1K0 likes19 downloads7mo agoHugging Face12orgrctera /beir_cqadupstack_programmers_test beir_cqadupstack_programmers_test BEIR CQADupStack/programmers test split Field Value Benchmark beir Sub-benchmark cqadupstack_programmers Type retrieval Items 876 Exported from Langfuse. textquestion-answeringn<1K0 likes12 downloads7mo agoHugging Face13dmrau /cqadupstack-programmers Dataset Card for "cqadupstack-programmers" More Information needed text10K<n<100K0 likes11 downloads3y agoHugging Face14clarin-knext /cqadupstack-programmers-pl-qrelsPart of BEIR-PL: Zero Shot Information Retrieval Benchmark for the Polish Language. Link to arxiv: https://arxiv.org/pdf/2305.19840.pdf Contact: konrad.wojtasik@pwr.edu.pl tabular1K<n<10K0 likes10 downloads2y agoHugging Face15mteb /CQADupstackProgrammers-NLtextn<1K0 likes10 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.