Team Ai
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hotchpotch /mmarco-hard-negatives-reranker-filtered mMARCO Reranker-Filtered Hard Negatives (Multilingual) Overview This dataset is built from mMARCO (multilingual MS MARCO) triplets for each language subset. For each (query, positive), hard negatives are bundled and then filtered using cross-encoder re-scoring. The goal is to remove negatives that are too strong or incorrect for training. The same procedure is applied to all language subsets. The dataset is published as mmarco-hard-negatives-reranker-filtered with… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/mmarco-hard-negatives-reranker-filtered.tabular10M<n<100M3 likes1.3k downloads4mo agoHugging Face02hotchpotch /hpprc_emb_reranker_score ⚠️ お知らせ よりスコア付したデータ件数とrerankerのバリエーションを増やしたデータセットのhotchpotch/hpprc_emb-scoresも公開しています。 hpprc/emb (便利なデータセットの公開、ありがとうございます)の collection と dataset がペアになっているデータに対し、negative を最大32個ランダムサンプリングしたものを、hotchpotch/japanese-bge-reranker-v2-m3-v1でスコア付けしたものです。 ライセンスは、subset ごとに hpprc/emb に記載のライセンスと同等とします。 スコア作成タイミングの revision に対してスコアを付与しているため、revision を変えると場合によって行ズレやデータ構造の変化が発生する可能性があることに注意が必要です。 例 from datasets import load_dataset # targets = ("auto-wiki-qa", "4feb2e2492")… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/hpprc_emb_reranker_score.tabular1M<n<10M4 likes453 downloads2y agoHugging Face03lightblue /reranker_continuous_filt_max7_train Reranker training data This data was generated using 4 steps: We gathered queries and corresponding text data from 35 high quality datasets covering more than 95 languages. For datasets which did not already have negative texts for queries, we mined hard negatives using the BAAI/bge-m3 embedding model. For each query, we selected one positive and one negative text and used Qwen/Qwen2.5-32B-Instruct-GPTQ-Int4 to rate the relatedness of each query-text pair using a token "1", "2"… See the full description on the dataset page: https://huggingface.co/datasets/lightblue/reranker_continuous_filt_max7_train.tabular1M<n<10M8 likes214 downloads2y agoHugging Face04manu /reranker-scorestabularn<1K0 likes81 downloads3y agoHugging Face05lightblue /reranker_fulltrain_scored_filteredtabular1M<n<10M0 likes72 downloads2y agoHugging Face06olaverse /reranker-general-en-llm-judged Reranker / Retriever Training Set (General-Purpose English, LLM-Judged) Graded query-passage relevance for training rerankers (cross-encoders) and retrievers (bi-encoders, ColBERT). English, commercial-use sources only. Configs pairs-graded (844k train / 8.5k test): query, passage, llm_grade (0-3), teacher_score, source, mining_method, role. For cross-encoder / reranker training. triplets (82k train / 819 test): query, positive, negative_1..5 with teacher scores.… See the full description on the dataset page: https://huggingface.co/datasets/olaverse/reranker-general-en-llm-judged.tabulartext-ranking100K<n<1M0 likes33 downloads4mo agoHugging Face07lgienapp /msmarco-triplets-qwen3-rerankertabular10M<n<100M0 likes24 downloads7mo agoHugging Face08maennyn /cqadupstack-reranker-datatabular10K<n<100K0 likes12 downloads1y agoHugging Face09Lo /rerankers-and-lexical-similarities Dataset Card for Re-ranker Evaluation Datasets This repo contains the evaluation datasets used in the paper "Language Model Re-rankers are Fooled by Lexical Similarities" accepted to FEVER 2025. Dataset Details The datasets in this repo are based on the NQ, LitQA2 (from LAB-Bench) and DRUID datasets. More details on the datasets can be found in our paper. Uses Evaluate re-rankers. Dataset Structure We release the NQ, LitQA2 and DRUID… See the full description on the dataset page: https://huggingface.co/datasets/Lo/rerankers-and-lexical-similarities.tabular10K<n<100K0 likes10 downloads1y agoHugging Face10ainewtrend07 /Evaluation_Qwen-Qwen3-Reranker-0.6Btabular10K<n<100K0 likes9 downloads1y agoHugging Face11iscchang /rag-ko-bge-reranker-v2-m3-top5-answer-evaltabularn<1K0 likes8 downloads1y agoHugging Face12UlrickBL /mm_reranker_rl_trainingimagen<1K0 likes7 downloads1y agoHugging Face13frankie137 /speaker_reranker_datatabular1M<n<10M0 likes7 downloads7mo agoHugging Face14lgienapp /msmarco-groups-qwen3-rerankertabular100K<n<1M0 likes7 downloads6mo agoHugging Face15Algokruti /thread-reranker-datatabular10K<n<100K0 likes7 downloads6mo agoHugging Face16iscchang /rag-ko-jina-reranker-v2-base-multilingual-top5-answer-evaltabularn<1K0 likes5 downloads1y agoHugging Face17ainewtrend07 /Evaluation_LoraQwen-Qwen3-Reranker-0.6Btabular10K<n<100K0 likes5 downloads1y agoHugging Face18UlrickBL /mm_reranker_rl_trainimage1K<n<10K0 likes5 downloads1y agoHugging Face19iscchang /rag-ko-bge-reranker-v2-m3-ko-top5-answer-evaltabularn<1K0 likes4 downloads1y agoHugging Face20juliadollis /energy-eval-rag-reranker20to5tabularn<1K0 likes4 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.