Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nyu-visionx /pisa-experiments Pisa Experiments This repository contains the PisaBench, training data, model checkpoints, introduced in PISA Experiments: Exploring Physics Post-Training for Video Diffusion Models by Watching Stuff Drop. PisaBench Real World Videos We curate a dataset comprising 361 videos demonstrating the dropping task.Each video begins with an object suspended by an invisible wire in the first frame. We cut the video clips to begin as soon as the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/pisa-experiments.n<1K2 likes1.9k downloads2y agoHugging Face02ServiceNow /PiSAs PiSAs: Benchmarking Contextual Integrity in Multi-User Agentic Systems Paper: arXiv:2607.05318 PiSAs (Privacy in Shared Agentic systems) is a benchmark for contextual privacy in multi-agent LLM systems. Each scenario puts an executor agent in an organisation, gives it a decision to make, and spreads the evidence it needs across colleagues — mixed in with private facts that are not legitimate inputs to that decision. A system is scored both on getting the decision right and on… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow/PiSAs.text-generation2 likes399 downloads23d agoHugging Face03pisa-engine /ciff-hub CIFF Hub Common Index File Format CIFF is an inverted index exchange format as defined as part of the Open-Source IR Replicability Challenge (OSIRRC) initiative. The Ciff Hub hosts many indexes and queries for a variety of collections and models. MS Marco The MS Marco passage ranking dataset consists of 8.8M passages. ESPLADE Lassance, Carlos, and Stéphane Clinchant. "An efficiency study for splade models." Proceedings of the 45th International ACM SIGIR… See the full description on the dataset page: https://huggingface.co/datasets/pisa-engine/ciff-hub.0 likes152 downloads8mo agoHugging Face04macavaney /msmarco-passage-v2-dedup.pisa msmarco-passage-v2-dedup.pisa Description TODO: What is the artifact? Usage # Load the artifact import pyterrier_alpha as pta artifact = pta.Artifact.from_hf('macavaney/msmarco-passage-v2-dedup.pisa') # TODO: Show how you use the artifact Benchmarks TODO: Provide benchmarks for the artifact. Reproduction # TODO: Show how you constructed the artifact. Metadata { "package_hint": "pyterrier_pisa" } text-retrieval0 likes108 downloads2y agoHugging Face05pyterrier /msmarco-passage-v2.splade-lg.pisa msmarco-passage-v2.splade-lg.pisa Description TODO: What is the artifact? Usage # Load the artifact import pyterrier as pt artifact = pt.Artifact.from_hf('pyterrier/msmarco-passage-v2.splade-lg.pisa') # TODO: Show how you use the artifact Benchmarks TODO: Provide benchmarks for the artifact. Reproduction # TODO: Show how you constructed the artifact. Metadata { "type": "sparse_index", "format":… See the full description on the dataset page: https://huggingface.co/datasets/pyterrier/msmarco-passage-v2.splade-lg.pisa.text-retrieval0 likes80 downloads3mo agoHugging Face06macavaney /cord19.pisa cord19.pisa Description TODO: What is the artifact? Usage # Load the artifact import pyterrier_alpha as pta artifact = pta.Artifact.from_hf('macavaney/cord19.pisa') # TODO: Show how you use the artifact Benchmarks TODO: Provide benchmarks for the artifact. Reproduction # TODO: Show how you constructed the artifact. Metadata { "type": "sparse_index", "format": "pisa", "package_hint": "pyterrier-pisa", "stemmer":… See the full description on the dataset page: https://huggingface.co/datasets/macavaney/cord19.pisa.text-retrieval0 likes62 downloads2y agoHugging Face07anusfoil /pisa-miditextn<1K1 likes62 downloads6mo agoHugging Face08DandyTian /ragwiki-splade-cocondenser-ensembledistil.pisa ragwiki-splade-cocondenser-ensembledistil.pisa Description TODO: What is the artifact? Usage # Load the artifact import pyterrier as pt artifact = pt.Artifact.from_hf('DandyTian/ragwiki-splade-cocondenser-ensembledistil.pisa') # TODO: Show how you use the artifact Benchmarks TODO: Provide benchmarks for the artifact. Reproduction # TODO: Show how you constructed the artifact. Metadata { "type":… See the full description on the dataset page: https://huggingface.co/datasets/DandyTian/ragwiki-splade-cocondenser-ensembledistil.pisa.text-retrieval1 likes56 downloads2d agoHugging Face09PisaBench /pisa-bench Dataset Card for PISA-Bench Paper: https://arxiv.org/abs/2510.24792Authors: Patrick Haller, Fabio Barth, Jonas Golde, Georg Rehm, Alan Akbik Dataset Summary PISA-Bench is a multilingual, multimodal benchmark constructed from expert-authored PISA exam questions.Each example is a human-created educational reasoning problem containing an image and a reading/math question, translated into six languages: English (EN) German (DE) Spanish (ES) French (FR) Italian (IT) Chinese… See the full description on the dataset page: https://huggingface.co/datasets/PisaBench/pisa-bench.imagen<1K1 likes46 downloads11mo agoHugging Face10macavaney /msmarco-passage-v2.pisa msmarco-passage-v2.pisa Description TODO: What is the artifact? Usage # Load the artifact import pyterrier as pt artifact = pt.Artifact.from_hf('macavaney/msmarco-passage-v2.pisa') # TODO: Show how you use the artifact Benchmarks TODO: Provide benchmarks for the artifact. Reproduction # TODO: Show how you constructed the artifact. Metadata { "type": "sparse_index", "format": "pisa", "package_hint": "pyterrier_pisa" } text-retrieval0 likes45 downloads2y agoHugging Face11pyterrier /scifact.pisa scifact.pisa Description A PISA index for the SciFact dataset Usage # Load the artifact import pyterrier as pt index = pt.Artifact.from_hf('pyterrier/scifact.pisa') index.bm25() # returns a BM25 retriever Benchmarks name nDCG@10 R@1000 bm25 0.6776 0.9733 dph 0.6735 0.97 Reproduction import pyterrier as pt from tqdm import tqdm import ir_datasets from pyterrier_pisa import PisaIndex index = PisaIndex("scifact.pisa"… See the full description on the dataset page: https://huggingface.co/datasets/pyterrier/scifact.pisa.text-retrieval0 likes38 downloads2y agoHugging Face12pyterrier /hotpotqa.pisa hotpotqa.pisa Description A PISA index for the Hotpot QA dataset Usage # Load the artifact import pyterrier as pt index = pt.Artifact.from_hf('pyterrier/hotpotqa.pisa') index.bm25() # returns a BM25 retriever Benchmarks hotpotqa/dev name nDCG@10 R@1000 bm25 0.6525 0.8909 dph 0.6445 0.8888 hotpotqa/test name nDCG@10 R@1000 bm25 0.6318 0.8851 dph 0.6246 0.8837 Reproduction import pyterrier as pt from… See the full description on the dataset page: https://huggingface.co/datasets/pyterrier/hotpotqa.pisa.text-retrieval0 likes37 downloads2y agoHugging Face13namawho /msmarco-segment-v2.1.pisa msmarco-segment-v2.1.pisa Description TODO: What is the artifact? Usage # Load the artifact import pyterrier as pt artifact = pt.Artifact.from_hf('namawho/msmarco-segment-v2.1.pisa') # TODO: Show how you use the artifact Benchmarks TODO: Provide benchmarks for the artifact. Reproduction # TODO: Show how you constructed the artifact. Metadata { "type": "sparse_index", "format": "pisa", "package_hint": "pyterrier-pisa"… See the full description on the dataset page: https://huggingface.co/datasets/namawho/msmarco-segment-v2.1.pisa.text-retrieval0 likes37 downloads1y agoHugging Face14pyterrier /fiqa.pisa fiqa.pisa Description A PISA index for the FIQA dataset Usage # Load the artifact import pyterrier as pt index = pt.Artifact.from_hf('pyterrier/fiqa.pisa') index.bm25() # returns a BM25 retriever Benchmarks fiqa/dev name nDCG@10 R@1000 bm25 0.263 0.7423 dph 0.2587 0.7497 fiqa/test name nDCG@10 R@1000 bm25 0.2411 0.7504 dph 0.2401 0.7615 Reproduction import pyterrier as pt from tqdm import tqdm import… See the full description on the dataset page: https://huggingface.co/datasets/pyterrier/fiqa.pisa.text-retrieval0 likes34 downloads2y agoHugging Face15pyterrier /fever.pisa fever.pisa Description A PISA index for the Fever dataset Usage # Load the artifact import pyterrier as pt index = pt.Artifact.from_hf('pyterrier/fever.pisa') index.bm25() # return a BM25 retriever Benchmarks fever/dev name nDCG@10 R@1000 bm25 0.6425 0.96 dph 0.6831 0.9604 fever/test name nDCG@10 R@1000 bm25 0.6305 0.9532 dph 0.6716 0.9573 Reproduction import pyterrier as pt from tqdm import tqdm… See the full description on the dataset page: https://huggingface.co/datasets/pyterrier/fever.pisa.text-retrieval0 likes32 downloads2y agoHugging Face16pyterrier /quora.pisa quora.pisa Description A PISA index for the Quora duplicate question dataset Usage # Load the artifact import pyterrier as pt index = pt.Artifact.from_hf('pyterrier/quora.pisa') index.bm25() # returns a BM25 retriever Benchmarks quora/dev name nDCG@10 R@1000 bm25 0.7195 0.9845 dph 0.5893 0.9711 quora/test name nDCG@10 R@1000 bm25 0.7122 0.9875 dph 0.5809 0.9729 Reproduction import pyterrier as pt from… See the full description on the dataset page: https://huggingface.co/datasets/pyterrier/quora.pisa.text-retrieval0 likes30 downloads2y agoHugging Face17pyterrier /trec-covid.pisa trec-covid.pisa Description A PISA Index for CORD19 (the corpus for the TREC-COVID query set) Usage # Load the artifact import pyterrier as pt index = pt.Artifact.from_hf('pyterrier/trec-covid.pisa') index.bm25() # returns a BM25 retriever Benchmarks name nDCG@10 R@1000 bm25 0.6254 0.4462 dph 0.6633 0.4136 Reproduction import pyterrier as pt from tqdm import tqdm import ir_datasets from pyterrier_pisa import PisaIndex… See the full description on the dataset page: https://huggingface.co/datasets/pyterrier/trec-covid.pisa.text-retrieval0 likes30 downloads2y agoHugging Face18pisakoz /parc2026-t2_v0040 likes27 downloads24d agoHugging Face19pyterrier /arguana.pisa arguana.pisa Description A PISA index for the Arguana dataset Usage # Load the artifact import pyterrier as pt index = pt.Artifact.from_hf('pyterrier/arguana.pisa') index.bm25() # returns a BM25 retriever Benchmarks name nDCG@10 R@1000 bm25 0.3436 0.9808 dph 0.3502 0.9815 Reproduction import pyterrier as pt from tqdm import tqdm import ir_datasets from pyterrier_pisa import PisaIndex index = PisaIndex("arguana.pisa"… See the full description on the dataset page: https://huggingface.co/datasets/pyterrier/arguana.pisa.text-retrieval0 likes26 downloads1y agoHugging Face20pyterrier-tutorial /bright.sustainable.splade.pisa bright.sustainable.splade.pisa Description TODO: What is the artifact? Usage # Load the artifact import pyterrier as pt artifact = pt.Artifact.from_hf('pyterrier-tutorial/bright.sustainable.splade.pisa') # TODO: Show how you use the artifact Benchmarks TODO: Provide benchmarks for the artifact. Reproduction # TODO: Show how you constructed the artifact. Metadata { "type": "sparse_index", "format":… See the full description on the dataset page: https://huggingface.co/datasets/pyterrier-tutorial/bright.sustainable.splade.pisa.text-retrieval0 likes23 downloads3mo agoHugging Face21pyterrier /nfcorpus.pisa nfcorpus.pisa Description A PISA index for the NFCorpus dataset Usage # Load the artifact import pyterrier as pt index = pt.Artifact.from_hf('pyterrier/nfcorpus.pisa') index.bm25() # returns a BM25 retriever Benchmarks nfcorpus/dev name nDCG@10 R@1000 bm25 0.2933 0.3299 dph 0.2912 0.334 nfcorpus/test name nDCG@10 R@1000 bm25 0.3271 0.3685 dph 0.3222 0.3672 Reproduction import pyterrier as pt from tqdm… See the full description on the dataset page: https://huggingface.co/datasets/pyterrier/nfcorpus.pisa.text-retrieval0 likes21 downloads2y agoHugging Face22macavaney /msmarco-passage.pisa MS MARCO PISA Index Description This is an index of the MS MARCO passage (v1) dataset with PISA. It can be used for passage retrieval using lexical methods. Usage >>> from pyterrier_pisa import PisaIndex >>> index = PisaIndex.from_hf('macavaney/msmarco-passage.pisa') >>> bm25 = index.bm25() >>> bm25.search('terrier breeds') qid query docno score rank 0 1 terrier breeds 1406578 22.686367 0 1 1 terrier breeds 5785957… See the full description on the dataset page: https://huggingface.co/datasets/macavaney/msmarco-passage.pisa.text-retrieval1 likes20 downloads2y agoHugging Face23pyterrier /scidocs.pisa scidocs.pisa Description A PISA index for the SciDocs dataset Usage # Load the artifact import pyterrier as pt index = pt.Artifact.from_hf('pyterrier/scidocs.pisa') index.bm25() # returns a BM25 retriever Benchmarks name nDCG@10 R@1000 bm25 0.1504 0.5637 dph 0.1512 0.5701 Reproduction import pyterrier as pt from tqdm import tqdm import ir_datasets from pyterrier_pisa import PisaIndex index = PisaIndex("scidocs.pisa"… See the full description on the dataset page: https://huggingface.co/datasets/pyterrier/scidocs.pisa.text-retrieval0 likes20 downloads2y agoHugging Face24kings-crown /FVELer_PISA_NotProventext1K<n<10K0 likes19 downloads2y agoHugging Face25kings-crown /FVELer_PISA_Proventext1K<n<10K0 likes17 downloads2y agoHugging Face26ClaireG /PiSAtext100K<n<1M0 likes16 downloads10mo agoHugging Face27kurtcobain1994 /PISA_20220 likes16 downloads3mo agoHugging Face28macavaney /my-index.pisa my-index.pisa Description TODO: What is the artifact? Usage # Load the artifact import pyterrier as pt artifact = pt.Artifact.from_hf('macavaney/my-index.pisa') # TODO: Show how you use the artifact Benchmarks TODO: Provide benchmarks for the artifact. Reproduction # TODO: Show how you constructed the artifact. Metadata { "type": "sparse_index", "format": "pisa", "package_hint": "pyterrier-pisa", "stemmer": "porter2" } text-retrieval0 likes15 downloads1y agoHugging Face29barthfab /PISA_tests Dataset Card: PISA Multimodal (Parallel & Not-Parallel) Summary This dataset contains 48 parallel multimodal samples (paired TXT↔PDF) derived from PISA studies up to 2012, plus 47 non-parallel samples (TXT-only or PDF-only). Each sample may include multiple questions. Content is available in German and English. Source & usage: Materials are published by the OECD and are provided here for non-commercial use only. Please verify that your usage complies with OECD terms.… See the full description on the dataset page: https://huggingface.co/datasets/barthfab/PISA_tests.documentn<1K0 likes15 downloads1y agoHugging Face30pyterrier /webis-touche2020.pisa webis-touche2020.pisa Description A PISA index for the Touche2020 dataset (version 2) Usage # Load the artifact import pyterrier as pt index = pt.Artifact.from_hf('pyterrier/webis-touche2020.pisa') index.bm25() # returns a BM25 retriever Benchmarks name nDCG@10 R@1000 bm25 0.6563 0.7163 dph 0.6785 0.7292 Reproduction import pyterrier as pt from tqdm import tqdm import ir_datasets from pyterrier_pisa import PisaIndex… See the full description on the dataset page: https://huggingface.co/datasets/pyterrier/webis-touche2020.pisa.text-retrieval0 likes14 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.