Team Ai
Datasetpublic

google/deepsearchqa

DeepSearchQA A 900-prompt factuality benchmark from Google DeepMind, designed to evaluate agents on difficult multi-step information-seeking tasks across 17 different fields. ▶ Google DeepMind Release Blog Post▶ DeepSearchQA Leaderboard on Kaggle▶ Technical Report▶ Evaluation Starter Code Benchmark DeepSearchQA is a 900-prompt benchmark for evaluating agents on difficult multi-step information-seeking tasks across 17 different fields. Unlike… See the full description on the dataset page: https://huggingface.co/datasets/google/deepsearchqa.

sourceHugging Faceapache-2.0updated 10mo agoView on Hugging Face
132likes22kdownloads
4 commits on main
b2623f810mo ago

Update README.md

geolocal
75c44ab10mo ago

Update README.md

geolocal
042aaaf10mo ago

Upload DSQA-full.csv

geolocal
286f28c10mo ago

initial commit

geolocal