Team Ai
Datasetpublic

StringNLP/longharness

LongHarness Bench: Stress-Testing Language Model Harnesses for Long-Context Reasoning LongHarness evaluates how language-model harnesses access and reason over long contexts. It is designed to distinguish context-access strategies, including direct reading, lexical and semantic retrieval, iterative agents, and recursive language-model harnesses. The benchmark contains 200 evaluation instances across four task suites. Project website Paper GitHub repository… See the full description on the dataset page: https://huggingface.co/datasets/StringNLP/longharness.

sourceHugging Faceupdated 10d agoView on Hugging Face
0likes2.4kdownloads
10 commits on main
95c54a110d ago

Add LongHarness logo to dataset card

quanghieupham
2da743110d ago

Link LongHarness arXiv paper

quanghieupham
952cce511d ago

Match dataset title to paper

quanghieupham
ef7841111d ago

Normalize dataset card formatting

quanghieupham
5fb868511d ago

Prepare dataset card for public release

quanghieupham
1aed8d612d ago

Align constraint task with match-only scoring

quanghieupham
59b2ab112d ago

Add dataset card metadata

quanghieupham
43729b212d ago

Publish LongHarness benchmark snapshot (part 2)

quanghieupham
6cf7c9712d ago

Publish LongHarness benchmark snapshot

quanghieupham
01275b512d ago

initial commit

quanghieupham