StringNLP/longharness
LongHarness Bench: Stress-Testing Language Model Harnesses for Long-Context Reasoning LongHarness evaluates how language-model harnesses access and reason over long contexts. It is designed to distinguish context-access strategies, including direct reading, lexical and semantic retrieval, iterative agents, and recursive language-model harnesses. The benchmark contains 200 evaluation instances across four task suites. Project website Paper GitHub repository… See the full description on the dataset page: https://huggingface.co/datasets/StringNLP/longharness.
This repository belongs to StringNLP on Hugging Face.
Team Ai never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
