arianhosseini/mt_puzzles
Multi-Turn Puzzles: Evaluating Interactive Reasoning and Strategic Dialogue in LLMs MT-Puzzles is a novel benchmark comprising a suite of multi-turn tasks each designed to test specific reasoning, interactive dialogue, and information-seeking abilities: Word Guess Guess the secret word in min attempts while environment gives feedback on how close the guess is at each turn. Movie Recommendation: Probe the user to decode the user preference function for N turns. Pick a movie for⦠See the full description on the dataset page: https://huggingface.co/datasets/arianhosseini/mt_puzzles.
Multi-Turn Puzzles: Evaluating Interactive Reasoning and Strategic Dialogue in LLMs
MT-Puzzles is a novel benchmark comprising a suite of multi-turn tasks each designed to test specific reasoning, interactive dialogue, and information-seeking abilities:
- Word Guess Guess the secret word in min attempts while environment gives feedback on how close the guess is at each turn.
- Movie Recommendation: Probe the user to decode the user preference function for N turns. Pick a movie for the same user at N+1 turn.
- Circuit Decoding: Probe the C different boolean circuits for N turns. Predict the joint truth table of all the circuits at N+1 turn.
- Word Chaining: Model and user take turns choosing allow-listed words that start with the last letter of the previously chosen word.
- Twenty Questions: Model chooses a secret word. The user asks questions to determine what the word is.
Instructions and Last Turn Prompts
Each task's instruction and last turn prompt (if applicable) can be found in the text file in the task's directory.
Word Guess: Instruction π
Movie Recommendation: Instruction π and Last Turn Prompt π
Circuit Decoding: Instruction π and Last Turn Prompt π
Word Chaining: Instruction π
Twenty Questions: Instruction π
Paper: arxiv.org/abs/2508.10142
