Team Ai
Datasetpublic

Center-Of-Advanced-Software-Technologies/IFEval-Multi-IF-hy

IFEval-Multi-IF-hy — Armenian IFEval & Multi-IF Dataset Summary We introduce IFEval & Multi-IF hy, an Armenian extension of Multi-IF, the benchmark for assessing LLMs' proficiency in following multi-turn and multilingual instructions. Multi-IF itself extends IFEval to multi-turn, multilingual conversations. Multi-IF covers eight languages and does not include Armenian. To build IFEval & Multi-IF hy, the English conversations were first split into two groups:… See the full description on the dataset page: https://huggingface.co/datasets/Center-Of-Advanced-Software-Technologies/IFEval-Multi-IF-hy.

sourceHugging Facecc-by-nc-2.0updated 11h agoView on Hugging Face
0likes30downloads
Dataset Card

IFEval-Multi-IF-hy — Armenian IFEval & Multi-IF

Dataset Summary

We introduce IFEval & Multi-IF hy, an Armenian extension of Multi-IF, the benchmark for assessing LLMs' proficiency in following multi-turn and multilingual instructions. Multi-IF itself extends IFEval to multi-turn, multilingual conversations. Multi-IF covers eight languages and does not include Armenian. To build IFEval & Multi-IF hy, the English conversations were first split into two groups: culturally neutral conversations, translated into Armenian using an agentic translation and review pipeline that combines LLM-based processing with human-expert review, and culturally sensitive conversations, adapted into Armenian manually by human experts. This results in 884 conversations. Each conversation carries an is_adapted flag indicating whether it was culturally adapted or translated directly without cultural modification: 235 were adapted and 649 were translated directly.

For each turn of the directly translated conversations, annotators were shown the English source and Armenian translation side by side and rated it using three labels:

  • —Good — the translation is correct and natural Armenian.
  • —Bad — the translation is substantially incorrect or fails to preserve the meaning of the English source.
  • —Bad Armenian — the meaning is correct, but the Armenian is unnatural, poorly formed, or otherwise does not read like natural Armenian.

Annotators could also leave comments identifying minor issues or suggesting small fixes needed to make an otherwise acceptable translation fully natural and correct.

Data Fields

  • —turns: Placehold for saving the history conversation in evaluation.
  • —responses: Placehold for saving the latest response in evaluation.
  • —turn_1_prompt: The user prompt at the first turn, which is the input for LLM generation.
  • —turn_1_instruction_id_list: The instructions of the user prompt at the first turn, which is needed in the evaluation script.
  • —turn_1_kwargs: The arguments of the first turn instructions, which is needed in the evaluation script.
  • —turn_2_prompt: The user prompt at the second turn, which is the input for LLM generation.
  • —turn_2_instruction_id_list: The instructions of the user prompt at the second turn, which is needed in the evaluation script.
  • —turn_2_kwargs: The arguments of the second turn instructions, which is needed in the evaluation script.
  • —turn_3_prompt: The user prompt at the third turn, which is the input for LLM generation. Null for the 12 conversations with no third turn.
  • —turn_3_instruction_id_list: The instructions of the user prompt at the third turn, which is needed in the evaluation script.
  • —turn_3_kwargs: The arguments of the third turn instructions, which is needed in the evaluation script.
  • —key: The key of each conversation
  • —turn_index: Placehold for saving the current turn index in evaluation.
  • —language: The language of each conversation
  • —is_adapted: True for the 235 conversations culturally adapted for an Armenian audience; False for the 649 translated directly.

The constraint values inside turn_*_kwargs (end_phrase, keywords, postscript_marker, section_spliter, ...) are in Armenian, as in every other Multi-IF language; numeric values (word counts, letter frequencies) match the English source.

Data Splits

SplitConversationsThree-turnTwo-turnInstructions
test884872126,647

How this dataset was built

Here's our pipeline, from the 909 English conversations to the 884 released — see the steps below:

[image]

Cultural adaptation

Some Multi-IF prompts touch on a cultural aspect - a place, institution, custom, or incidental name tied to a particular culture. The 248 flagged conversations were rewritten by human experts, replacing that reference with an Armenian equivalent. The instruction being tested is never changed - an adapted prompt tests the same constraint on an Armenian subject. Example (key 1203:9:hy, turn 1):

English: What happened when the Tang dynasty of China was in power? Make sure to use the word war at least 8 times, and the word peace at least 10 times. Armenian, adapted: Ի՞նչ պատահեց, երբ Հայաստանում թագավորում էր Բագրատունիների արքայատոհմը։ Համոզվի՛ր, որ «պատերազմ» բառը օգտագործում ես առնվազն 8 անգամ, իսկ «խաղաղություն» բառը՝ առնվազն 10 անգամ։

Benchmark results

56 model configurations evaluated on the full 884-conversation set with the upstream Multi-IF metrics. All eight are 0–1 scores where higher is better (↑).

Each model's scores are also broken out by the is_adapted split, using the dataset's own vocabulary:

  • —Culturally Adapted — scored only on the 235 culturally adapted conversations.
  • —Culturally Neutral — scored only on the 649 directly translated conversations, with no cultural rewrite.
  • —Merged — scored on all 884 conversations together; this is the headline number used to rank the leaderboard.
MetricMeaning
P-strPrompt-level, strict: fraction of conversations where every instruction was followed, checked exactly as written.
I-strInstruction-level, strict: fraction of individual instructions followed, checked exactly as written.
P-looPrompt-level, loose: same as P-str, but forgiving of minor formatting (e.g. markdown wrappers, extra whitespace).
I-looInstruction-level, loose: same as I-str, but forgiving of minor formatting.
OverallAverage of the four scores above — the headline number, used to sort the leaderboard.
T1 / T2 / T3Overall, computed separately on turn 1, 2 and 3 of each conversation — shows how accuracy degrades as the conversation goes on.

[image]

👉 See the full results on the [IFEval & Multi-IF hy Leaderboard](https://huggingface.co/spaces/Center-Of-Advanced-Software-Technologies/IFEval-Multi-IF-hy-Leaderboard).

Acknowledgements

IFEval & Multi-IF hy is a joint work of Center for Advanced Software Technologies (CAST) and Armenian Language Technology Research Laboratory (RLALT).

The research was supported by the Higher Education and Science Committee of MESCS RA (Research project № 27TARGET-6B173).