kaustubhg73/multilingual-multitask-refusal
Multilingual Multitask Refusal A multilingual, multitask prompt dataset for analysing refusal behaviour across language, task wrapper, and harmful / harmless labels. English seed spans and labels come from the previous Multitask Multilingual Refusal dataset. Non-English content is produced with Google Sheets GOOGLETRANSLATE. Task instructions are language-localized via templates_localized.json. Rows 211,320 English seeds 1,761 Languages 15 Tasks 8 Product 1… See the full description on the dataset page: https://huggingface.co/datasets/kaustubhg73/multilingual-multitask-refusal.
Multilingual Multitask Refusal
A multilingual, multitask prompt dataset for analysing refusal behaviour across language, task wrapper, and harmful / harmless labels.
English seed spans and labels come from the previous Multitask Multilingual Refusal dataset. Non-English content is produced with Google Sheets GOOGLETRANSLATE. Task instructions are language-localized via templates_localized.json.
Languages
Tasks (intended_task)
English template skeletons (other languages use the localized shells in templates_localized.json):
Splits
Inherited from the English seed split membership in the previous dataset.
Schema
Load
from datasets import load_dataset
ds = load_dataset("kaustubhg73/multilingual-multitask-refusal")
print(ds)
row = ds["test"][0]
print(row["lang"], row["intended_task"], row["prompt"][:120])Filter example:
hi_count = ds["test"].filter(
lambda r: r["lang"] == "hi" and r["intended_task"] == "character_counting"
)Reproduce
Dependencies: pandas, pyarrow, datasets.
Work from this repository root (the folder that contains scripts/, templates_localized.json, and data/).
Step 1 — emit Google Sheets CSV
python scripts/dataset_pipeline.py --prepare-sheetsThis writes sheets/translate_formulas.csv (content_id, English en, and GOOGLETRANSLATE formulas for the 14 non-English languages).
- Upload the CSV to Google Drive and open it with Google Sheets.
- Wait until every language column shows resolved text (no leading
=). - Download as CSV and save as
sheets/translate_filled.csv.
Optional smoke: --limit N caps the number of English seeds.
Step 2 — import and wrap
python scripts/dataset_pipeline.py --import-sheets sheets/translate_filled.csvThis:
- Loads English seeds from the previous Multitask Multilingual Refusal dataset (same
content_idscheme). - Joins filled Sheet translations for all non-English languages.
- Applies localized task templates from
templates_localized.json. - Writes
build/multilingual-multitask-refusal.parquetand Hugging Face split files underdata/.
Repository layout
multilingual-multitask-refusal/
├── README.md
├── templates_localized.json
├── scripts/dataset_pipeline.py
├── data/ # train / val / test parquet
├── sheets/ # formula CSV + filled import (local build)
└── build/ # seeds + full parquet (local build)Source
- English seeds and labels: previous Multitask Multilingual Refusal dataset
- Non-English content: Google Sheets
GOOGLETRANSLATE - Task shells:
templates_localized.json(already presentlanguages from the previous dataset;newly addedforhi,id,tr,sw,ne)
