Algorithmic-Human-Development-Group/Multilingual-Therapy-Dialogues
Dataset Summary Multilingual Therapy Dialogues is a diverse and bilingual dataset consisting of paired dialogues between patients and therapists in both Persian and English. Dataset Statistics Number of samples: 7,179 English: Average tokens per sentence: 101.30 Maximum tokens in a sentence: 939 Average characters per sentence: 567.85 Number of unique tokens: 32,968 Persian: Average tokens per sentence: 100.06 Maximum tokens in a sentence: 1,413… See the full description on the dataset page: https://huggingface.co/datasets/Algorithmic-Human-Development-Group/Multilingual-Therapy-Dialogues.
033
1---2license: mit3language:4 - en5 - fa6tags:7 - therapy8pretty_name: Multilingual Therapy Dialogues9size_categories:10 - 1K<n<10K11---12 13[]() []()14 15## Dataset Summary16**Multilingual Therapy Dialogues** is a diverse and bilingual dataset consisting of paired dialogues between patients and therapists in both Persian and English.17 18## Dataset Statistics19- Number of samples: 7,179 20 21**English:**221. Average tokens per sentence: 101.30 232. Maximum tokens in a sentence: 939 243. Average characters per sentence: 567.85 254. Number of unique tokens: 32,968 26 27**Persian:**281. Average tokens per sentence: 100.06 292. Maximum tokens in a sentence: 1,413 303. Average characters per sentence: 516.57 314. Number of unique tokens: 33,298 32 33## Dataset Fields341. **Patient**: Original English text spoken by the patient. 352. **Therapist**: Original English text spoken by the therapist. 363. **Translated Patient**: Persian translation of the patient's text. 374. **Translated Therapist**: Persian translation of the therapist's text. 38 39## Dataset Generation Pipeline40 41The dataset was constructed using the following steps:42 431. **Data Collection**: Dialogues were collected from various public sources, including:44 - [Mental Health Counseling Conversations](https://huggingface.co/datasets/Amod/mental_health_counseling_conversations)45 - [Mental Health CSV Dataset](https://www.kaggle.com/datasets/zuhairhasanshaik/datacsv)46 - [Mental Health Conversational Data](https://www.kaggle.com/datasets/elvis23/mental-health-conversational-data)47 - Additional manually curated sources48 492. **Translation**: English dialogues were translated into Persian using the [SeamlessM4T model](https://github.com/facebookresearch/seamless_communication) by Meta AI.50 513. **Refinement**: Translations were revised and enhanced in three steps using GPT-4o:52 - First pass to make the tone more natural and emotionally sympathetic to be more likely to real world scenarios.53 - Second pass to improve fluency and human-likeness.54 - Final pass for consistency and correction of subtle translation errors.55 564. **Filtering**: Only meaningful and conte57 58 59## Usage Instructions60 61### Option 1: Manual Download62 63Visit the [dataset repository](https://huggingface.co/datasets/Algorithmic-Human-Development-Group/Multilingual-Therapy-Dialogues/tree/main) and download the `SAT_dataset.csv` file.64 65### Option 2: Programmatic Download66 67Use the `huggingface_hub` library to download the dataset programmatically:68 69```python70from huggingface_hub import hf_hub_download71import pandas as pd72 73dataset = hf_hub_download(74 repo_id="Algorithmic-Human-Development-Group/Multilingual-Therapy-Dialogues",75 filename="SAT_dataset.csv",76 repo_type="dataset"77)78df = pd.read_csv(dataset)79df.head()80```81 82## Citations83If you find our paper, code, data, or models useful, please cite the paper: 84```85To be updated once the paper is published.86```87 88## Contact89If you have questions, please email sinaaelahimanesh@gmail.com or mahdi.abootorabi2@gmail.com.