Team Ai
Datasetpublic

Algorithmic-Human-Development-Group/Multilingual-Therapy-Dialogues

Dataset Summary Multilingual Therapy Dialogues is a diverse and bilingual dataset consisting of paired dialogues between patients and therapists in both Persian and English. Dataset Statistics Number of samples: 7,179 English: Average tokens per sentence: 101.30 Maximum tokens in a sentence: 939 Average characters per sentence: 567.85 Number of unique tokens: 32,968 Persian: Average tokens per sentence: 100.06 Maximum tokens in a sentence: 1,413… See the full description on the dataset page: https://huggingface.co/datasets/Algorithmic-Human-Development-Group/Multilingual-Therapy-Dialogues.

sourceHugging Facemitupdated 1y agoView on Hugging Face
0likes33downloads
README.md89 linesDownload Raw Back to root
1---2license: mit3language:4  - en5  - fa6tags:7  - therapy8pretty_name: Multilingual Therapy Dialogues9size_categories:10  - 1K<n<10K11---12 13[![arXiv](https://img.shields.io/badge/arXiv-Paper-<COLOR>.svg)]() [![GitHub](https://img.shields.io/badge/GitHub-Code-181717?logo=github)]()14 15## Dataset Summary16**Multilingual Therapy Dialogues** is a diverse and bilingual dataset consisting of paired dialogues between patients and therapists in both Persian and English.17 18## Dataset Statistics19- Number of samples: 7,179  20 21**English:**221. Average tokens per sentence: 101.30  232. Maximum tokens in a sentence: 939  243. Average characters per sentence: 567.85  254. Number of unique tokens: 32,968  26 27**Persian:**281. Average tokens per sentence: 100.06  292. Maximum tokens in a sentence: 1,413  303. Average characters per sentence: 516.57  314. Number of unique tokens: 33,298  32 33## Dataset Fields341. **Patient**: Original English text spoken by the patient.  352. **Therapist**: Original English text spoken by the therapist.  363. **Translated Patient**: Persian translation of the patient's text.  374. **Translated Therapist**: Persian translation of the therapist's text.  38 39## Dataset Generation Pipeline40 41The dataset was constructed using the following steps:42 431. **Data Collection**: Dialogues were collected from various public sources, including:44   - [Mental Health Counseling Conversations](https://huggingface.co/datasets/Amod/mental_health_counseling_conversations)45   - [Mental Health CSV Dataset](https://www.kaggle.com/datasets/zuhairhasanshaik/datacsv)46   - [Mental Health Conversational Data](https://www.kaggle.com/datasets/elvis23/mental-health-conversational-data)47   - Additional manually curated sources48 492. **Translation**: English dialogues were translated into Persian using the [SeamlessM4T model](https://github.com/facebookresearch/seamless_communication) by Meta AI.50 513. **Refinement**: Translations were revised and enhanced in three steps using GPT-4o:52   - First pass to make the tone more natural and emotionally sympathetic to be more likely to real world scenarios.53   - Second pass to improve fluency and human-likeness.54   - Final pass for consistency and correction of subtle translation errors.55 564. **Filtering**: Only meaningful and conte57 58 59## Usage Instructions60 61### Option 1: Manual Download62 63Visit the [dataset repository](https://huggingface.co/datasets/Algorithmic-Human-Development-Group/Multilingual-Therapy-Dialogues/tree/main) and download the `SAT_dataset.csv` file.64 65### Option 2: Programmatic Download66 67Use the `huggingface_hub` library to download the dataset programmatically:68 69```python70from huggingface_hub import hf_hub_download71import pandas as pd72 73dataset = hf_hub_download(74    repo_id="Algorithmic-Human-Development-Group/Multilingual-Therapy-Dialogues",75    filename="SAT_dataset.csv",76    repo_type="dataset"77)78df = pd.read_csv(dataset)79df.head()80```81 82## Citations83If you find our paper, code, data, or models useful, please cite the paper:  84```85To be updated once the paper is published.86```87 88## Contact89If you have questions, please email sinaaelahimanesh@gmail.com or mahdi.abootorabi2@gmail.com.