Toorintan/T-PID
In the name of Allah Sample of Toorintan-Persian Informal Dataset (T-PID) Sample data is available in the train, dev, and test folders. Corresponding metadata can be found in the train.tsv, dev.tsv, and test.tsv files, which are available for download. The full dataset is currently under review for publication as part of an academic paper. Once accepted, we will release the complete dataset for research purposes. π© To request access to the full dataset in the meantimeβ¦ See the full description on the dataset page: https://huggingface.co/datasets/Toorintan/T-PID.
In the name of Allah
Sample of Toorintan-Persian Informal Dataset (T-PID)
Sample data is available in the train, dev, and test folders. Corresponding metadata can be found in the train.tsv, dev.tsv, and test.tsv files, which are available for download.
The full dataset is currently under review for publication as part of an academic paper. Once accepted, we will release the complete dataset for research purposes.
π© To request access to the full dataset in the meantime, please email:
info@toorintan.com
---
# TPID Text Normalization Tool
This project provides a script for performing advanced text normalization on TSV files, with a focus on Persian (Farsi) and Arabic text. It supports converting numbers and time expressions to text, correcting formatting issues, and removing unwanted characters or symbols β making it ideal for preparing clean, consistent data for speech or NLP models. (For use this download Normalizer)
---
## π§ Features
- π Convert **numeric dates/times** (e.g., `2021/10/27`, `22:57:11`) into their **spoken equivalents**
- π Normalize **Arabic expressions** such as religious abbreviations (e.g., `(Ψ΅)`)
- βοΈ Correct **semi-spaces** in Persian (common writing error)
- β Replace **repeated punctuation** (e.g., `!!!` β `!`)
- β Remove **all punctuation** if needed
- π§Ή Strip **Persian/Arabic vowels** (diacritics) from the text
- π’ Convert **numbers** into full Persian word equivalents
---
## π Input Format
The script works on `.tsv` files (tab-separated values). It assumes there's a column containing text that needs normalization.
---
## π» Setup Instructions
> π Requires **Python 3.8**
### 1. Install Python 3.8 (if not already installed)
On Ubuntu:sudo apt update sudo apt install python3.8 python3.8-venv python3.8-dev
### 2. Create a Virtual Environment
python3.8 -m venv venv source venv/bin/activate
### 3. Install Requirements
pip install -r requirements.txt
---
## π How to Use
Run the script like this:
python Generalnormarg.py -i /path/to/input.tsv -o /path/to/output.tsv
### β
Example
python Generalnormarg.py -i /content/test2.tsv -o com_9.tsv
---
## βοΈ Command-Line Options
| Argument | Description |
| ------------------------------- | ---------------------------------------------- |
| `-i` | Input `.tsv` file (required) |
| `-o` | Output `.tsv` file (required) |
| `-time_to_text` | Convert times and dates to text |
| `-semi_space_correction` | Correct Persian semi-space usage |
| `-arabic_correction` | Normalize Arabic religious expressions |
| `-replace_repeated_punctuation` | Replace repeated punctuation characters |
| `-remove_punctuations` | Remove punctuation symbols from the text |
| `-remove_vowels` | Remove short vowels (diacritics) from text |
| `-convert_numbers_to_text` | Convert numbers (e.g., `123`) to Persian words |
All flags default to `True`. You can toggle them as needed.
---
## π License
This project is licensed under the **[CC BY-NC 4.0 License](https://creativecommons.org/licenses/by-nc/4.0/)**.
This means:
* β
You can use it for **academic and non-commercial purposes**
* β
You must **credit the original author**
* β You may **not use it commercially** without permission
---
## π§βπ» Maintainer
To get access to the full dataset, please email us. For questions, issues, or contributions, feel free to open an issue or contact the maintainer via Hugging Face.
Email: info@toorintan.com
