datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
scriptspali-tripitaka-thai-script-siamrath-version
📚 พระไตรปิฎกภาษาบาลี อักษรไทยฉบับสยามรัฐ (๔๕ เล่ม)
ต้นฉบับสามารถเข้าถึงได้ที่: 84000 พระธรรมขันธ์ 84000.org พระไตรปิฎก
🧾 รายการพระไตรปิฎก
📘 เล่ม ๑–๘: วินัยปิฎก
เล่ม ๑: มหาวิภงฺโค (๑)
เล่ม ๒: มหาวิภงฺโค (๒)
เล่ม ๓: ภิกฺขุนีวิภงฺโค
เล่ม ๔: มหาวคฺโค (๑)
เล่ม ๕: มหาวคฺโค (๒)
เล่ม ๖: จุลฺลวคฺโค (๑)
เล่ม ๗: จุลฺลวคฺโค (๒)
เล่ม ๘: ปริวาโร
📗 เล่ม ๙–๒๕: สุตตันตปิฎก
เล่ม ๙–๑๑: ทีฆนิกาย
เล่ม ๑๒–๑๔: มัชฌิมนิกาย
เล่ม ๑๕–๑๙: สังยุตตนิกาย… See the full description on the dataset page: https://huggingface.co/datasets/mgprogm/pali-tripitaka-thai-script-siamrath-version.pali-commentary-thai-script-siamrath-version
Multi-File CSV Dataset
คำอธิบาย
อรรถกถาบาลี อักษรไทยฉบับสยามรัฏฐ จำนวน ๔๘ เล่ม
ชุดข้อมูลนี้ประกอบด้วยไฟล์ CSV หลายไฟล์
01/010001.csv: เล่ม 1 หน้า 1
01/010002.csv: เล่ม 1 หน้า 2
...
02/020001.csv: เล่ม 2 หน้า 1
คำอธิบายของแต่ละเล่ม
เล่ม ๑: วินยฏฺกถา (สมนฺตปาสาทิกา ๑)
เล่ม ๒: วินยฏฺกถา (สมนฺตปาสาทิกา ๒)
เล่ม ๓: วินยฏฺกถา (สมนฺตปาสาทิกา ๓)
เล่ม ๔: ทีฆนิกายฏฺกถา (สุมงฺคลวิลาสินี ๑)
เล่ม ๕: ทีฆนิกายฏฺกถา (สุมงฺคลวิลาสินี ๒)
เล่ม ๖: ทีฆนิกายฏฺกถา… See the full description on the dataset page: https://huggingface.co/datasets/uisp/pali-commentary-thai-script-siamrath-version.pali-tripitaka-thai-script-siamrath-version
Multi-File CSV Dataset
คำอธิบาย
พระไตรปิฎกภาษาบาลี อักษรไทยฉบับสยามรัฏฐ จำนวน ๔๕ เล่ม
ชุดข้อมูลนี้ประกอบด้วยไฟล์ CSV หลายไฟล์
01/010001.csv: เล่ม 1 หน้า 1
01/010002.csv: เล่ม 1 หน้า 2
...
02/020001.csv: เล่ม 2 หน้า 1
...
คำอธิบายของแต่ละเล่ม
เล่ม ๑: วินย. มหาวิภงฺโค (๑)
เล่ม ๒: วินย. มหาวิภงฺโค (๒)
เล่ม ๓: วินย. ภิกฺขุนีวิภงฺโค
เล่ม ๔: วินย. มหาวคฺโค (๑)
เล่ม ๕: วินย. มหาวคฺโค (๒)
เล่ม ๖: วินย. จุลฺลวคฺโค (๑)
เล่ม ๗: วินย. จุลฺลวคฺโค (๒)
เล่ม ๘: วินย.… See the full description on the dataset page: https://huggingface.co/datasets/uisp/pali-tripitaka-thai-script-siamrath-version.script-fidelity-benchmark
Script fidelity benchmark
Anonymous supplement for the paper "Script collapse in multilingual ASR:
A reference-free metric and 100-pair benchmark."
Script Fidelity Rate (SFR) measures the fraction of ASR hypothesis characters
that belong to the expected target script. WER measures word edits, while SFR
checks whether the output is written in the target orthography.
Related resources:
PyPI package: https://pypi.org/project/script-fidelity/
Hugging Face Evaluate metric:… See the full description on the dataset page: https://huggingface.co/datasets/themechanism/script-fidelity-benchmark.tipitaka_pali_in_15_scripts
Tipitaka Pali in 15 Scripts
Dataset Summary
This dataset contains the Pali Tipitaka (The Pali Canon), Commentaries (Aṭṭhakathā), Sub-commentaries (Ṭīkā), and related texts (Anya). It covers the fundamental scriptures of Theravada Buddhism.
This repository serves as a Hugging Face mirror and processed version of the open-source XML data provided by the Vipassana Research Institute (VRI). The texts are available in various scripts (including Roman and Myanmar) and are… See the full description on the dataset page: https://huggingface.co/datasets/freococo/tipitaka_pali_in_15_scripts.NLLB-Seed_Tamasheq-Tifinagh-ScriptSinhala-Script-LangID-BenchmarkSCRIPTS_results_newTREAT-Script-DatabseThe Largest Database of Movies and Shows Scripts sourced from IMSDB
Used for TREAT (Trigger Recognition for Enjoyable and Appropriate Television)
All data is (mostly) updated until Febuary 2025
license: apache-2.0
rick-and-morty-scripts-llama-2sindhi-parallel-scripts
Sindhi Parallel Scripts Dataset
Dataset Description
This dataset is a parallel-script dataset for Sindhi, containing the same Sindhi sentences represented in multiple writing systems and transliteration schemes.
The dataset contains four corresponding columns:
devanagari: Sindhi sentences written in Devanagari script. This serves as the ground-truth source representation.
persoarabic: Sindhi sentences represented in the Perso-Arabic script.
khudabadi: Sindhi… See the full description on the dataset page: https://huggingface.co/datasets/fahadqazi/sindhi-parallel-scripts.ACT_therapy_scriptsExamples patient conversations using ACT (Acceptance & Commitment therapy). Scripts come from 17 ACT books.
NLLB-Seed_Tamasheq-Latin-Scripttest-ds-script-ssrf
Test
vietnamese_nom_scriptsa10_command_scriptscriptscriptum250521-scriptumBlueArchive-Aris-Scriptsscripts
