datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pali-commentary-thai-script-siamrath-version
Multi-File CSV Dataset
คำอธิบาย
อรรถกถาบาลี อักษรไทยฉบับสยามรัฏฐ จำนวน ๔๘ เล่ม
ชุดข้อมูลนี้ประกอบด้วยไฟล์ CSV หลายไฟล์
01/010001.csv: เล่ม 1 หน้า 1
01/010002.csv: เล่ม 1 หน้า 2
...
02/020001.csv: เล่ม 2 หน้า 1
คำอธิบายของแต่ละเล่ม
เล่ม ๑: วินยฏฺกถา (สมนฺตปาสาทิกา ๑)
เล่ม ๒: วินยฏฺกถา (สมนฺตปาสาทิกา ๒)
เล่ม ๓: วินยฏฺกถา (สมนฺตปาสาทิกา ๓)
เล่ม ๔: ทีฆนิกายฏฺกถา (สุมงฺคลวิลาสินี ๑)
เล่ม ๕: ทีฆนิกายฏฺกถา (สุมงฺคลวิลาสินี ๒)
เล่ม ๖: ทีฆนิกายฏฺกถา… See the full description on the dataset page: https://huggingface.co/datasets/uisp/pali-commentary-thai-script-siamrath-version.tripitaka-siamrath
Multi-File CSV Dataset
คำอธิบาย
พระไตรปิฎกภาษาไทยฉบับสยามรัฏฐ จำนวน 45 เล่ม
ชุดข้อมูลนี้ประกอบด้วยไฟล์ CSV หลายไฟล์
01/010001.csv: เล่ม 1 หน้า 1
01/010002.csv: เล่ม 1 หน้า 2
...
02/020001.csv: เล่ม 2 หน้า 1
คำอธิบายของแต่ละเล่ม
เล่ม 1 (754 หน้า): พระวินัยปิฎก เล่ม ๑ มหาวิภังค์ ปฐมภาค
เล่ม 2 (717 หน้า): พระวินัยปิฎก เล่ม ๒ มหาวิภังค์ ทุติภาค
เล่ม 3 (328 หน้า): พระวินัยปิฎก เล่ม ๓ ภิกขุณี วิภังค์
เล่ม 4 (304 หน้า): พระวินัยปิฎก เล่ม ๔ มหาวรรคภาค ๑
เล่ม 5 (278… See the full description on the dataset page: https://huggingface.co/datasets/uisp/tripitaka-siamrath.vihsd
Dataset Card for Dataset Name
ViHSD - Vietnamese Hate Speech Detection Dataset
Dataset Details
Dataset Description
The dataset contains about 33K annotated comments from social networks. Each has one of three labels: HATE, OFFENSIVE, CLEAN
Uses
Use directly from the Hugging face dataset loader
Direct Use
from datasets import load_dataset
train = load_dataset("sonlam1102/vihsd", split="train")
dev = load_dataset("sonlam1102/vihsd"… See the full description on the dataset page: https://huggingface.co/datasets/uitnlp/vihsd.pali-tripitaka-thai-script-siamrath-version
Multi-File CSV Dataset
คำอธิบาย
พระไตรปิฎกภาษาบาลี อักษรไทยฉบับสยามรัฏฐ จำนวน ๔๕ เล่ม
ชุดข้อมูลนี้ประกอบด้วยไฟล์ CSV หลายไฟล์
01/010001.csv: เล่ม 1 หน้า 1
01/010002.csv: เล่ม 1 หน้า 2
...
02/020001.csv: เล่ม 2 หน้า 1
...
คำอธิบายของแต่ละเล่ม
เล่ม ๑: วินย. มหาวิภงฺโค (๑)
เล่ม ๒: วินย. มหาวิภงฺโค (๒)
เล่ม ๓: วินย. ภิกฺขุนีวิภงฺโค
เล่ม ๔: วินย. มหาวคฺโค (๑)
เล่ม ๕: วินย. มหาวคฺโค (๒)
เล่ม ๖: วินย. จุลฺลวคฺโค (๑)
เล่ม ๗: วินย. จุลฺลวคฺโค (๒)
เล่ม ๘: วินย.… See the full description on the dataset page: https://huggingface.co/datasets/uisp/pali-tripitaka-thai-script-siamrath-version.UiPad
UiPad - UI Parsing and Accessibility Dataset
📌 Dataset status: stable release. UiPad was built for the IASA Champ 2024 Challenge and is a complete, fixed research artifact. No further updates are planned.
Curated by: MacPaw Way Ltd.
Language(s): Mostly EN, UA
License: MIT
Overview
UiPad is a dataset created for the IASA Champ 2024 Challenge, focusing on the accessibility and interface understanding of MacOS applications. With growing interest in AI-driven user interface… See the full description on the dataset page: https://huggingface.co/datasets/macpaw-research/UiPad.UIT-ViHSD
UIT-ViHSD Dataset
This is a copy instance of the original dataset provided by UIT. Please visit https://nlp.uit.edu.vn/datasets to obtain a usage permission before using this dataset.
UIT-VSFC
UIT-VSFC Dataset
This is a copy instance of dataset provided in the below paper. If you use this dataset, please cite the original work.
Paper: Kiet Van Nguyen, Vu Duc Nguyen, Phu Xuan-Vinh Nguyen, Tham Thi-Hong Truong, Ngan Luu-Thuy Nguyen, UIT-VSFC: Vietnamese Students' Feedback Corpus for Sentiment Analysis, 2018 10th International Conference on Knowledge and Systems Engineering (KSE 2018), November 1-3, 2018, Ho Chi Minh City, Vietnam.
vigoemotions
ViGoEmotions: A Benchmark Dataset For Fine-grained Emotion Detection on Vietnamese Texts
Datasize: 20,664 comments from social network sites.
Label: 27 categories of emotion, each comment can have multiple emotions.
Label mapping:
Index
Emotion
0
amusement
1
excitement
2
joy
3
love
4
desire
5
optimism
6
caring
7
pride
8
admiration
9
gratitude
10relief
11
approval
12
realization
13
surprise
14
curiosity
15
confusion
16
fear
17
nervousness… See the full description on the dataset page: https://huggingface.co/datasets/uitnlp/vigoemotions.ldt-latents-uint8UIS-QA
UIS-QA: A Benchmark for Unindexed Information Seeking
Figure 1. UIS problem. Standard agents (bottom) rely on indexed information and often fail or hallucinate; UIS-capable agents (top) use additional tools to excavate unindexed information and solve UIS tasks.
If .figs do not load, see the paper.
🔔 News
[2026.03.10] 🎉 We release the UIS-QA dataset and the paper (ICLR 2026, arXiv) today!
📋 Dataset Description
Homepage
Paper… See the full description on the dataset page: https://huggingface.co/datasets/UIS-Digger/UIS-QA.UI-Simulator_web_dataJuICE
JuICE
Sources
Repository: https://anonymous.4open.science/r/JuICE
HuggingFace: juice-cultural-eval/JuiCE
About
We present JuICE (Benchmark for LLM-Judge in Identifying Cultural Errors), a multilingual dataset of 7,470 span-level annotations of cultural and linguistic errors, collected from native speakers in long-form LLM responses. It covers 1,050 query-response pairs from four countries (the United States, South Korea, Indonesia, and Bangladesh)… See the full description on the dataset page: https://huggingface.co/datasets/uilab/JuICE.UIT-VSMEC
Dataset Card for UIT‑VSMEC
1. Dataset Summary
UIT‑VSMEC (Vietnamese Social Media Emotion Corpus) is a benchmark corpus for emotion recognition in Vietnamese social media comments. It consists of 6,927 human‐annotated sentences, each labeled with one of six basic emotions, plus an Other category.
2. Supported Tasks and Leaderboard
Primary Task: Text classification – emotion recognition
Metrics: Accuracy, F1‑score
No public leaderboard yet; contributions… See the full description on the dataset page: https://huggingface.co/datasets/visolex/UIT-VSMEC.UIT-VSMEC
Model description
This data from UIT aka University of Information Technology
It contain 7 class 'Other', 'Disgust', 'Enjoyment', 'Anger', 'Surprise', 'Sadness', 'Fear'
Contributions
Thanks to ViDataset - Vietnamese Datasets for Natural Language Processing for sharing this dataset.
UI-Simulator_android_datauit-sentiment-dataset-reddit-2000-balancedUINAUILUIT-VSMECumit_txtclass_dsimdb_th
รีวิว sentimental imdb ภาษาไทย
ตั้งต้นจาก https://huggingface.co/datasets/stanfordnlp/imdb
label
0 neg
1 pos
train.csv
test.csv
ตัวอย่างการใช้งาน
from datasets import load_dataset
# Specify the data files
data_files = {
"test": "test.csv",
"train": "train.csv"
}
dataset = load_dataset("uisp/ag_news_th", data_files=data_files)
print("Keys in loaded dataset:", dataset.keys()) # Should show keys for splits, like {'test', 'train'}
# Convert a… See the full description on the dataset page: https://huggingface.co/datasets/uisp/imdb_th.HTML-CSS-UIwikipedia-tr-llm-finetuneUIT-VSMEC
UIT-VSMEC Dataset
This is a copy instance of dataset provided in below paper. If you use this dataset, please cite the original work.
Vong Ho, Duong Nguyen, Danh Nguyen, Linh Pham, Kiet Nguyen and Ngan Nguyen, Emotion Recognition for Vietnamese Social Media Text, 2019 16th International Conference of the Pacific Association for Computational Linguistics (PACLING 2019), October 11-13, 2019, Ha Noi, Vietnam.
tweet_review_uidvi-uit-vsfc-classificationref: https://huggingface.co/datasets/SEACrowd/uit_vsfc
yelp_review_uidag_news_th
หมวด ข่าว ภาษาไทย
ตั้งต้นจาก https://huggingface.co/datasets/fancyzhx/ag_news
label
0 World
1 Sport
2 Business
3 Sci/Tech
ยังแปลไม่ครบ แต่มีแปลแล้วประมาณหนึ่ง
train.csv: แปลแล้ว 17,271 records จาก 120,000 records
test.csv: แปลทั้งหมดแล้ว 7,600 records
ตัวอย่างการใช้งาน
from datasets import load_dataset
# Specify the data files
data_files = {
"test": "test.csv",
"train": "train.csv"
}
dataset = load_dataset("uisp/ag_news_th"… See the full description on the dataset page: https://huggingface.co/datasets/uisp/ag_news_th.imdb_review_uiddialog_uid_gpt2embeddings_getstart
