Team Ai
Datasetpublic

uitnlp/ViANLI

Dataset Card for “ViANLI” Dataset Summary ViANLI (Vietnamese Adversarial Natural Language Inference) is the first adversarial benchmark dataset for Vietnamese NLI, designed to evaluate model robustness against complex linguistic phenomena. The dataset was constructed using a human-and-machine-in-the-loop approach with multi-round adversarial generation and dual human–machine verification. ViANLI contains over 10,000 high-quality premise–hypothesis pairs across 13… See the full description on the dataset page: https://huggingface.co/datasets/uitnlp/ViANLI.

sourceHugging Facemitupdated 11mo agoView on Hugging Face
3likes280downloads
Dataset Card

Dataset Card for “ViANLI”

Dataset Summary

ViANLI (Vietnamese Adversarial Natural Language Inference) is the first adversarial benchmark dataset for Vietnamese NLI, designed to evaluate model robustness against complex linguistic phenomena. The dataset was constructed using a human-and-machine-in-the-loop approach with multi-round adversarial generation and dual human–machine verification. ViANLI contains over 10,000 high-quality premise–hypothesis pairs across 13 diverse domains from Vietnamese news articles.

Each pair is labeled as entailment, contradiction, or neutral, following the standard NLI framework. The dataset provides a challenging benchmark for evaluating reasoning robustness and supports research in both Vietnamese and multilingual NLI.


Languages

  • —Vietnamese (vi)

Dataset Structure

Data Instances

Each instance in ViANLI is a JSON line with the following fields:

FieldTypeDescription
uidstringUnique identifier for each instance
premisestringThe premise sentence extracted from Vietnamese news
hypothesisstringThe hypothesis sentence written by annotators
labelstringOne of three classes: entailment, neutral, contradiction

Data Splits

SplitSize
train8,012
validation1,000
test1,000

Dataset License

  • —License: CC BY-NC-SA 4.0 (Creative Commons Attribution–NonCommercial–ShareAlike 4.0 International License)

Citation Information

If you use this dataset, please cite:

bibtex
@article{HUYNH2025130109,
title = {A New Benchmark Dataset and Mixture-of-Experts Language Models for Adversarial Natural Language Inference in Vietnamese},
journal = {Expert Systems with Applications},
pages = {130109},
year = {2025},
issn = {0957-4174},
doi = {https://doi.org/10.1016/j.eswa.2025.130109},
url = {https://www.sciencedirect.com/science/article/pii/S095741742503725X},
author = {Tin Van Huynh, Kiet Van Nguyen and Ngan Luu-Thuy Nguyen},
}