datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TextSegmentationtajik-text-segmentationThis dataset contains texts in Tajik language with sentence annotations. It can be used to train and evaluate sentence-wise text segmentation algorithms.
The dataset contains more than 100 short and long texts and more than 3000 annotated sentences. The texts were carefully selected from different catergories
such as news, articles, novels, classical texts, poetry, and religious texts. It deliberately contains more of "hard" passages where splitting them by period "." characters would result… See the full description on the dataset page: https://huggingface.co/datasets/sobir-hf/tajik-text-segmentation.myanmar-text-segmentation-datasetPlease visit to the GitHub repository for other Myanmar Langauge datasets.
Myanmar Text Segmentation Dataset
A token classification dataset for Myanmar (Burmese) chunk segmentation, formatted for sequence labeling tasks using the BIO tagging scheme.
📓 Dataset Creation Notebook: myanmar-text-segmentation-dataset.ipynb
📓 Fine-Tuning Notebook: myanmar-text-segmentation-fine-tuning.ipynb (based on the HuggingFace Token Classification Guide)
🚀 Try it out: Myanmar Text Segmentation… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-text-segmentation-dataset.text-conditioned-sar-optical-segmentationtext-segmentation
