Team Ai
20 results

multilingual

AmazonScience /MultilingualMultiModalClassification Additional Information To load the dataset, import datasets ds = datasets.load_dataset("AmazonScience/MultilingualMultiModalClassification", data_dir="wiki-doc-ar-merged") print(ds) DatasetDict({ train: Dataset({ features: ['image', 'filename', 'words', 'ocr_bboxes', 'label'], num_rows: 8129 }) validation: Dataset({ features: ['image', 'filename', 'words', 'ocr_bboxes', 'label'], num_rows: 1742 }) test: Dataset({ features:… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/MultilingualMultiModalClassification.2 likes61k downloads2y agoHugging FaceSWE-bench /SWE-bench_Multilingual SWE-bench Multilingual Dataset Summary SWE-bench Multilingual is a dataset that tests systems' ability to resolve real-world GitHub issues across a broad range of programming languages. The original SWE-bench is Python-only; this dataset extends the same task format to 9 languages drawn from 41 popular repositories. The dataset collects 300 test Issue-Pull Request pairs. Evaluation is performed by unit test verification, using post-PR behavior as the reference solution. The… See the full description on the dataset page: https://huggingface.co/datasets/SWE-bench/SWE-bench_Multilingual.textn<1K28 likes51k downloads2mo agoHugging Facefacebook /multilingual_librispeech Dataset Card for MultiLingual LibriSpeech Dataset Summary This is a streamable version of the Multilingual LibriSpeech (MLS) dataset. The data archives were restructured from the original ones from OpenSLR to make it easier to stream. MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish.… See the full description on the dataset page: https://huggingface.co/datasets/facebook/multilingual_librispeech.audioautomatic-speech-recognition1M<n<10M191 likes44k downloads2y agoHugging FaceMultilingual-NLP /M-ABSA M-ABSA This repo contains the data for our paper M-ABSA: A Multilingual Dataset for Aspect-Based Sentiment Analysis. Data Description: This is a dataset suitable for the multilingual ABSA task with triplet extraction. All datasets are stored in the data/ folder: All dataset contains 7 domains. domains = ["coursera", "hotel", "laptop", "restaurant", "phone", "sight", "food"] Each dataset contains 21 languages. langs = ["ar", "da", "de", "en", "es", "fr", "hi"… See the full description on the dataset page: https://huggingface.co/datasets/Multilingual-NLP/M-ABSA.texttoken-classification100K<n<1M10 likes29k downloads4mo agoHugging Facehysi-lab /unwebtv-multilingual-audio-archive UN WebTV Multilingual Audio Archive This public dataset is a reproducibility-oriented archive of multilingual audio tracks collected from publicly accessible institutional media pages. Files are grouped by source and published as source-level archives. The accompanying manifests preserve source URLs, language labels, extraction status, and verification metadata. The archive is intended for research and engineering evaluation; downstream users must respect the terms, licenses… See the full description on the dataset page: https://huggingface.co/datasets/hysi-lab/unwebtv-multilingual-audio-archive.0 likes23k downloads8m agoHugging Facenvidia /OCR-Synthetic-Multilingual-v1 OCR-Synthetic-Multilingual-v1 Dataset Description Large-scale synthetically generated OCR training dataset for multilingual text detection and recognition. The data was produced using a heavily modified and extended version of SynthDoG (Synthetic Document Generator), originally introduced in the Donut project by Kim et al. This dataset was used to train Nemotron OCR v2, a state-of-the-art multilingual OCR model that is part of the NVIDIA NeMo Retriever collection.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/OCR-Synthetic-Multilingual-v1.object-detection10M<n<100M53 likes14k downloads6mo agoHugging Face