Team Ai
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01morten-j /medhie-tokenized-dataset-mBERT10M<n<100M0 likes177 downloads2y agoHugging Face02crystina-z /mbert-mrtydi-corpustext10M<n<100M0 likes56 downloads5y agoHugging Face03Inabia-AI /mBERT-large-claim-agent-v10 mBERT-large Claim Agent — Training Dataset v10 Sentence-level binary classification data used to fine-tune mBERT-large for claim detection in medical-aesthetics promotional material. A claim is a statement of product efficacy, safety, indication, or market performance that requires substantiation against an approved claims matrix. Schema column type description id int Unique row id, 0..4717 sentence str The extracted sentence label int 1 = claim, 0… See the full description on the dataset page: https://huggingface.co/datasets/Inabia-AI/mBERT-large-claim-agent-v10.tabulartext-classification1K<n<10K0 likes54 downloads19d agoHugging Face04Kashif786 /sindhi-gold-corpus-mlm-tokenized-mbert1M<n<10M0 likes45 downloads1mo agoHugging Face05crystina-z /mbert-mrtyditext10K<n<100K0 likes41 downloads5y agoHugging Face06MayaGalvez /linguistic_representation_mBERTThis dataset obtains genealogical and typological information for the 104 languages used for pre-training of the language model multilingual BERT (Devlin et al., 2019). The genealogical information covers the language family and the genus for each language. For typological description of the pre-training languages, 36 features from WALS (Dryer & Haspelmath, 2013) were used. The information provided here can be used, among other things, to investigate how the pre-training corpus is structured… See the full description on the dataset page: https://huggingface.co/datasets/MayaGalvez/linguistic_representation_mBERT.documentn<1K0 likes39 downloads4y agoHugging Face07kobybar /tokenized_sefaria_dataset_mberttext1M<n<10M0 likes32 downloads2y agoHugging Face08tchubakov /kgz_dataset_chunked_medium_mbert1M<n<10M0 likes19 downloads9mo agoHugging Face09morten-j /medhie-tokenized-dataset-sefaria-mBERT1M<n<10M0 likes18 downloads2y agoHugging Face10morten-j /medhie-tokenized-dataset-roots_ar_b50-mBERT1M<n<10M0 likes16 downloads2y agoHugging Face11morten-j /medhie-tokenized-dataset-roots_ar-mBERT10M<n<100M0 likes16 downloads2y agoHugging Face12morten-j /medhie-tokenized-dataset-roots_ar_f50-mBERT1M<n<10M0 likes13 downloads2y agoHugging Face13romjansen /mbert-base-cased-NER-NL-legislation-refs-datagated Dataset description This dataset was created for fine-tuning the model mbert-base-cased-NER-NL-legislation-refs and consists of 512 token long examples which each contain one or more legislation references. These examples were created from a weakly labelled corpus of Dutch case law which was scraped from Linked Data Overheid, pre-tokenized and labelled (biluo_tags_from_offsets) through spaCy and further tokenized through applying Hugging Face's AutoTokenizer.from_pretrained() for… See the full description on the dataset page: https://huggingface.co/datasets/romjansen/mbert-base-cased-NER-NL-legislation-refs-data.texttoken-classification10K<n<100K0 likes10 downloads4y agoHugging Face14versokuro /mbert-circuit-outputsimagen<1K0 likes9 downloads5mo agoHugging Face15kobybar /tokenized_roots_dataset_mberttext1K<n<10K0 likes8 downloads2y agoHugging Face16fernandabufon /results_mbert_basetabularn<1K0 likes5 downloads2y agoHugging Face17ardauzunoglu /mbert_reward_dpoed_model_ckpt20_c4_lowq_200m2b_subsample20m_grpo_prompttabular100K<n<1M0 likes5 downloads4mo agoHugging Face18mbertrand7905 /social-dataset Social Audio Video Data Notes Dataset summary This repository contains a preparation pipeline and a small metadata sample for Social work with Audio Video inputs. It does not claim to be a complete benchmark release; the loader documents how source data is normalized and validated. Included material dataset.py — loading, cleaning, and split preparation code. dataset_infos.json — schema and split metadata. metadata_sample.jsonl — small… See the full description on the dataset page: https://huggingface.co/datasets/mbertrand7905/social-dataset.0 likes16h agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.