Team Ai
20 results

mbert

morten-j /medhie-tokenized-dataset-mBERT10M<n<100M0 likes177 downloads2y agoHugging Facecrystina-z /mbert-mrtydi-corpustext10M<n<100M0 likes56 downloads5y agoHugging FaceInabia-AI /mBERT-large-claim-agent-v10 mBERT-large Claim Agent — Training Dataset v10 Sentence-level binary classification data used to fine-tune mBERT-large for claim detection in medical-aesthetics promotional material. A claim is a statement of product efficacy, safety, indication, or market performance that requires substantiation against an approved claims matrix. Schema column type description id int Unique row id, 0..4717 sentence str The extracted sentence label int 1 = claim, 0… See the full description on the dataset page: https://huggingface.co/datasets/Inabia-AI/mBERT-large-claim-agent-v10.tabulartext-classification1K<n<10K0 likes54 downloads18d agoHugging FaceKashif786 /sindhi-gold-corpus-mlm-tokenized-mbert1M<n<10M0 likes45 downloads1mo agoHugging Facecrystina-z /mbert-mrtyditext10K<n<100K0 likes41 downloads5y agoHugging FaceMayaGalvez /linguistic_representation_mBERTThis dataset obtains genealogical and typological information for the 104 languages used for pre-training of the language model multilingual BERT (Devlin et al., 2019). The genealogical information covers the language family and the genus for each language. For typological description of the pre-training languages, 36 features from WALS (Dryer & Haspelmath, 2013) were used. The information provided here can be used, among other things, to investigate how the pre-training corpus is structured… See the full description on the dataset page: https://huggingface.co/datasets/MayaGalvez/linguistic_representation_mBERT.documentn<1K0 likes39 downloads4y agoHugging Face