tomaarsen/span-marker-mbert-base-multinerd
SpanMarker for Multilingual Named Entity Recognition
This is a SpanMarker model that can be used for multilingual Named Entity Recognition trained on the MultiNERD dataset. In particular, this SpanMarker model uses bert-base-multilingual-cased as the underlying encoder. See train.py for the training script.
Is your data not (always) capitalized correctly? Then consider using this uncased variant of this model by @lxyuan for better performance: lxyuan/span-marker-bert-base-multilingual-uncased-multinerd.
Metrics
Label set
Usage
To use this model for inference, first install the span_marker library:
pip install span_markerYou can then run inference with this model like so:
from span_marker import SpanMarkerModel
# Download from the 🤗 Hub
model = SpanMarkerModel.from_pretrained("tomaarsen/span-marker-mbert-base-multinerd")
# Run inference
entities = model.predict("Amelia Earhart flew her single engine Lockheed Vega 5B across the Atlantic to Paris.")See the SpanMarker repository for documentation and additional information on this library.
Training procedure
Training hyperparameters
The following hyperparameters were used during training:
- learning_rate: 5e-05
- trainbatchsize: 32
- evalbatchsize: 32
- seed: 42
- optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
- lrschedulertype: linear
- lrschedulerwarmup_ratio: 0.1
- num_epochs: 1
Training results
Framework versions
- SpanMarker 1.2.4
- Transformers 4.28.1
- Pytorch 1.13.1+cu117
- Datasets 2.12.0
- Tokenizers 0.13.2
See also
- lxyuan/span-marker-bert-base-multilingual-cased-multinerd is similar to this model, but trained on 3 epochs instead of 2. It reaches better performance on 7 out of the 10 languages.
- lxyuan/span-marker-bert-base-multilingual-uncased-multinerd is a strong uncased variant of this model, also trained on 3 epochs instead of 2.
Contributions
Many thanks to Simone Tedeschi from Babelscape for his insight when training this model and his involvement in the creation of the training dataset.
