Team Ai
Modelpublic

KM4STfulltext/SSCI-SciBERT-e4

sourceHugging Faceapache-2.0updated 4y agoView on Hugging Face
5likes43downloads
Model Card

SSCI-BERT: A pretrained language model for social scientific text

Introduction

The research for social science texts needs the support natural language processing tools.

The pre-trained language model has greatly improved the accuracy of text mining in general texts. At present, there is an urgent need for a pre-trained language model specifically for the automatic processing of scientific texts in social science.

We used the abstract of social science research as the training set. Based on the deep language model framework of BERT, we constructed SSCI-BERT and SSCI-SciBERT pre-training language models by transformers/run_mlm.py.

We designed four downstream tasks of Text Classification on different social scientific article corpus to verify the performance of the model.

  • —SSCI-BERT and SSCI-SciBERT are trained on the abstract of articles published in SSCI journals from 1986 to 2021. The training set involved in the experiment included a total of 503910614 words.
  • —Based on the idea of Domain-Adaptive Pretraining, SSCI-BERT and SSCI-SciBERT combine a large amount of abstracts of scientific articles based on the BERT structure, and continue to train the BERT and SSCI-SciBERT models respectively to obtain pre-training models for the automatic processing of Social science research texts.

News

  • —2022-03-24 : SSCIBERT and SSCI-SciBERT has been put forward for the first time.

How to use

Huggingface Transformers

The from_pretrained method based on Huggingface Transformers can directly obtain SSCI-BERT and SSCI-SciBERT models online.

  • —SSCI-BERT
python
from transformers import AutoTokenizer, AutoModel

tokenizer = AutoTokenizer.from_pretrained("KM4STfulltext/SSCI-BERT-e2")

model = AutoModel.from_pretrained("KM4STfulltext/SSCI-BERT-e2")
  • —SSCI-SciBERT
python
from transformers import AutoTokenizer, AutoModel

tokenizer = AutoTokenizer.from_pretrained("KM4STfulltext/SSCI-SciBERT-e2")

model = AutoModel.from_pretrained("KM4STfulltext/SSCI-SciBERT-e2")

Download Models

  • —The version of the model we provide is PyTorch.

From Huggingface

From Google Drive

We have put the model on Google Drive for users.

ModelDATASET(year)Base Model
SSCI-BERT-e21986-2021Bert-base-cased
SSCI-SciBERT-e2 (recommended)1986-2021Scibert-scivocab-cased
SSCI-BERT-e41986-2021Bert-base-cased
SSCI-SciBERT-e41986-2021Scibert-scivocab-cased

Evaluation & Results

  • —We use SSCI-BERT and SSCI-SciBERT to perform Text Classificationon different social science research corpus. The experimental results are as follows. Relevant data sets are available for download in the Verification task datasets folder of this project.
JCR Title Classify Dataset
Modelaccuracymacro avgweighted avg
Bert-base-cased28.4322.0621.86
Scibert-scivocab-cased38.4833.8933.92
SSCI-BERT-e240.4335.3735.33
SSCI-SciBERT-e241.3537.2737.25
SSCI-BERT-e440.6535.4935.40
SSCI-SciBERT-e441.1336.9636.94
Support230023002300
JCR Abstract Classify Dataset
Modelaccuracymacro avgweighted avg
Bert-base-cased48.5942.842.82
Scibert-scivocab-cased55.5951.451.81
SSCI-BERT-e258.0553.3153.73
SSCI-SciBERT-e259.9556.5157.12
SSCI-BERT-e459.0054.9755.59
SSCI-SciBERT-e460.0056.3856.90
Support220022002200
JCR Mixed Titles and Abstracts Dataset
**Model****accuracy****macro avg****weighted avg**
Bert-base-cased58.2457.2757.25
Scibert-scivocab-cased59.5858.6558.68
SSCI-BERT-e260.8960.2460.30
SSCI-SciBERT-e260.9660.5460.51
SSCI-BERT-e461.0060.4860.43
SSCI-SciBERT-e461.2460.7160.75
Support450045004500
SSCI Abstract Structural Function Recognition (Classify Dataset)
Bert-base-casedSSCI-BERT-e2SSCI-BERT-e4support
B63.7764.2964.63224
P53.6657.1457.9995
M87.6388.4389.06323
R86.8188.2888.47419
C78.3279.8278.95316
accuracy79.5980.980.971377
macro avg74.0475.5975.821377
weighted avg79.0280.3280.441377
Scibert-scivocab-casedSSCI-SciBERT-e2SSCI-SciBERT-e4support
B69.9870.9570.95224
P58.8960.1258.9695
M89.3790.1288.11323
R87.6688.0787.44419
C80.782.6182.94316
accuracy81.6382.7282.061377
macro avg77.3278.3777.681377
weighted avg81.682.5881.921377

Cited

  • —If our content is helpful for your research work, please quote our research in your article.
  • —https://link.springer.com/article/10.1007/s11192-022-04602-4

Disclaimer

  • —The experimental results presented in the report only show the performance under a specific data set and hyperparameter combination, and cannot represent the essence of each model. The experimental results may change due to random number seeds and computing equipment.
  • —Users can use the model arbitrarily within the scope of the license, but we are not responsible for the direct or indirect losses caused by using the content of the project.

Acknowledgment

  • —SSCI-BERT was trained based on BERT-Base-Cased](https://github.com/google-research/bert)).
  • —SSCI-SciBERT was trained based on scibert-scivocab-cased](https://github.com/allenai/scibert))