Team Ai
Datasetpublic

turkish-nlp-suite/Treebank-Benchmarking

Turkish Treebank Benchmarking This is the repo for Turkish treebank benchmarking, namely evaluating Tranformer models on POS-Dep-Morph task. For the data, we used two treebank, IMST and BOUN. We converted conllu format to json lines for being compatible to HF dataset formats. Here are treebank sizes at a glance: Dataset train lines dev lines test lines BOUN 7803 979 979 IMST 3435 1100 1100 A typical instance from the dataset looks like: { "id": "ins_1267"… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/Treebank-Benchmarking.

sourceHugging Facecc-by-sa-4.0updated 8mo agoView on Hugging Face
0likes34downloads
Dataset Card

<img src="https://raw.githubusercontent.com/turkish-nlp-suite/.github/main/profile/TreeBench.png" width="30%" height="30%">

Turkish Treebank Benchmarking

This is the repo for Turkish treebank benchmarking, namely evaluating Tranformer models on POS-Dep-Morph task. For the data, we used two treebank, IMST and BOUN. We converted conllu format to json lines for being compatible to HF dataset formats.

Here are treebank sizes at a glance:

Datasettrain linesdev linestest lines
BOUN7803979979
IMST343511001100

A typical instance from the dataset looks like:

{
  "id": "ins_1267",
  "tokens": [
    "Rüzgâr",
    "yine",
    "güçlü",
    "esiyor",
    "du",
    "."
  ],
  "upos": [
    "NOUN",
    "ADV",
    "ADV",
    "VERB",
    "AUX",
    "PUNCT"
  ],
  "heads": [
    4,
    4,
    4,
    0,
    4,
    4
  ],
  "rels": [
    "nsubj",
    "advmod",
    "advmod",
    "root",
    "cop",
    "punct"
  ],
  "feats": [
    "Case=Nom|Number=Sing|Person=3",
    "_",
    "_",
    "Aspect=Imp|Polarity=Pos|VerbForm=Part",
    "Aspect=Perf|Evident=Fh|Number=Sing|Person=3|Tense=Past",
    "_"
  ],
  "text": "Rüzgâr yine güçlü esiyor du .",
  "feats_dict_json": [
    "{\"Case\":\"Nom\",\"Number\":\"Sing\",\"Person\":\"3\"}",
    "{}",
    "{}",
    "{\"Aspect\":\"Imp\",\"Polarity\":\"Pos\",\"VerbForm\":\"Part\"}",
    "{\"Aspect\":\"Perf\",\"Evident\":\"Fh\",\"Number\":\"Sing\",\"Person\":\"3\",\"Tense\":\"Past\"}",
    "{}"
  ]
}

Benchmarking

Benchmarking is done by scripts on accompanying Github repo. Please proceed to this repo for running the experiments. Here are the benchmarking results for BERTurk with our scripts:

MetricBOUNIMST
pos_acc0.92630.9377
uas0.81510.7680
las0.74590.6960
morphAbbracc0.46570.6705
morphAspectacc0.11410.1152
morphCaseacc0.11960.0586
morphEchoacc0.42610.4875
morphEvidentacc0.30720.3953
morphMoodacc0.06540.0651
morphNumTypeacc0.26940.2991
morphNumberacc0.39860.4782
morphNumber[psor]acc0.43480.2333
morphPersonacc0.40210.4726
morphPerson[psor]acc0.24900.0671
morphPolarityacc0.33500.1674
morphPronTypeacc0.15350.2680
morphReflexacc0.56200.7051
morphTenseacc0.21490.1241
morphTypoacc0.5081—
morphVerbFormacc0.49120.2364
morphVoiceacc0.02010.2602
morphPoliteacc—0.1436
morphmicroacc0.30760.2915

Notes:

  • —— means that metric wasn’t present in that dataset’s reported results (e.g., morph_Typo_acc only in BOUN; morph_Polite_acc only in IMST).

Acknowledgments

This research was supported with Cloud TPUs from Google's TPU Research Cloud (TRC), like most of our projects. Many thanks to TRC team once again.