Team Ai
Datasetpublic

german-tokenizer-benchmark/germeval14

GermEval 2014: Tokenized Sentences This datasets hosts a sentence-tokenized version of the GermEval 2014 NER dataset. Creation The following script can be used to reproduce the creation of the dataset: import json from flair.datasets import NER_GERMAN_GERMEVAL corpus = NER_GERMAN_GERMEVAL() with open("./germeval14/train.jsonl", "wt") as f_out: for sentence in germeval_corpus.train: current_example = { "text":… See the full description on the dataset page: https://huggingface.co/datasets/german-tokenizer-benchmark/germeval14.

sourceHugging Facecc-by-4.0updated 11mo agoView on Hugging Face
0likes12downloads
Dataset Card

GermEval 2014: Tokenized Sentences

This datasets hosts a sentence-tokenized version of the GermEval 2014 NER dataset.

Creation

The following script can be used to reproduce the creation of the dataset:

python
import json

from flair.datasets import NER_GERMAN_GERMEVAL

corpus = NER_GERMAN_GERMEVAL()

with open("./germeval14/train.jsonl", "wt") as f_out:
    for sentence in germeval_corpus.train:
        current_example = {
            "text": sentence.to_tokenized_string()
        }
        f_out.write(json.dumps(current_example) + "\n")

The extracted dataset has 24,000 sentences.