Team Ai
Datasetpublic

sello-ralethe/SA-KnowledgeBases

SA-KnowledgeBases ConceptNet and DBpedia knowledge projected into isiZulu, isiXhosa, Sesotho and Sepedi using LeNS-Align. Produced for the doctoral thesis Injecting Commonsense Knowledge into Pretrained Language Models for Low Resource Languages (University of Cape Town, 2026). Code at https://github.com/sello-ralethe/SA-knowledge Structure Six configurations. conceptnet and dbpedia hold the projected triples; conceptnet_verbalized and dbpedia_verbalized hold the… See the full description on the dataset page: https://huggingface.co/datasets/sello-ralethe/SA-KnowledgeBases.

sourceHugging Facecc-by-sa-4.0updated 1mo agoView on Hugging Face
0likes99downloads
Dataset Card

SA-KnowledgeBases

ConceptNet and DBpedia knowledge projected into isiZulu, isiXhosa, Sesotho and Sepedi using LeNS-Align.

Produced for the doctoral thesis Injecting Commonsense Knowledge into Pretrained Language Models for Low Resource Languages (University of Cape Town, 2026). Code at https://github.com/sello-ralethe/SA-knowledge

Structure

Six configurations. conceptnet and dbpedia hold the projected triples; conceptnet_verbalized and dbpedia_verbalized hold the same knowledge rendered as sentences, which is the form consumed by the text-conditioned model in Chapter 5; conceptnet_validation and dbpedia_validation hold the human judgements. Each configuration has one split per language.

Triple files use the columns subject, predicate, object. The source files were inconsistent in both column naming and column order, and have been normalised so that a single schema covers every language and both knowledge bases.

Loading

python
from datasets import load_dataset

triples = load_dataset("sello-ralethe/SA-KnowledgeBases", "conceptnet",
                       split="isizulu")
facts   = load_dataset("sello-ralethe/SA-KnowledgeBases",
                       "conceptnet_verbalized", split="isizulu")
judged  = load_dataset("sello-ralethe/SA-KnowledgeBases",
                       "conceptnet_validation", split="isizulu")

The validation configurations are the human-verified material. Everything else in this dataset was produced automatically.

Human validation

A sample of the projected triples was judged by first-language speakers, and these are the judgements behind the accuracy figures reported in Section 4.7.2.

Each row carries the English triple, its projection into the target language, the verdict in is_factual, and where the verdict is negative an error_type recording which part of the triple went wrong. The categories distinguish a wrong subject, a wrong object and a wrong translation, which is the categorization used in the error analysis in Section 4.7.2.

The dominant failure mode differs by source. Projection of DBpedia triples fails most often on the subject, which is typically a named entity and therefore dependent on the entity alignment component. ConceptNet projection fails more often on the object and on translation quality, reflecting its everyday vocabulary and the greater ambiguity of common nouns.

Annotators saw the English triple and its projection and decided whether the projected triple was factually correct in the target language, recording a reason where it was not.

Coverage

Projection operates on entity and relation labels, so graph topology is preserved and the projected graphs remain alignable to their English originals. Coverage is therefore the coverage of the source knowledge bases. Concepts without English equivalents are structurally absent, and no improvement in projection accuracy recovers them.

Licence

CC BY-SA 4.0, inherited from ConceptNet 5 and DBpedia. The share-alike obligation propagates to derivative work, including to the verbalized files and to the validation records, which reproduce the triples they judge.