Team Ai
Datasetpublic

jensjepsen/danish-wiki-closedqa-v1

Danish Wiki closed-QA v1 Factual question-and-answer pairs generated from Danish Wikipedia via google/gemma-3-12b-it. Each row is one question + short factual answer grounded in a specific Wikipedia article. Article selection Articles were scored for general-knowledge salience and filtered before generation. Signals combined: page_len — article byte length (from page.sql.gz) langlinks — number of other language editions of the article (from langlinks.sql.gz) —… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/danish-wiki-closedqa-v1.

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes35downloads
Dataset Card

Danish Wiki closed-QA v1

Factual question-and-answer pairs generated from Danish Wikipedia via google/gemma-3-12b-it. Each row is one question + short factual answer grounded in a specific Wikipedia article.

Article selection

Articles were scored for general-knowledge salience and filtered before generation. Signals combined:

  • —page_len — article byte length (from page.sql.gz)
  • —langlinks — number of other language editions of the article (from langlinks.sql.gz) — signals universal salience
  • —inbound wikilinks — number of other da.wiki articles that link INTO this one (from pagelinks.sql.gz joined with linktarget.sql.gz) — signals Danish-community relevance, filters out globally-popular-but- not-Danish-relevant topics
  • —danish_relevance ratio — inbound ≥ langlinks / 10, to demote entries with lots of foreign cross-wiki interest but little linking from other Danish articles
  • —stub-heavy category demotion — any article whose direct categories match sport-year / league-season / individual-athlete-by-nationality patterns is demoted regardless of the other signals
  • —title heuristics — year-only pages, "Liste over", parish stubs, Portal: pages excluded

Articles passing all filters were tiered T1universal / T2mainstream. Only T1+T2 (11,624 articles) were used as seed for this dataset.

Generation

Each article's ~1500-char intro was passed to Gemma with instructions to produce 12 distinct factual Q/A pairs grounded in the text, with all questions phrased to stand alone (no references to "the article", "the text", "the source" etc.).