Team Ai
Datasetpublic

Ashfaq2000/machine-learning-wikipedia-dataset

Machine Learning Wikipedia Dataset Dataset Description This dataset contains 50 curated Wikipedia articles related to the Machine Learning niche, including core concepts, key algorithms, conferences, and tools within the field. Each record includes the article title, source URL, a summary, full text content, Wikipedia categories, outbound references, and associated images. Dataset Structure Each line in the .jsonl file is a JSON object with the… See the full description on the dataset page: https://huggingface.co/datasets/Ashfaq2000/machine-learning-wikipedia-dataset.

sourceHugging Facecc-by-sa-4.0updated 21d agoView on Hugging Face
1likes47downloads
Dataset Card

Machine Learning Wikipedia Dataset

Dataset Description

This dataset contains 50 curated Wikipedia articles related to the Machine Learning niche, including core concepts, key algorithms, conferences, and tools within the field.

Each record includes the article title, source URL, a summary, full text content, Wikipedia categories, outbound references, and associated images.

Dataset Structure

Each line in the .jsonl file is a JSON object with the following fields:

FieldDescription
titleWikipedia article title
urlLink to the original Wikipedia article
summaryShort summary of the article
contentFull extracted text content
categoriesWikipedia categories the article belongs to
referencesExternal reference links cited in the article
imagesImage URLs found in the article
sub_nicheThe related sub-topic/seed article this entry was grouped under

Example record

json
{
  "title": "Adversarial machine learning",
  "url": "https://en.wikipedia.org/wiki/Adversarial_machine_learning",
  "summary": "Adversarial machine learning is the study of the attacks on machine learning algorithms...",
  "content": "...",
  "categories": ["..."],
  "references": ["..."],
  "images": [],
  "sub_niche": "Adversarial machine learning"
}

Source and Licensing

All content is sourced from Wikipedia and is distributed under the CC-BY-SA 4.0 license. Attribution to Wikipedia and its contributors is required for any reuse or redistribution, and derivative works must be shared under the same license.

Intended Use

This dataset is intended for:

  • —Fine-tuning or evaluating language models on machine learning topics
  • —Building educational tools (quizzes, flashcards, glossaries)
  • —Knowledge graph and RAG (retrieval-augmented generation) experiments
  • —NLP research and coursework

Limitations

  • —Article distribution across sub-topics is uneven (some sub-topics have more articles than others).
  • —Content reflects Wikipedia at the time of collection and may not include the latest edits.