xml
Datasets
All datasets matching “xml”Dutch-Basisbestandwetten-Legislation-Laws-XML-Cleanstackexchange_xmlThis is a dump of the files from
https://archive.org/details/stackexchange
downloaded via torrent on 2021-07-01.
Publication date 2021-06-07 Usage Attribution-ShareAlike 4.0 International Creative Commons License by sa Topics Stack Exchange Data Dump Contributor Stack Exchange Community
Please see the license information at:
https://archive.org/details/stackexchange
The dataset has been split into following for cleaner formatting.… See the full description on the dataset page: https://huggingface.co/datasets/flax-sentence-embeddings/stackexchange_xml.pmc_open_access_xml
Dataset Card for PMC Open Access XML
Dataset Summary
The XML Open Access includes more than 3.4 million journal articles and preprints that are made available under
license terms that allow reuse.
Not all articles in PMC are available for text mining and other reuse, many have copyright protection, however articles
in the PMC Open Access Subset are made available under Creative Commons or similar licenses that generally allow more
liberal redistribution and reuse than a… See the full description on the dataset page: https://huggingface.co/datasets/TomTBT/pmc_open_access_xml.WM-ORG-XML-DUMP_FR_2026.08
Jeux De Données : Dump WikiMedia Français Août 2026 : Extraction et Nettoyage Complet
Description
Ce JDD contient des articles extraits du dump complet de Wikimedia d'août 2026, nettoyés et structurés pour l'entraînement de modèles d'apprentissage automatique.
Source originale : https://www.wikimedia.org/
Licence : https://creativecommons.org/licenses/by-sa/4.0/deed.en
Dump source : https://dumps.wikimedia.org/frwiki/latest/
Fichiers traités (4 fichiers «… See the full description on the dataset page: https://huggingface.co/datasets/MisterAI/WM-ORG-XML-DUMP_FR_2026.08.patent-spec-xmlCowpea-Architecture-XML
Cowpea-Architecture-XML-WDS
This dataset contains simulated images of Cowpea plants paired with organ-level architecture representations in XML format, packaged in WebDataset (.tar) format for efficient high-performance training.
Dataset Structure
The dataset is sharded into .tar files, each containing up to 10,000 samples.
Each sample consists of:
.jpeg: The plant image
.xml: The organ-level architecture representation
.json: (Optional) Metadata
Usage with… See the full description on the dataset page: https://huggingface.co/datasets/heesup/Cowpea-Architecture-XML.
