Team Ai
Datasetpublic

oklenAI/UDM_cleaned_docs

UDM cleaned docs 6,029,052 web pages reduced to just their mathematical content, extracted verbatim by oklenAI/udm_doc_extract_qwen3.5_2B — a 2B model distilled from GPT-5.6. Every row is model output, not human-curated text. The extract field is what the model returned for that page; the source page text is not included. Read Two repetition flags below before filtering — the obvious flag is not the one you want. How it was built step pages… See the full description on the dataset page: https://huggingface.co/datasets/oklenAI/UDM_cleaned_docs.

sourceHugging Faceupdated 1mo agoView on Hugging Face
0likes898downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

Team Ai shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
oklenAI/UDM_cleaned_docs · Team Ai