datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Intel-WebCorpus-forms
💻 Intel WebCorpus Forms (Enterprise Hardware Q&A)
This dataset is a massive, high-fidelity archive of 176,472 technical troubleshooting discussions (containing nearly 1 million individual messages) scraped from the official Intel Community Forums.
It has been meticulously engineered for Large Language Model (LLM) training. Instead of a raw, messy dump of isolated posts, the data has been reconstructed into chronological conversation threads, noise-filtered, deduplicated, and… See the full description on the dataset page: https://huggingface.co/datasets/sphita/Intel-WebCorpus-forms.legal-templates-multilingual
Forms Legal — Multilingual Legal Templates Corpus
19,419 legal document templates across 36 jurisdictions in 12 languages, expanded to 22,876 (document × locale) rows. Released under CC-BY-4.0 by forms-legal.com.
Quick description
A multilingual corpus of structured legal document templates spanning 36 jurisdictions. Each document includes a multi-section editorial brief (whatIs, whenNeeded, keyElements, howToFill, legalRequirements, commonMistakes), 5-8… See the full description on the dataset page: https://huggingface.co/datasets/forms-legal/legal-templates-multilingual.Sanskrit-verb-forms
Sanskrit Verb Forms Dataset (संस्कृत धातु रूप संग्रह)
Overview
description: |
A comprehensive dataset containing Sanskrit verb conjugations (dhatu roop) with 10,348 unique entries.
Each entry provides the complete information about a Sanskrit verb form, including:
धातु (Dhatu): The root verb
पद (Pada): Voice of the verb (परस्मैपद/आत्मनेपद)
लकार (Lakara): Tense/mood of the verb
पुरुष (Purusha): Person (प्रथम/मध्यम/उत्तम)
वचन (Vachana): Number (एकवचन/द्विवचन/बहुवचन)… See the full description on the dataset page: https://huggingface.co/datasets/Process-Venue/Sanskrit-verb-forms.
