datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LegalRAG
Kazakh Legal Text Chunks
Dataset Summary
Kazakh Legal Text Chunks is a processed corpus of official legal texts of the Republic of Kazakhstan, prepared for retrieval-augmented generation (RAG), legal information retrieval, and grounded legal question answering in the Kazakh language.
The dataset contains structure-preserving text chunks derived from publicly available legal and normative documents. It is intended for research and development in:
legal retrieval,
legal QA… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/LegalRAG.oi_docs_datasetoi_docs_synthetic_alpacahungary_history_alpacajozsef_attila_osszes
