bubbles01/legal-document-summarization
# Legal Document Summarization Dataset This dataset was prepared for a Legal Document Summarization project using Indian legal judgments. Dataset The dataset contains training and test records created from legal documents. Each record consists of a legal text chunk and its corresponding aligned summary. Split Records Train 1,287 Test 675 Total 1,962 Fields doc_id — Document identifier chunk_id — Chunk identifier within the document section —… See the full description on the dataset page: https://huggingface.co/datasets/bubbles01/legal-document-summarization.
Legal Document Summarization Dataset
This dataset was prepared for a Legal Document Summarization project using Indian legal judgments.
Dataset
The dataset contains training and test records created from legal documents. Each record consists of a legal text chunk and its corresponding aligned summary.
Fields
- doc_id — Document identifier
- chunk_id — Chunk identifier within the document
- section — Section associated with the chunk
- chunk_text — Input legal text
- ligned_target — Target summary for the chunk
- split — Dataset split ( rain or est)
- para_range — Paragraph range covered by the chunk
- matched_sentences — Summary sentences aligned to the chunk
- lignment_scores — Cross-Encoder alignment scores
- matched_labels — Labels associated with the aligned summary sentences
- structure_fallback — Whether a structural fallback was used
- chunk_index — Index of the chunk
- chunktokens_est — Estimated number of input tokens
- targettokens_est — Estimated number of target tokens
Data Preparation
The legal documents were processed using section detection and chunking. Summary sentences were then aligned with relevant document chunks using a Cross-Encoder.
The resulting pairs are used as supervised data for training a sequence-to-sequence summarization model such as BART.
Task
The main task is legal document summarization:
Legal document chunk → Concise aligned summary
Splits
The training and test documents are kept separate during dataset construction. The dataset pipeline also performs split-isolation checks to prevent cross-split duplicate text.
Intended Use
This dataset is intended for academic experimentation and evaluation of legal document summarization models.
Repository Contents
- rain.jsonl — Training records
- est.jsonl — Test records
Acknowledgement
The source legal documents and reference summaries are used for academic NLP experimentation.
