Team Ai
Datasetpublic

bubbles01/legal-document-summarization

# Legal Document Summarization Dataset This dataset was prepared for a Legal Document Summarization project using Indian legal judgments. Dataset The dataset contains training and test records created from legal documents. Each record consists of a legal text chunk and its corresponding aligned summary. Split Records Train 1,287 Test 675 Total 1,962 Fields doc_id — Document identifier chunk_id — Chunk identifier within the document section —… See the full description on the dataset page: https://huggingface.co/datasets/bubbles01/legal-document-summarization.

sourceHugging Faceupdated 10d agoView on Hugging Face
1likes43downloads
Dataset Card

Legal Document Summarization Dataset

This dataset was prepared for a Legal Document Summarization project using Indian legal judgments.

Dataset

The dataset contains training and test records created from legal documents. Each record consists of a legal text chunk and its corresponding aligned summary.

SplitRecords
Train1,287
Test675
Total1,962

Fields

  • —doc_id — Document identifier
  • —chunk_id — Chunk identifier within the document
  • —section — Section associated with the chunk
  • —chunk_text — Input legal text
  • —ligned_target — Target summary for the chunk
  • —split — Dataset split ( rain or est)
  • —para_range — Paragraph range covered by the chunk
  • —matched_sentences — Summary sentences aligned to the chunk
  • —lignment_scores — Cross-Encoder alignment scores
  • —matched_labels — Labels associated with the aligned summary sentences
  • —structure_fallback — Whether a structural fallback was used
  • —chunk_index — Index of the chunk
  • — chunktokens_est — Estimated number of input tokens
  • — targettokens_est — Estimated number of target tokens

Data Preparation

The legal documents were processed using section detection and chunking. Summary sentences were then aligned with relevant document chunks using a Cross-Encoder.

The resulting pairs are used as supervised data for training a sequence-to-sequence summarization model such as BART.

Task

The main task is legal document summarization:

Legal document chunk → Concise aligned summary

Splits

The training and test documents are kept separate during dataset construction. The dataset pipeline also performs split-isolation checks to prevent cross-split duplicate text.

Intended Use

This dataset is intended for academic experimentation and evaluation of legal document summarization models.

Repository Contents

  • —rain.jsonl — Training records
  • —est.jsonl — Test records

Acknowledgement

The source legal documents and reference summaries are used for academic NLP experimentation.