sandeepsaga/Transcript_Summerizer
016
1---2license: apache-2.03language:4- en5tags:6- summarization7- pegasus8- nlp9pipeline_tag: summarization10base_model: google/pegasus-large11---12 13# Fine-Tuned Pegasus Transcript Summarizer14 15## Model Details16 17### Model Description18This model is a fine-tuned version of `google/pegasus-large` optimized for text summarization, specifically designed for generating concise and structured summaries from spoken transcripts. It captures key takeaways, removes filler content, and retains core semantic coherence.19 20- **Developed by:** Sandeep Ganga21- **Model type:** Transformer-based sequence-to-sequence summarization model22- **Language(s):** English23- **License:** Apache 2.024- **Finetuned from model:** google/pegasus-large25 26## Model Sources27 28- **Repository:** [https://huggingface.co/sandeepsaga/Transcript_Summerizer](https://huggingface.co/sandeepsaga/Transcript_Summerizer)29 30---31 32## Uses33 34### Direct Use35* Meeting and lecture transcript summarization36* Podcast and interview summaries37* Long-form unstructured text condensor38 39### Downstream Use40The model can be fine-tuned further for:41* Domain-specific summarization (e.g., medical, legal, educational transcripts)42* Integration into AI-powered note-taking tools43 44### Out-of-Scope Use45* Generating fictional or creative stories46* Processing highly degraded or unsegmented audio transcripts47* Producing legally binding or medical summaries without human review48 49---50 51## Bias, Risks, and Limitations52 53* **Dataset Bias:** The model may reflect domain biases present in the fine-tuning transcript dataset.54* **Loss of Context:** Highly complex or lengthy transcripts may experience minor context compression.55* **Accuracy:** Generated output should be validated for mission-critical applications to avoid hallucinations.56 57---58 59## How to Get Started with the Model60 61Use the code below to run inference with PyTorch:62 63```python64from transformers import PegasusForConditionalGeneration, PegasusTokenizer65 66model_id = "sandeepsaga/Transcript_Summerizer"67 68tokenizer = PegasusTokenizer.from_pretrained(model_id)69model = PegasusForConditionalGeneration.from_pretrained(model_id)70 71def summarize_text(text):72 inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=512, padding="longest")73 summary_ids = model.generate(**inputs, max_length=150, min_length=30, length_penalty=2.0)74 return tokenizer.decode(summary_ids[0], skip_special_tokens=True)75 76# Example usage77sample_transcript = "The team met to review quarterly progress. Overall performance met expectations..."78print(summarize_text(sample_transcript))