Team Ai
Apppublic

KarinaAulia/Movie_Review_Summarizer

sourceHugging Facemitupdated 11mo agoView on Hugging Face
0likes
App README

🎬 Smart Text Summarizer: TextRank vs T5-small

<div align="center">

Python Gradio Transformers License

Domain-aware text summarization system comparing Extractive and Abstractive approaches

Try Demo β€’ Report Bug β€’ Documentation

</div>


🌟 Overview

This project implements and compares two fundamental text summarization approaches:

  • β€”TextRank (Extractive): Graph-based algorithm that extracts key sentences from the original text
  • β€”T5-small (Abstractive): Transformer-based model that generates new paraphrased summaries

Key Features

✨ Real-time ROUGE Scoring - Automatically calculates quality metrics for every input 🎯 Domain Detection - Intelligently identifies movie reviews vs general text πŸ“Š Three Comparison Modes - Multiple perspectives on summary quality πŸš€ Universal Application - Works on any text type, optimized for movie reviews ⚑ Fast Processing - Instant results with efficient algorithms 🎨 Interactive UI - Beautiful, user-friendly Gradio interface


🎯 Demo

Try different text types to see how both models perform:

  1. 1.Movie Reviews (Positive/Negative) - Primary domain
  2. 2.News Articles - General news content
  3. 3.Academic Papers - Technical/scientific text
  4. 4.Stories & Narratives - Creative writing
  5. 5.Your Own Text - Paste anything!

πŸ”¬ Methodology

TextRank (Extractive Summarization)

Original Text β†’ Sentence Tokenization β†’ Similarity Matrix β†’ 
PageRank Algorithm β†’ Ranked Sentences β†’ Extract Top-K β†’ Summary

Algorithm:

  • β€”Graph-based ranking using PageRank
  • β€”Sentence similarity via word overlap
  • β€”Preserves original phrasing
  • β€”Fast and efficient

Pros:

  • β€”βœ… Grammatically correct (uses original sentences)
  • β€”βœ… Factually accurate
  • β€”βœ… Fast processing
  • β€”βœ… No training required

Cons:

  • β€”βŒ Less fluent connections
  • β€”βŒ May miss context
  • β€”βŒ Limited paraphrasing

T5-small (Abstractive Summarization)

Original Text β†’ Tokenization β†’ T5 Encoder β†’ 
Transformer Layers β†’ T5 Decoder β†’ Generated Summary

Model:

  • β€”Pre-trained T5-small (60M parameters)
  • β€”Text-to-Text Transfer Transformer
  • β€”Beam search decoding (num_beams=4)
  • β€”Max length: 150 tokens

Pros:

  • β€”βœ… More fluent and natural
  • β€”βœ… Better paraphrasing
  • β€”βœ… Captures abstract concepts
  • β€”βœ… Flexible output

Cons:

  • β€”βŒ May introduce errors
  • β€”βŒ Slower processing
  • β€”βŒ Requires more resources

πŸ“Š Real-time ROUGE Calculation

This app calculates ROUGE scores in real-time for every input!

Three Comparison Modes:

1️⃣ Summaries vs Original Text

Measures content retention and overlap

TextRank vs Original β†’ Shows extractive quality
T5 vs Original β†’ Shows abstractive quality
2️⃣ Model Agreement (TextRank vs T5)

Shows consensus between approaches

High Agreement β†’ Both identify similar key points
Low Agreement β†’ Different focus/perspectives
3️⃣ Interpretation Guide
ROUGE ScoreQuality Level
0.5 - 1.0⭐⭐⭐⭐⭐ Excellent
0.3 - 0.5⭐⭐⭐⭐ Good
0.1 - 0.3⭐⭐⭐ Moderate
< 0.1⭐⭐ Needs improvement

πŸ“ˆ ROUGE Metrics Explained

  • β€”ROUGE-1: Unigram (single word) overlap
  • β€”ROUGE-2: Bigram (two consecutive words) overlap
  • β€”ROUGE-L: Longest Common Subsequence

πŸŽ“ Domain Detection

Smart algorithm that identifies text type:

Detection Logic

python
Strong Keywords: film, movie, director, actor, screenplay
Medium Keywords: plot, character, scene, performance
Weak Keywords: good, bad, great, entertaining

Score = (Strong Γ— 3) + (Medium Γ— 2) + (Weak Γ— 1)

Confidence Levels

BadgeConfidenceMeaning
🎬 Greenβ‰₯70%Movie review detected
⚠️ Orange40-70%Possible movie review
πŸ“ Blue<40%General text

πŸ“š Evaluation Results

Dataset Evaluation (Colab)

Evaluated on 100 movie reviews with ground truth summaries:

ModelROUGE-1ROUGE-2ROUGE-LAvg
TextRank0.42 Β± 0.050.18 Β± 0.040.35 Β± 0.050.32
T5-small0.51 Β± 0.040.26 Β± 0.030.44 Β± 0.040.40

Key Findings:

  • β€”βœ… T5 outperforms TextRank by 21.4% in ROUGE-1
  • β€”βœ… T5 shows 44.4% improvement in ROUGE-2 (better phrase capture)
  • β€”βœ… TextRank is ~10x faster than T5
  • β€”βœ… TextRank better for factual accuracy, T5 for fluency

Performance Metrics

MetricTextRankT5-small
Speed~0.1s~1.5s
Memory50MB250MB
AccuracyHighVery High
FluencyMediumHigh
ResourceLowMedium

πŸ› οΈ Technical Stack

Core Libraries

python
transformers==4.35.0      # T5 model
torch==2.1.0              # PyTorch backend
nltk==3.8.1               # Text processing
networkx==3.2.1           # Graph algorithms
rouge-score==0.1.2        # Evaluation metrics
gradio==4.7.1             # Web interface
scikit-learn==1.3.2       # ML utilities

System Requirements

  • β€”Python: 3.8+
  • β€”RAM: 2GB minimum (4GB recommended)
  • β€”GPU: Optional (CPU works fine)
  • β€”Storage: ~500MB for models

πŸš€ Usage Examples

Example 1: Movie Review

Input:

"This film is a masterpiece of modern cinema. The director's vision is crystal clear throughout..."

TextRank Output:

"This film is a masterpiece of modern cinema. The performances are outstanding. The soundtrack complements the narrative perfectly."

T5 Output:

"the film is a masterpiece with breathtaking cinematography and outstanding performances. while pacing may feel slow, it serves the story well."

ROUGE Scores:

  • β€”TextRank: R-1: 0.45, R-2: 0.22, R-L: 0.38
  • β€”T5: R-1: 0.52, R-2: 0.28, R-L: 0.45

Example 2: News Article

Input:

"The technology sector experienced significant volatility today..."

TextRank Output:

"The technology sector experienced significant volatility. Share prices fluctuated dramatically. Investors remain cautious but optimistic."

T5 Output:

"major companies announced quarterly earnings, causing significant volatility in the technology sector amid supply chain challenges."

🎯 Use Cases

Educational

  • β€”πŸ“š Teaching NLP concepts
  • β€”πŸŽ“ Comparing summarization approaches
  • β€”πŸ”¬ Research demonstrations

Practical

  • β€”πŸ“° News article summarization
  • β€”πŸŽ¬ Movie review digests
  • β€”πŸ“„ Document processing
  • β€”πŸ“§ Email summarization

Development

  • β€”πŸ§ͺ Testing summarization models
  • β€”πŸ“Š Benchmarking algorithms
  • β€”πŸ”„ Comparing extractive vs abstractive

πŸ” How It Works

Step-by-Step Process

  1. 1.Input Processing
   User Input β†’ Domain Detection β†’ Preprocessing
  1. 1.Parallel Summarization
   TextRank: Text β†’ Sentences β†’ Graph β†’ PageRank β†’ Summary
   T5: Text β†’ Tokenize β†’ Encode β†’ Generate β†’ Summary
  1. 1.Real-time Evaluation
   ROUGE Calculation β†’ Statistics β†’ Display Results
  1. 1.Interactive Display
   Summaries + Metrics + Visualizations β†’ User Interface

πŸ“Š Statistics Provided

For each summary, you get:

Summary Metrics

  • β€”Word count (original vs summary)
  • β€”Sentence count
  • β€”Character count
  • β€”Compression ratio (%)

ROUGE Scores

  • β€”ROUGE-1 (unigram overlap)
  • β€”ROUGE-2 (bigram overlap)
  • β€”ROUGE-L (longest common subsequence)

Model Info

  • β€”Processing method
  • β€”Algorithm type
  • β€”Speed/performance
  • β€”Quality indicators

🎨 Interface Features

Clean & Modern Design

  • β€”πŸŽ¨ Purple gradient theme
  • β€”πŸ“± Responsive layout
  • β€”πŸŒ“ Professional styling
  • β€”βœ¨ Smooth animations

User-Friendly

  • β€”πŸŽ― Clear instructions
  • β€”πŸ“‹ Sample texts included
  • β€”πŸ”„ One-click testing
  • β€”πŸ“Š Visual comparisons

Interactive Elements

  • β€”βœ… Real-time processing
  • β€”πŸ“ˆ Live statistics
  • β€”πŸŽ¬ Domain indicators
  • β€”πŸ—‘οΈ Easy reset

πŸ§ͺ Testing

Try these scenarios to test the system:

  1. 1.Long Reviews (200+ words) - Test compression
  2. 2.Short Reviews (50 words) - Test edge cases
  3. 3.Technical Text - Test domain detection
  4. 4.Mixed Content - Test robustness
  5. 5.Multiple Languages - Test limitations (English only)

πŸ“– Academic Context

Project Information

Course: Natural Language Processing Topic: Text Summarization Comparison Approach: Extractive vs Abstractive Evaluation: ROUGE Metrics

Learning Objectives

βœ… Understanding graph-based algorithms (TextRank) βœ… Working with transformer models (T5) βœ… Implementing evaluation metrics (ROUGE) βœ… Comparing summarization approaches βœ… Building interactive NLP applications

Research Questions

  1. 1.How do extractive and abstractive methods compare?
  2. 2.When is each approach more suitable?
  3. 3.What are the trade-offs in speed vs quality?
  4. 4.How does domain affect summarization quality?

🀝 Contributing

This is an educational project, but suggestions are welcome!

How to Contribute

  1. 1.Fork the repository
  2. 2.Create your feature branch
  3. 3.Test your changes
  4. 4.Submit a pull request

Areas for Improvement

  • β€”[ ] Add more domain detection categories
  • β€”[ ] Implement fine-tuned T5 for movie reviews
  • β€”[ ] Add more evaluation metrics
  • β€”[ ] Support for other languages
  • β€”[ ] Batch processing capability

πŸ“ Citation

If you use this project in your research or education, please cite:

bibtex
@software{smart_text_summarizer,
  title={Smart Text Summarizer: TextRank vs T5-small},
  author={Your Name},
  year={2024},
  url={https://huggingface.co/spaces/YOUR_USERNAME/YOUR_SPACE}
}

πŸ”— References

Papers

  • β€”TextRank: Mihalcea & Tarau (2004) - "TextRank: Bringing Order into Text"
  • β€”T5: Raffel et al. (2020) - "Exploring the Limits of Transfer Learning"
  • β€”ROUGE: Lin (2004) - "ROUGE: A Package for Automatic Evaluation of Summaries"

Resources


πŸ“œ License

This project is licensed under the MIT License - see the LICENSE file for details.

MIT License

Copyright (c) 2024

Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction...

πŸ‘€ Author

Your Name


πŸ™ Acknowledgments

  • β€”Hugging Face for hosting and transformers library
  • β€”Google for T5 pre-trained model
  • β€”NLTK community for text processing tools
  • β€”Gradio team for the amazing interface framework
  • β€”Course Instructor for guidance and support

πŸ“ž Support

Having issues or questions?

  • β€”πŸ“§ Email: your.email@example.com
  • β€”πŸ’¬ Issues: GitHub Issues
  • β€”πŸ“š Documentation: Wiki

πŸ—ΊοΈ Roadmap

Version 1.0 βœ…

  • β€”[x] TextRank implementation
  • β€”[x] T5 integration
  • β€”[x] Real-time ROUGE scoring
  • β€”[x] Domain detection
  • β€”[x] Gradio interface

Version 2.0 (Planned)

  • β€”[ ] Fine-tuned T5 for movie reviews
  • β€”[ ] Multi-language support
  • β€”[ ] Batch processing
  • β€”[ ] Advanced metrics (BERTScore)
  • β€”[ ] Export functionality

<div align="center">

⭐ Star this project if you find it helpful!

🎬 Made with ❀️ for NLP Education

⬆ Back to Top

</div>