sanyab/SEO-Intelligence-Tool
0
๐ SEO Content Intelligence Tool
An AI-powered web application that analyzes web articles and blog content to generate SEO insights โ including keyword extraction, topic detection, and meta tag suggestions.
๐ Problem Statement
In the digital marketing and content publishing space, optimizing content for search engines is critical for visibility. However, identifying relevant keywords, extracting topics, and generating effective meta titles/descriptions is often time-consuming and manual.
Challenge: Build a lightweight, deployable SEO tool that can automate this process using modern NLP techniques.
โ Solution Overview
This tool enables users to extract SEO-relevant insights from:
- Raw text content
- Uploaded documents (
.txt,.pdf,.doc,.docx) - Blog/article URLs (via web scraping)
It performs the following:
- Keyword extraction using KeyBERT (transformer-based)
- Named Entity Recognition (NER) and noun chunk detection via spaCy
- SEO meta title and description generation
- Support for OCR-based scanned PDFs using Tesseract
- Unified interface built with Streamlit, deployed on Hugging Face Spaces
๐ ๏ธ How It Works
๐ Input Options
- Raw Text: Direct paste input
- File Upload: Accepts
.txt,.pdf,.doc,.docx - URL Input: Scrapes article content using BeautifulSoup
๐ง AI-Powered Analysis
๐ผ๏ธ Application Demo
๐ Live Demo: https://huggingface.co/spaces/your-username/seo-intelligence-tool
๐ Project Structure
.
โโโ src
| โโโ app.py # Streamlit UI and main controller
| โโโ nlp_utils.py # NLP functions for keyword, entity, and meta extraction
| โโโ file_utils.py # File/URL scraping and OCR handling
โโโ requirements.txt # Python package dependencies
โโโ .huggingface.yml # Config for installing system packages (OCR support)
โโโ README.md # Project documentationโ๏ธ Installation & Local Setup
# Clone the repository
git clone https://github.com/sanyab1801/seo-intelligence-tool.git
cd seo-intelligence-tool
# Install Python dependencies
pip install -r requirements.txt
# Optional: install OCR dependencies (for .pdf scans, .doc support)
sudo apt install tesseract-ocr poppler-utils
# Run the app
streamlit run app.py๐ On Windows, install Tesseract OCR and Poppler manually, and add to PATH.
๐ฆ Requirements
streamlit
spacy
keybert
sentence-transformers
PyPDF2
python-docx
textract
pytesseract
pdf2image
beautifulsoup4
requestsThe app uses the en_core_web_sm spaCy model. It will auto-download on first run.๐ Supported File Types
๐ง Future Improvements (Optional Enhancements)
- โ
Export results to
.csvor.json - โ Keyword density and readability scoring
- โ SEO similarity comparison with competitor URLs
- โ Language detection and support for multilingual input
๐จโ๐ป Author
Developed by \[Sanya Behera] GitHub: github.com/sanyab1801 Deployed on: Hugging Face Spaces
