Team Ai
Apppublic

sanyab/SEO-Intelligence-Tool

sourceHugging Facemitupdated 1y agoView on Hugging Face
0likes
App README

๐Ÿ“˜ SEO Content Intelligence Tool

An AI-powered web application that analyzes web articles and blog content to generate SEO insights โ€” including keyword extraction, topic detection, and meta tag suggestions.

๐Ÿ“Œ Problem Statement

In the digital marketing and content publishing space, optimizing content for search engines is critical for visibility. However, identifying relevant keywords, extracting topics, and generating effective meta titles/descriptions is often time-consuming and manual.

Challenge: Build a lightweight, deployable SEO tool that can automate this process using modern NLP techniques.

โœ… Solution Overview

This tool enables users to extract SEO-relevant insights from:

  • โ€”Raw text content
  • โ€”Uploaded documents (.txt, .pdf, .doc, .docx)
  • โ€”Blog/article URLs (via web scraping)

It performs the following:

  • โ€”Keyword extraction using KeyBERT (transformer-based)
  • โ€”Named Entity Recognition (NER) and noun chunk detection via spaCy
  • โ€”SEO meta title and description generation
  • โ€”Support for OCR-based scanned PDFs using Tesseract
  • โ€”Unified interface built with Streamlit, deployed on Hugging Face Spaces

๐Ÿ› ๏ธ How It Works

๐Ÿ‘‡ Input Options
  • โ€”Raw Text: Direct paste input
  • โ€”File Upload: Accepts .txt, .pdf, .doc, .docx
  • โ€”URL Input: Scrapes article content using BeautifulSoup
๐Ÿง  AI-Powered Analysis
FeatureDescription
Top KeywordsExtracted via KeyBERT using transformer embeddings
Named EntitiesExtracted via spaCy
Topics / Noun PhrasesBased on noun chunks from spaCyโ€™s parser
Meta Title & DescriptionHeuristically generated from top-ranked keywords and leading sentences

๐Ÿ–ผ๏ธ Application Demo

๐Ÿ”— Live Demo: https://huggingface.co/spaces/your-username/seo-intelligence-tool

Demo Example Video

๐Ÿ“ Project Structure

bash
.
โ”œโ”€โ”€ src
|     โ”œโ”€โ”€ app.py           # Streamlit UI and main controller
|     โ”œโ”€โ”€ nlp_utils.py     # NLP functions for keyword, entity, and meta extraction
|     โ”œโ”€โ”€ file_utils.py    # File/URL scraping and OCR handling
โ”œโ”€โ”€ requirements.txt       # Python package dependencies
โ”œโ”€โ”€ .huggingface.yml       # Config for installing system packages (OCR support)
โ””โ”€โ”€ README.md              # Project documentation

โš™๏ธ Installation & Local Setup

bash
# Clone the repository
git clone https://github.com/sanyab1801/seo-intelligence-tool.git
cd seo-intelligence-tool

# Install Python dependencies
pip install -r requirements.txt

# Optional: install OCR dependencies (for .pdf scans, .doc support)
sudo apt install tesseract-ocr poppler-utils

# Run the app
streamlit run app.py
๐Ÿ“ On Windows, install Tesseract OCR and Poppler manually, and add to PATH.

๐Ÿ“ฆ Requirements

txt
streamlit
spacy
keybert
sentence-transformers
PyPDF2
python-docx
textract
pytesseract
pdf2image
beautifulsoup4
requests
The app uses the en_core_web_sm spaCy model. It will auto-download on first run.

๐Ÿ“‚ Supported File Types

FormatSupportNotes
.txtโœ…Direct text read
.pdfโœ…Uses PyPDF2 or Tesseract OCR if scanned
.docxโœ…Parsed using python-docx
.docโœ…Extracted using textract
Scanned .pdfโœ…OCR performed via pytesseract

๐Ÿง  Future Improvements (Optional Enhancements)

  • โ€”โœ… Export results to .csv or .json
  • โ€”โœ… Keyword density and readability scoring
  • โ€”โœ… SEO similarity comparison with competitor URLs
  • โ€”โœ… Language detection and support for multilingual input

๐Ÿ‘จโ€๐Ÿ’ป Author

Developed by \[Sanya Behera] GitHub: github.com/sanyab1801 Deployed on: Hugging Face Spaces