arcane-bit/book-finder-v2
Book Finder โ Big Data Engineering
End-to-End ETL Pipeline + Google Books Enrichment + FastAPI Service
Phase 2 Semantic Search Project
๐ Live Deployment: Book Finder v2 on Hugging Face Spaces
1. Introduction & Motivation
Modern semantic search systems rely on clean, structured, and enriched data. Our university library records (OPAC) are often:
- Sparse (basic title/author only)
- Missing rich metadata (descriptions, categories, thumbnails)
- Inconsistent for machine learning tasks
This project addresses those challenges by building a production-style data pipeline that transforms raw accession registers into a queryable, enriched database, and exposes the data through a FAST API for downstream applications such as:
- Semantic search (Phase 2)
- Recommendation systems
- Library analytics dashboards
The project follows real-world data engineering principles:
- Clear separation of pipeline stages
- Deterministic and resume-safe processing
- CLI-driven configuration
- Database-backed persistence
- API-based data access
2. High-Level Capabilities
This system provides:
- Robust Ingestion: Asynchronous fetching from Google Books API
- Smart Synchronization: Altcha-aware scraping of Koha OPAC
- Data Transformation: Deduplication, normalization, and cleaning
- Persistent Storage: SQLAlchemy + SQLite architecture
- Pipeline Orchestration: Unified controller for all stages
- FastAPI Service: REST interface for browsing and triggering syncs
- Fully Self-Documenting CLI:
--helpsupport for every script
3. Project Structure & Responsibilities
Each folder has one clear responsibility, mirroring how production pipelines are organized.
Book-Finder/
โ
โโโ api/
โ โโโ serving.py # FastAPI service & DB models
โโโ data/
โ โโโ raw/ # CSVs and BibTeX files
โ โโโ processed/ # Intermediate JSONL files
โโโ ingestion/
โ โโโ ingestion.py # Async Google Books fetcher
โโโ logs/
โ โโโ project_log.md # Technical development log
โโโ analysis/
โ โโโ metrics_analysis.py # Data quality reporting
โโโ storage/
โ โโโ storage.py # Database loading logic
โโโ Transformation/
โ โโโ transformation.py # Cleaning & Deduplication
โโโ main.py # Pipeline Orchestrator
โโโ sync_pipeline.py # OPAC Synchronization
โโโ README.mdThis structure ensures:
- Clear data lineage
- Easy debugging
- Independent execution of each stage
4. ๐ฝ Pipeline Architecture (End-to-End Flow)
The pipeline is linear, deterministic, and restartable.
โโโโโโโโโโโโโโโโโโโโโโโโ
โโโโโโโโโโโโโโโโโโโโโโโโ
โ Raw CSV Registers โ โ OPAC (Koha) โ
โ (Existing Data) โ โ (New Arrivals) โ
โโโโโโโโโโโโฌโโโโโโโโโโโโ โโโโโโโโโโโโฌโโโโโโโโโโโโ
โ โ
โผ โผ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ SYNC & MERGE โ
โ - Crawl New Arrivals (Altcha-aware) โ
โ - Merge with Accession Register โ
โ - Detect Incremental Changes โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโ
โ
โผ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ INGESTION (Google Books) โ
โ - Async I/O (aiohttp) โ
โ - Rate Limiting & Backoff โ
โ - Fetch Metadata (ISBN, Desc, Thumbnails) โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโ
โ
โผ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ TRANSFORMATION โ
โ - Normalize Titles/Authors โ
โ - Deduplicate (ISBN > Google ID > Title match) โ
โ - Merge Metadata Conflicts โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโ
โ
โผ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ STORAGE โ
โ - JSONL โ SQLite โ
โ - SQLAlchemy ORM โ
โ - Integrity Checks โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโ
โ
โผ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ FASTAPI SERVICE โ
โ - Browse & Paginate โ
โ - Search (Title/Author/ISBN) โ
โ - Trigger Pipeline Sync โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ5. Running the Complete Pipeline
To execute all ETL stages in order, run:
python main.pyYou can also skip specific stages or set limits:
python main.py --skip-sync --ingest-limit 100Final Artifact
book_finder.dbThis database becomes the single source of truth for the API.
6. Detailed Stage Explanations
6.1 Sync Stage (sync_pipeline.py)
Goal: Keep the local dataset aligned with the physical library's "New Arrivals".
Key Design Choices
- Altcha-Awareness: Detects anti-bot protection and degrades gracefully (warns user instead of crashing).
- Incremental Logic: Checks for new IDs against existing CSV to minimize redundant processing.
- BibTeX Parsing: Converts library standard format to project CSV schema.
Default Run
python sync_pipeline.py6.2 Ingestion Stage (ingestion.py)
Goal: Enrich sparse CSV records with rich metadata from Google Books.
Key Design Choices
- Asyncio + Aiohttp: High-throughput fetching.
- Semaphore Handling: Limits concurrency to prevent IP bans.
- Resumable State: Skips already processed IDs in
outputJSONL.
Default Run
python ingestion/ingestion.py --limit 100Custom Input / Output
python ingestion/ingestion.py \
--input "data/raw/custom_list.csv" \
--output "data/processed/my_enrichment.jsonl"6.3 Transformation Stage (transformation.py)
Goal: Clean, normalize, and deduplicate the noisy API results.
Operations
- Step 1 (Transform): Merges CSV metadata (Book No, Publisher) with Google Metadata.
- Step 2 (Dedup): Resolves duplicates using a hierarchy:
ISBN-13>Google ID>Title+Author.
Why this matters: API results often return the same book for slightly different queries. This stage ensures uniqueness.
Default Run
python Transformation/transformation.py6.4 Storage Stage (storage.py)
Goal: Persist final records into a structured Relational Database.
Design Choices
- SQLAlchemy: ORM-based access for future portability (e.g., to PostgreSQL).
- Batch Commits: Improves write performance.
- Integrity Handling: Skips duplicates at the database level if pre-checks fail.
Default Run
python storage/storage.py7. ๐ง Phase 2: Semantic Search & ML Implementation
Phase 1 focused on data engineering and enrichment. Phase 2 transforms that data into a searchable knowledge base using Vector Embeddings.
7.1 Model Choice: all-MiniLM-L6-v2
We use the all-MiniLM-L6-v2 model from the sentence-transformers library.
- Dimensionality: 384-dimensional dense vectors.
- Why this model?: It offers an excellent balance between speed and performance. It is small (~80MB) and fast enough for real-time CPU-based search on Hugging Face Spaces while maintaining high semantic accuracy.
- Embedding Generation: Captured in
ml/embeddings.pyusing theSentenceTransformerclass.
7.2 Vector Store: ChromaDB
- Technology: ChromaDB is used as our vector database. It is an open-source embedding database suited for local deployments.
- Configuration: Uses
hnsw:space: "cosine"to calculate similarity between the query and the library records. - Persistence: Data is saved locally in
data/vector_store/, allowing the index to be reusable without re-generation.
7.3 Indexing Pipeline (ml/index_books.py)
To build the searchable index, we perform the following:
- Context Assembly: For every book, we combine fields into a single "document" string:
Title | Subtitle | Authors | Categories | Description - Batch Processing: Books are processed in batches (default: 100) to optimize memory usage and indexing speed.
- Unique Identifiers: We use
ISBN-13as the primary key in ChromaDB to ensure a 1:1 mapping back to the SQLite records.
7.4 UI Integration (app.py)
The semantic search is seamlessly integrated into the Streamlit dashboard:
- On-the-fly Embedding: When a user types a query (e.g., "History of Space travel"), the app generates a vector for that query in real-time.
- Relevance Ranking: ChromaDB returns the Top 300 matches based on cosine distance.
- Threshold Filtering: We apply a
DISTANCE_THRESHOLD(0.7) to ensure only highly relevant books are displayed to the user.
8. FastAPI & UI Service
The FastAPI layer provides read-only access to the final dataset and control access to the pipeline.
Start API Server
python api/serving.py --reloador via the Streamlit UI:
streamlit run app.py -- --api-url http://localhost:8000GET /books/โ Paginated book listingGET /books/{isbn}โ ISBN lookupGET /search/?q=termโ Partial match searchGET /sync/โ Trigger background pipeline run
Swagger UI: http://127.0.0.1:8000/docs
9. Pipeline Statistics
All statistics can be reproduced using analysis/metrics_analysis.py or the analysis/project_metrics.ipynb notebook.
9.1 Data Flow Summary
9.2 Enrichment Quality
Observation
The pipeline successfully enriches over 90% of the library catalog, proving the effectiveness of the fuzzy matching strategy used during ingestion.
10. Data Dictionary (Core Fields)
11. Design Philosophy
This project emphasizes:
- Separation of concerns: Each module does one thing well.
- Fail-Safe Operation: Network errors or API limits do not crash the pipeline.
- Reproducibility: Everything is code-defined and scriptable.
- Transparency: Extensive logging (
logs/project_log.md) tracks all decisions.
12. Conclusion
This project demonstrates a complete, production-style data pipeline:
- Quantifiable data-quality improvements
- Deterministic ETL stages
- Resume-safe enrichment
- Persistent storage
- API-based data access
It bridges the gap between 'operational' library lists and 'analytical' datasets, forming a strong foundation for Phase 2: Semantic Search.
Authors
202518053 : Falak Parmar 202518035 : Aditya Jana
DA-IICT โ Big Data Engineering
