Team Ai
Apppublic

arcane-bit/book-finder-v2

sourceHugging Faceupdated 8mo agoView on Hugging Face
1likes
App README

Book Finder โ€” Big Data Engineering

End-to-End ETL Pipeline + Google Books Enrichment + FastAPI Service

Phase 2 Semantic Search Project

๐Ÿš€ Live Deployment: Book Finder v2 on Hugging Face Spaces

1. Introduction & Motivation

Modern semantic search systems rely on clean, structured, and enriched data. Our university library records (OPAC) are often:

  • โ€”Sparse (basic title/author only)
  • โ€”Missing rich metadata (descriptions, categories, thumbnails)
  • โ€”Inconsistent for machine learning tasks

This project addresses those challenges by building a production-style data pipeline that transforms raw accession registers into a queryable, enriched database, and exposes the data through a FAST API for downstream applications such as:

  • โ€”Semantic search (Phase 2)
  • โ€”Recommendation systems
  • โ€”Library analytics dashboards

The project follows real-world data engineering principles:

  • โ€”Clear separation of pipeline stages
  • โ€”Deterministic and resume-safe processing
  • โ€”CLI-driven configuration
  • โ€”Database-backed persistence
  • โ€”API-based data access

2. High-Level Capabilities

This system provides:

  • โ€”Robust Ingestion: Asynchronous fetching from Google Books API
  • โ€”Smart Synchronization: Altcha-aware scraping of Koha OPAC
  • โ€”Data Transformation: Deduplication, normalization, and cleaning
  • โ€”Persistent Storage: SQLAlchemy + SQLite architecture
  • โ€”Pipeline Orchestration: Unified controller for all stages
  • โ€”FastAPI Service: REST interface for browsing and triggering syncs
  • โ€”Fully Self-Documenting CLI: --help support for every script

3. Project Structure & Responsibilities

Each folder has one clear responsibility, mirroring how production pipelines are organized.

Book-Finder/
โ”‚
โ”œโ”€โ”€ api/
โ”‚   โ””โ”€โ”€ serving.py          # FastAPI service & DB models
โ”œโ”€โ”€ data/
โ”‚   โ”œโ”€โ”€ raw/                # CSVs and BibTeX files
โ”‚   โ”œโ”€โ”€ processed/          # Intermediate JSONL files
โ”œโ”€โ”€ ingestion/
โ”‚   โ””โ”€โ”€ ingestion.py        # Async Google Books fetcher
โ”œโ”€โ”€ logs/
โ”‚   โ””โ”€โ”€ project_log.md      # Technical development log
โ”œโ”€โ”€ analysis/
โ”‚   โ””โ”€โ”€ metrics_analysis.py # Data quality reporting
โ”œโ”€โ”€ storage/
โ”‚   โ””โ”€โ”€ storage.py          # Database loading logic
โ”œโ”€โ”€ Transformation/
โ”‚   โ””โ”€โ”€ transformation.py   # Cleaning & Deduplication
โ”œโ”€โ”€ main.py                 # Pipeline Orchestrator
โ”œโ”€โ”€ sync_pipeline.py        # OPAC Synchronization
โ””โ”€โ”€ README.md

This structure ensures:

  • โ€”Clear data lineage
  • โ€”Easy debugging
  • โ€”Independent execution of each stage

4. ๐Ÿ”ฝ Pipeline Architecture (End-to-End Flow)

The pipeline is linear, deterministic, and restartable.

โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚   Raw CSV Registers  โ”‚     โ”‚   OPAC (Koha)        โ”‚
โ”‚  (Existing Data)     โ”‚     โ”‚   (New Arrivals)     โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜     โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
           โ”‚                            โ”‚
           โ–ผ                            โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚ SYNC & MERGE                                      โ”‚
โ”‚ - Crawl New Arrivals (Altcha-aware)               โ”‚
โ”‚ - Merge with Accession Register                   โ”‚
โ”‚ - Detect Incremental Changes                      โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                           โ”‚
                           โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚ INGESTION (Google Books)                          โ”‚
โ”‚ - Async I/O (aiohttp)                             โ”‚
โ”‚ - Rate Limiting & Backoff                         โ”‚
โ”‚ - Fetch Metadata (ISBN, Desc, Thumbnails)         โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                           โ”‚
                           โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚ TRANSFORMATION                                    โ”‚
โ”‚ - Normalize Titles/Authors                        โ”‚
โ”‚ - Deduplicate (ISBN > Google ID > Title match)    โ”‚
โ”‚ - Merge Metadata Conflicts                        โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                           โ”‚
                           โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚ STORAGE                                           โ”‚
โ”‚ - JSONL โ†’ SQLite                                  โ”‚
โ”‚ - SQLAlchemy ORM                                  โ”‚
โ”‚ - Integrity Checks                                โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                           โ”‚
                           โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚ FASTAPI SERVICE                                   โ”‚
โ”‚ - Browse & Paginate                               โ”‚
โ”‚ - Search (Title/Author/ISBN)                      โ”‚
โ”‚ - Trigger Pipeline Sync                           โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

5. Running the Complete Pipeline

To execute all ETL stages in order, run:

bash
python main.py

You can also skip specific stages or set limits:

bash
python main.py --skip-sync --ingest-limit 100

Final Artifact

book_finder.db

This database becomes the single source of truth for the API.


6. Detailed Stage Explanations


6.1 Sync Stage (sync_pipeline.py)

Goal: Keep the local dataset aligned with the physical library's "New Arrivals".

Key Design Choices

  • โ€”Altcha-Awareness: Detects anti-bot protection and degrades gracefully (warns user instead of crashing).
  • โ€”Incremental Logic: Checks for new IDs against existing CSV to minimize redundant processing.
  • โ€”BibTeX Parsing: Converts library standard format to project CSV schema.

Default Run

bash
python sync_pipeline.py

6.2 Ingestion Stage (ingestion.py)

Goal: Enrich sparse CSV records with rich metadata from Google Books.

Key Design Choices

  • โ€”Asyncio + Aiohttp: High-throughput fetching.
  • โ€”Semaphore Handling: Limits concurrency to prevent IP bans.
  • โ€”Resumable State: Skips already processed IDs in output JSONL.

Default Run

bash
python ingestion/ingestion.py --limit 100

Custom Input / Output

bash
python ingestion/ingestion.py \
  --input "data/raw/custom_list.csv" \
  --output "data/processed/my_enrichment.jsonl"

6.3 Transformation Stage (transformation.py)

Goal: Clean, normalize, and deduplicate the noisy API results.

Operations

  • โ€”Step 1 (Transform): Merges CSV metadata (Book No, Publisher) with Google Metadata.
  • โ€”Step 2 (Dedup): Resolves duplicates using a hierarchy: ISBN-13 > Google ID > Title+Author.

Why this matters: API results often return the same book for slightly different queries. This stage ensures uniqueness.

Default Run

bash
python Transformation/transformation.py

6.4 Storage Stage (storage.py)

Goal: Persist final records into a structured Relational Database.

Design Choices

  • โ€”SQLAlchemy: ORM-based access for future portability (e.g., to PostgreSQL).
  • โ€”Batch Commits: Improves write performance.
  • โ€”Integrity Handling: Skips duplicates at the database level if pre-checks fail.

Default Run

bash
python storage/storage.py

7. ๐Ÿง  Phase 2: Semantic Search & ML Implementation

Phase 1 focused on data engineering and enrichment. Phase 2 transforms that data into a searchable knowledge base using Vector Embeddings.

7.1 Model Choice: all-MiniLM-L6-v2

We use the all-MiniLM-L6-v2 model from the sentence-transformers library.

  • โ€”Dimensionality: 384-dimensional dense vectors.
  • โ€”Why this model?: It offers an excellent balance between speed and performance. It is small (~80MB) and fast enough for real-time CPU-based search on Hugging Face Spaces while maintaining high semantic accuracy.
  • โ€”Embedding Generation: Captured in ml/embeddings.py using the SentenceTransformer class.

7.2 Vector Store: ChromaDB

  • โ€”Technology: ChromaDB is used as our vector database. It is an open-source embedding database suited for local deployments.
  • โ€”Configuration: Uses hnsw:space: "cosine" to calculate similarity between the query and the library records.
  • โ€”Persistence: Data is saved locally in data/vector_store/, allowing the index to be reusable without re-generation.

7.3 Indexing Pipeline (ml/index_books.py)

To build the searchable index, we perform the following:

  1. 1.Context Assembly: For every book, we combine fields into a single "document" string: Title | Subtitle | Authors | Categories | Description
  2. 2.Batch Processing: Books are processed in batches (default: 100) to optimize memory usage and indexing speed.
  3. 3.Unique Identifiers: We use ISBN-13 as the primary key in ChromaDB to ensure a 1:1 mapping back to the SQLite records.

7.4 UI Integration (app.py)

The semantic search is seamlessly integrated into the Streamlit dashboard:

  • โ€”On-the-fly Embedding: When a user types a query (e.g., "History of Space travel"), the app generates a vector for that query in real-time.
  • โ€”Relevance Ranking: ChromaDB returns the Top 300 matches based on cosine distance.
  • โ€”Threshold Filtering: We apply a DISTANCE_THRESHOLD (0.7) to ensure only highly relevant books are displayed to the user.

8. FastAPI & UI Service

The FastAPI layer provides read-only access to the final dataset and control access to the pipeline.

Start API Server

bash
python api/serving.py --reload

or via the Streamlit UI:

bash
streamlit run app.py -- --api-url http://localhost:8000
  • โ€”GET /books/ โ€“ Paginated book listing
  • โ€”GET /books/{isbn} โ€“ ISBN lookup
  • โ€”GET /search/?q=term โ€“ Partial match search
  • โ€”GET /sync/ โ€“ Trigger background pipeline run

Swagger UI: http://127.0.0.1:8000/docs


9. Pipeline Statistics

All statistics can be reproduced using analysis/metrics_analysis.py or the analysis/project_metrics.ipynb notebook.

9.1 Data Flow Summary

StageInput RecordsOutput RecordsSuccess Rate
Source (CSV)36,358--
Ingestion~36,35833,50292.1%
Transformation33,50226,17378.1% (Dedup)
Final DB26,17326,173100%

9.2 Enrichment Quality

MetricValue
Total Source Records36,358
Successful Matches33,502
Final Unique Books26,173
Database Size~41 MB

Observation

The pipeline successfully enriches over 90% of the library catalog, proving the effectiveness of the fuzzy matching strategy used during ingestion.

10. Data Dictionary (Core Fields)

FieldDescription
idInternal DB Primary Key
isbn_1313-digit ISBN (Primary Identifier)
titleBook Title (from Google Books)
authorsComma-separated list of authors
descriptionFull text summary/blurb
categoriesGenre/Subject tags
thumbnailURL to cover image
average_ratingGoogle Books rating
book_noOriginal Library Call Number

11. Design Philosophy

This project emphasizes:

  • โ€”Separation of concerns: Each module does one thing well.
  • โ€”Fail-Safe Operation: Network errors or API limits do not crash the pipeline.
  • โ€”Reproducibility: Everything is code-defined and scriptable.
  • โ€”Transparency: Extensive logging (logs/project_log.md) tracks all decisions.

12. Conclusion

This project demonstrates a complete, production-style data pipeline:

  • โ€”Quantifiable data-quality improvements
  • โ€”Deterministic ETL stages
  • โ€”Resume-safe enrichment
  • โ€”Persistent storage
  • โ€”API-based data access

It bridges the gap between 'operational' library lists and 'analytical' datasets, forming a strong foundation for Phase 2: Semantic Search.


Authors

202518053 : Falak Parmar 202518035 : Aditya Jana

DA-IICT โ€” Big Data Engineering