Team Ai
Apppublic

build-small-hackathon/wikipedia-companion

sourceHugging Facemitupdated 4mo agoView on Hugging Face
1likes
App README

๐Ÿ—ฃ๏ธ Wikipedia Companion

Growing up and working in Uganda, I learned that having internet doesn't mean having equal internet. Data is expensive and slow. Text is cheap to load; rich educational video isn't, because it eats gigabytes and buffers endlessly. So I built the Wikipedia Companion: ask a question out loud or by typing, and a friendly animated face reads the answer back while the words light up so you can follow along. The trick is that all the AI runs on-device: speech recognition, the language model, the voice, and the lip-sync, all on small models. The only thing it pulls from the internet is the article's text, a few kilobytes. You get the warmth and engagement of a video tutor at a tiny fraction of the data, built for the community I come from.

All AI runs locally; only Wikipedia's text is fetched over the internet.

โ–ถ๏ธ Watch the Video Demo!

  • โ€”Speech-to-text: faster-whisper (local)
  • โ€”Query understanding + answers: Qwen2.5-3B via llama.cpp (local)
  • โ€”Semantic retrieval: all-MiniLM-L6-v2 (local)
  • โ€”Text normalization: NVIDIA NeMo
  • โ€”Speech synthesis: Kokoro (local)
  • โ€”Follow-along reading & Lip-sync: Wav2Vec2-CTC forced alignment (local)

Built for the Build Small Hackathon. Source code, demo video, HuggingFace space, blog post, and social post:

How it works

[image]

The whole system is a single Gradio app. While the diagram above maps out the 18 technical interactions, the pipeline boils down to 7 core stages where every AI model runs locally:

  1. 1.Listen: your voice is transcribed locally with Whisper.
  2. 2.Understand the question: the local LLM reads "how big is the moon" and figures out the subject is the Moon. It proposes candidate article titles, which are verified against Wikipedia before use.
  3. 3.Fetch: the chosen article's text is pulled live from the Wikipedia API (the only network call!).
  4. 4.Retrieve: a tiny embedding model semantically ranks the article's sections to find the passage that actually answers the question.
  5. 5.Explain: the local LLM turns that passage into a short, warm, spoken answer, grounded only in what the passage says.
  6. 6.Speak: the text is normalized for natural speech (numbers, units, abbreviations expanded via NVIDIA NeMo), then synthesized to audio with Kokoro.
  7. 7.Come alive: forced alignment (Wav2Vec2-CTC) produces millisecond word-level timings, which drive both the animated face's lip movements and the follow-along transcript.

๐Ÿ“– Follow-along reading experience

For learners, especially in communities where the language of instruction is a second language, just listening isn't enough. The app extracts millisecond word-level timestamps using forced alignment to create a "karaoke-style" transcript. As the companion speaks, the text highlights word-by-word on the screen, actively bridging the gap between spoken and written language to improve reading comprehension and literacy.

๐Ÿ˜Š A zero-bandwidth animated face (Off Brand)

To provide the psychological engagement of a video tutor without the massive bandwidth tax of video streaming, the interface centers on an animated SVG character:

  • โ€”Breathing idle animation so it feels alive while waiting.
  • โ€”Blinking eyelids on a natural timer.
  • โ€”Mood states (rest, thinking, and speaking) with a soft glow that signals when it's working versus talking.
  • โ€”Six visemes (ah, ee, oh, oo, mb, rest): distinct mouth shapes that are driven by real word-level timings from Wav2Vec2-CTC forced alignment. The mouth actually lip-syncs to the spoken answer rather than flapping randomly.

The lip-sync data flows straight from the Python pipeline into the browser runtime (companion/static/companion.js) on each answer, with no DOM scraping, a direct data channel.

๐Ÿ”‹ Fully local inference (Off the Grid & Tiny Titan)

No cloud AI. Every model runs locally on the machine in front of you. This is what allows the app to transform a tiny, cheap kilobyte text payload into a rich multimedia experience.

Why fetch text instead of bundling Wikipedia offline? A full local Wikipedia database dump is tens of gigabytes to download, instantly eating up phone/laptop storage, and is frozen at download time. By fetching just the specific article's text live on-demand, the app guarantees the answers use today's facts, not last year's snapshot, while still only costing a fraction of a cent in data, unlike streaming video.

Every model stays small, well under 4B parameters each:

ComponentModel~Params
Query understanding + answersQwen2.5-3B-Instruct (GGUF, q4km)~3B
Speech-to-textfaster-whisper (base)~74M
Semantic retrievalall-MiniLM-L6-v2~22M
Speech synthesisKokoro~82M
Lip-sync alignmentWav2Vec2-CTC (base)~95M

The largest model is ~3B and every model is โ‰ค 4B, qualifying for the Tiny Titan badge: small weights, doing real work.

๐Ÿฆ™ Runs on llama.cpp (Llama Champion)

The language model, Qwen2.5-3B-Instruct (GGUF, q4km), is served entirely through the llama.cpp runtime via llama-cpp-python. On the HuggingFace GPU Space, it runs as a CUDA build with all 37/37 transformer layers offloaded to the GPU, so a full grounded answer is generated in roughly two seconds. The same code path falls back to CPU inference for fully local, offline use.