build-small-hackathon/wikipedia-companion
๐ฃ๏ธ Wikipedia Companion
Growing up and working in Uganda, I learned that having internet doesn't mean having equal internet. Data is expensive and slow. Text is cheap to load; rich educational video isn't, because it eats gigabytes and buffers endlessly. So I built the Wikipedia Companion: ask a question out loud or by typing, and a friendly animated face reads the answer back while the words light up so you can follow along. The trick is that all the AI runs on-device: speech recognition, the language model, the voice, and the lip-sync, all on small models. The only thing it pulls from the internet is the article's text, a few kilobytes. You get the warmth and engagement of a video tutor at a tiny fraction of the data, built for the community I come from.
All AI runs locally; only Wikipedia's text is fetched over the internet.
- Speech-to-text: faster-whisper (local)
- Query understanding + answers: Qwen2.5-3B via llama.cpp (local)
- Semantic retrieval: all-MiniLM-L6-v2 (local)
- Text normalization: NVIDIA NeMo
- Speech synthesis: Kokoro (local)
- Follow-along reading & Lip-sync: Wav2Vec2-CTC forced alignment (local)
Built for the Build Small Hackathon. Source code, demo video, HuggingFace space, blog post, and social post:
- Source code (automated tests, linting, and CI authored with OpenAI Codex)
- Demo video
- HuggingFace space
- Blog post (includes field notes and instructions on how to run this companion locally)
- Social post
How it works
The whole system is a single Gradio app. While the diagram above maps out the 18 technical interactions, the pipeline boils down to 7 core stages where every AI model runs locally:
- Listen: your voice is transcribed locally with Whisper.
- Understand the question: the local LLM reads "how big is the moon" and figures out the subject is the Moon. It proposes candidate article titles, which are verified against Wikipedia before use.
- Fetch: the chosen article's text is pulled live from the Wikipedia API (the only network call!).
- Retrieve: a tiny embedding model semantically ranks the article's sections to find the passage that actually answers the question.
- Explain: the local LLM turns that passage into a short, warm, spoken answer, grounded only in what the passage says.
- Speak: the text is normalized for natural speech (numbers, units, abbreviations expanded via NVIDIA NeMo), then synthesized to audio with Kokoro.
- Come alive: forced alignment (Wav2Vec2-CTC) produces millisecond word-level timings, which drive both the animated face's lip movements and the follow-along transcript.
๐ Follow-along reading experience
For learners, especially in communities where the language of instruction is a second language, just listening isn't enough. The app extracts millisecond word-level timestamps using forced alignment to create a "karaoke-style" transcript. As the companion speaks, the text highlights word-by-word on the screen, actively bridging the gap between spoken and written language to improve reading comprehension and literacy.
๐ A zero-bandwidth animated face (Off Brand)
To provide the psychological engagement of a video tutor without the massive bandwidth tax of video streaming, the interface centers on an animated SVG character:
- Breathing idle animation so it feels alive while waiting.
- Blinking eyelids on a natural timer.
- Mood states (
rest,thinking, andspeaking) with a soft glow that signals when it's working versus talking. - Six visemes (
ah,ee,oh,oo,mb,rest): distinct mouth shapes that are driven by real word-level timings from Wav2Vec2-CTC forced alignment. The mouth actually lip-syncs to the spoken answer rather than flapping randomly.
The lip-sync data flows straight from the Python pipeline into the browser runtime (companion/static/companion.js) on each answer, with no DOM scraping, a direct data channel.
๐ Fully local inference (Off the Grid & Tiny Titan)
No cloud AI. Every model runs locally on the machine in front of you. This is what allows the app to transform a tiny, cheap kilobyte text payload into a rich multimedia experience.
Why fetch text instead of bundling Wikipedia offline? A full local Wikipedia database dump is tens of gigabytes to download, instantly eating up phone/laptop storage, and is frozen at download time. By fetching just the specific article's text live on-demand, the app guarantees the answers use today's facts, not last year's snapshot, while still only costing a fraction of a cent in data, unlike streaming video.
Every model stays small, well under 4B parameters each:
The largest model is ~3B and every model is โค 4B, qualifying for the Tiny Titan badge: small weights, doing real work.
๐ฆ Runs on llama.cpp (Llama Champion)
The language model, Qwen2.5-3B-Instruct (GGUF, q4km), is served entirely through the llama.cpp runtime via llama-cpp-python. On the HuggingFace GPU Space, it runs as a CUDA build with all 37/37 transformer layers offloaded to the GPU, so a full grounded answer is generated in roughly two seconds. The same code path falls back to CPU inference for fully local, offline use.
