Team Ai
Apppublic

build-small-hackathon/CodeFlow

sourceHugging Facemitupdated 4mo agoView on Hugging Face
10likes
App README

πŸ“Š CodeFlow

Paste code β†’ read its logic as a flowchart. A 30B coder model runs entirely on CPU via llama.cpp to translate source code into a clean, animated Mermaid.js control-flow diagram β€” with each node wired back to the exact lines it came from.

πŸ”— Links

[πŸš€ Live Space][space] Β· [▢️ Demo Video][video] Β· [🐦 Social Post][social] Β· [πŸ““ Field Notes (blog)][blog] Β· [πŸ” Agent Traces][traces] Β· [πŸŽ›οΈ Fine-Tuned Model][model]

[space]: https://huggingface.co/spaces/build-small-hackathon/CodeFlow "Hugging Face Space" [video]: https://youtu.be/R5GbpN9FVxo "Demo video" [social]: https://www.linkedin.com/feed/update/urn:li:share:7471327684539785217/ "Social post" [blog]: https://huggingface.co/blog/build-small-hackathon/codeflow-field-notes "Field notes / blog post" [traces]: https://huggingface.co/datasets/build-small-hackathon/codeflow-agent-traces "Agent traces dataset" [model]: https://huggingface.co/build-small-hackathon/codeflow-qwen-3-finetuning "Fine-tuned model"


❓ The Problem

Reading unfamiliar code means simulating its control flow in your head β€” chasing branches, loops, and early returns line by line. That's slow, error-prone, and gets worse the deeper the nesting. Existing "code β†’ diagram" tools are usually rigid AST parsers (brittle, language-locked) or cloud LLM APIs (your code leaves the building).

CodeFlow turns any snippet into a scannable flowchart you can audit at a glance β€” generated by a real language model that runs 100% locally, so nothing is sent to an external API.

βš™οΈ How It Works

 Paste code ──▢ Generate ──▢ POST /generate_flowchart        (Gradio API)
                                    β”‚
                    number the source lines + structured system prompt
                                    β”‚
          CodeFlow fine-tune of Qwen3-Coder-30B-A3B  (llama.cpp Β· CPU)
                                    β”‚
                 <thinking> …reasoning… </thinking>
                 graph TD … nodes & edges …
                 <linemap> A:1  B:2  C:3-4 </linemap>
                                    β”‚
        strip reasoning Β· parse + validate the line-map Β· sanitize labels
                                    β”‚
                  { mermaid, linemap }  ──▢  append agent_traces.jsonl
                                    β”‚
   Mermaid render + "trace-the-path" reveal + node ↔ code linking
  1. 1.You paste code (or pick a pre-rendered example) into the CodeMirror editor and hit Generate.
  2. 2.The backend numbers the source lines and sends them with a strict system prompt to the CodeFlow fine-tune of Qwen3-Coder running on llama.cpp.
  3. 3.The model returns hidden <thinking>, the Mermaid graph, and a <linemap> mapping every node to its source line(s).
  4. 4.The server strips the reasoning, validates the line-map against the source, sanitizes labels for Mermaid, and returns { mermaid, linemap }.
  5. 5.The frontend renders the diagram with a trace-the-path reveal that flows out of a persistent Start node while the canvas scrolls along in real time.
  6. 6.Node ↔ code linking: hover a node to highlight its source lines, click a node to jump-and-edit them, or move your cursor over a line to light up the matching node.
  7. 7.Every generation is captured as a structured agent trace (/traces).

πŸŽ›οΈ Fine-Tuning

CodeFlow runs a [LoRA fine-tune][model] of Qwen3-Coder-30B-A3B-Instruct (β‰ˆ30.5B params), specialized for the code β†’ Mermaid + <linemap> task rather than relying on the base model's general coding ability.

  • β€”Data: 2,400 synthetic examples (2,208 train / 192 val β€” 8% holdout), built from 22 control-flow templates across Python, JavaScript, C++, and C.
  • β€”Method: LoRA r=16, Ξ±=32 on the attention + MLP projections, bf16, cosine schedule β€” then merged and exported to a Q3_K_L GGUF for CPU inference.
  • β€”Validation: the holdout is hard-validated β€” generated outputs are syntax-checked / compiled, not just eyeballed.

See the [model card][model] for the full data engine, finetune.py options, and dataset preview.

🧰 Tech Stack

LayerWhat it isUsed for
Model[CodeFlow fine-tune][model] of Qwen3-Coder-30B-A3B-Instruct (Mixture-of-Experts)Code β†’ Mermaid + line-map generation
Fine-tuningLoRA SFT (r=16, Ξ±=32) on attention + MLP projections, merged to GGUFSpecializes the base model for the code β†’ Mermaid + line-map task
QuantizationQ3_K_L GGUF (~3-bit)Shrinks the 30B model to run on CPU
Inference`llama-cpp-python` (llama.cpp)Local CPU inference (n_ctx=4096)
Model fetchhuggingface_hubDownloads the GGUF on first run
ServerGradio gr.Server + FastAPI/generate_flowchart API, / UI, /traces
FrontendA single self-contained frontend.html (vanilla JS + CSS custom properties)Editor, diagram, animation, theming
EditorCodeMirror 6 β€” vendored bundle (static/cm.bundle.js)Syntax-highlighted code input
DiagramsMermaid.js 10 β€” vendored UMD (static/mermaid.min.js)Flowchart rendering
AnimationWeb Animations APITrace-the-path reveal + theme crossfade
TypeFraunces Β· Hanken Grotesk Β· JetBrains Mono β€” vendored woff2 (static/fonts/)Custom, non-default look
AssetsAll JS/CSS/fonts bundled into static/ (no CDN at runtime)True offline operation
ObservabilityHand-rolled JSONL agent tracesOne trace per generation, served at /traces
Testssmoke-test.sh (headless Chrome)13 build/render checks
DeployHugging Face SpacesHosting

πŸ”’ Total Parameters

CodeFlow is driven by a [LoRA fine-tune][model] of Qwen3-Coder-30B-A3B-Instruct β€” a Mixture-of-Experts model with:

  • β€”β‰ˆ 30.5 billion total parameters (well under the 32B cap)
  • β€”β‰ˆ 3.3 billion active parameters per token (128 experts, 8 activated)

It's served as a ~3-bit (Q3_K_L) GGUF, which compresses those 30B weights to a CPU-runnable footprint (~13 GB on disk) β€” letting a 30B-class model generate diagrams off the grid, with no GPU and no external API.

πŸ… Badges (6 / 6)

These map to the Space tags above.

BadgeHow CodeFlow earns it
πŸ”Œ Off the GridNo external API or CDN at runtime β€” period. The model runs fully locally (Qwen3-Coder GGUF on CPU via llama.cpp), and every frontend asset (Mermaid, CodeMirror, the Gradio client, all fonts) is vendored into static/. The Gradio share tunnel is off (share=False). The only network call in the whole project is the one-time model download at startup. The UI even runs fully offline from file://.
🎨 Off-BrandZero default-Gradio look. A bespoke single-file UI: custom "Pine & Sage" palette (one-word rust fallback), Fraunces + Hanken Grotesk type, a hand-drawn decision-node logo, restyled Mermaid nodes, and a trace-the-path reveal animation β€” deliberately designed not to look templated.
πŸ““ Field NotesSee the [blog post][blog].
🀝 Sharing is CaringOpen-source under MIT, a public Space, plus a [social post][social] sharing the process and learnings.
πŸ€– AgenticEvery model generation is captured as a structured agent trace (input code, the model's reasoning, output, token usage, latency), downloadable at [/traces][traces].
πŸŽ›οΈ Well-TunedA [LoRA fine-tune][model] of Qwen3-Coder-30B-A3B-Instruct (β‰ˆ30.5B params β€” under the 32B cap), specialized for the code β†’ Mermaid + <linemap> task and shipped as the GGUF the Space actually runs.

πŸŽ₯ Demo

▢️ [Watch the demo video][video] β€” a full walkthrough of CodeFlow in action.

πŸ’» Run It Locally

First launch downloads the ~13 GB GGUF from Hugging Face. CPU inference is slow (cold generations can take minutes) β€” the built-in examples render instantly because their diagrams are pre-computed.
bash
# 1. Clone
git clone https://huggingface.co/spaces/build-small-hackathon/CodeFlow CodeFlow
cd CodeFlow

# 2. Create a virtual env
python -m venv .venv
source .venv/bin/activate        # Windows: .venv\Scripts\activate

# 3. Install deps (uses a prebuilt CPU wheel for llama-cpp-python)
pip install -r requirements.txt

# 4. Run β€” opens a local Gradio URL
python app.py

Then open the printed URL. Preview the UI without the model by opening frontend.html directly in a browser (file://) β€” fully offline, since all assets are vendored in static/; the example presets render their diagrams instantly.

Rebuilding the vendored bundles (optional): the CodeMirror + Gradio-client bundles in static/ are produced by build/build.sh (needs Node). Mermaid and the fonts are downloaded into static/ as well. You never need this to run the app β€” only to regenerate the bundles.

Endpoints: / (UI) Β· /generate_flowchart (API) Β· /traces (download all agent traces as JSONL).

πŸ—‚οΈ Repository Structure

CodeFlow/
β”œβ”€β”€ app.py             # Gradio + FastAPI server: loads the model and exposes
β”‚                      #   /generate_flowchart (API), / (UI), /static, /traces
β”œβ”€β”€ frontend.html      # Self-contained UI β€” CodeMirror editor, Mermaid render,
β”‚                      #   trace-the-path animation, node↔code linking, theming
β”œβ”€β”€ static/            # Vendored frontend assets β€” NO CDN at runtime
β”‚   β”œβ”€β”€ mermaid.min.js #   Mermaid (UMD, ~3.2 MB)
β”‚   β”œβ”€β”€ cm.bundle.js   #   CodeMirror 6 (single IIFE bundle)
β”‚   β”œβ”€β”€ gradio-client.js #  @gradio/client (IIFE bundle)
β”‚   β”œβ”€β”€ fonts.css      #   @font-face β†’ local woff2
β”‚   └── fonts/         #   Fraunces Β· Hanken Grotesk Β· JetBrains Mono (woff2)
β”œβ”€β”€ build/             # Reproducible bundle build (Node) β€” build.sh + entry files
β”œβ”€β”€ requirements.txt   # Python deps (CPU llama-cpp-python wheel, gradio, hub)
β”œβ”€β”€ smoke-test.sh      # Headless-Chrome smoke test (13 checks)
β”œβ”€β”€ notes-for-blog.md  # Field Notes β€” the full build log
β”œβ”€β”€ README.md          # You are here
└── LICENSE            # MIT

⚠️ Limitations

  • β€”CPU inference is slow. A 30B model on CPU means cold generations can take minutes; the demo leans on pre-rendered examples for instant feedback.
  • β€”3-bit quantization trades some fidelity for the ability to run a 30B model at all β€” occasional imperfect diagrams.
  • β€”4096-token context β€” very large files won't fit; works best on functions/snippets.
  • β€”Line-map depends on the model. The <linemap> is LLM-generated; the server validates and drops bad entries, so node↔code links can be partial on tricky code.
  • β€”Paraphrased labels. Nodes describe logic in plain words (no raw code), so they read cleanly but aren't verbatim.
  • β€”Mermaid parse failures on unusual syntax are possible (the raw output is shown so nothing is lost).
  • β€”Ephemeral traces on Spaces. agent_traces.jsonl lives on the runtime filesystem and resets on restart/rebuild β€” download it before then.

πŸ™ Credits

πŸ“„ License

Released under the MIT License β€” see `LICENSE`. Β© 2026 Rishi Jain.