Team Ai
Apppublic

bep40/vnews-image-zerogpu

sourceHugging Faceupdated 7d agoView on Hugging Face
0likes
App README

VNEWS Image — SenseNova-U1.5-8B-MoT on ZeroGPU

Image generation AND image editing for the VNEWS news app + Zalo bot (`bep40/vaistudio-zalo-bot`).

Uses `sensenova/SenseNova-U1.5-8B-MoT` — the FULL bf16 checkpoint (35 GB), identical to the official hugging-apps/sensenova-sensenova-u1-5-8b-mot demo, via the ZeroGPU execution model: model loaded at module level (CPU-backed "cuda" outside @spaces.GPU), packed to disk at startup, moved onto the real GPU (H200 ~69 GiB / Blackwell 48 GiB on the current zero-a10g pool) per call. No GGUF, no offload.

Reference-quality sampling recipe from the official demo:

  • —cfg_scale=4.0, timestep_shift=3.0, cfg_interval=(0.0, 1.0)
  • —num_steps default 28 (max 50, model-card reference)
  • —native T2I resolution buckets (3:4 → 1760×2368, 1:1 → 2048×2048, …) then LANCZOS resize to the caller's requested size
  • —Image editing (`it2i_generate`): same recipe + img_cfg_scale=1.0, output keeps the input aspect ratio (smart_resize to ~2048², factor 32)

sensenova_u1 is vendored locally (same package as the official U1.5 demo).

API

  • —POST /gradio_api/call/generate_bg — {"data": ["prompt", 1024, 1536, 28, -1]} → returns event id; poll GET .../generate_bg/<id> (status) and GET .../generate_bg/<id>/ (JSON result {"data":[...]}).
  • —POST /gradio_api/call/generate_text — {"data": ["topic", 4]} → fresh Vietnamese text generated from the topic (no news-title regurgitation).
  • —POST /gradio_api/call/edit_image — {"data": ["<edit prompt>", <image>, 28, -1]}. The image slot accepts the server-local path returned by POST /gradio_api/upload ({"files": [...]} → ["/tmp/gradio/.../x.png"]). Returns the edited image (GR_Image), same aspect ratio as the input.

ZeroGPU notes

Concurrency is serialized (default_concurrency_limit=1): a second 35 GB worker on a shared 48 GB Blackwell half is evicted instantly. The edit path reserves max(scaled+60, t2i_budget)+120s so a cold start (35 GB transfer + fork) never evicts a run mid-sampling. Clients should retry once on the transient event: error, data: null that ZeroGPU occasionally emits while a worker is being forked.