Brunobkr/llama.cpp_AlgMor24_github
ΩFFFΣLLIa • llama.cpp • AlgMor24 ██████╗ ███████╗███████╗███████╗██╗ ██╗ ██╗ █████╗ ██╔═══██╗██╔════╝██╔════╝██╔════╝██║ ██║ ██║██╔══██╗ ██║ ██║█████╗ █████╗ █████╗ ██║ ██║ ██║███████║ ██║ ██║██╔══╝ ██╔══╝ ██╔══╝ ██║ ██║ ██║██╔══██║ ╚██████╔╝██║ ██║ ███████╗███████╗███████╗██║██║ ██║ ╚═════╝ ╚═╝ ╚═╝ ╚══════╝╚══════╝╚══════╝╚═╝╚═╝ ╚═╝ High-Performance LLM / VLM Inference & Autonomous Agentic Ecosystem… See the full description on the dataset page: https://huggingface.co/datasets/Brunobkr/llama.cpp_AlgMor24_github.
03.1k
1# Agentic thread perf harness2 3Two tiers, both reusing the existing vitest projects (see `vite.config.ts`).4 5## Tier 1 - `agentic-stream.perf.svelte.test.ts` (project: `client`, real Chromium)6 7Mounts `ChatMessageAgenticContent` and replays a stream, replacing the message8object on each chunk exactly as the real pipeline does:9 10- `chat.svelte.ts` `updateStreamingUI()` runs per SSE chunk11- `conversations.svelte.ts` `updateMessageAtIndex` does `{ ...old, ...updates }`12 13That new object identity is the thing under test: it cascades through14`deriveAgenticSections` (which returns fresh `AgenticSection` objects) into every15tool-call block in the message, including completed ones.16 17```18npx vitest --project=client --run tests/client/agentic-stream.perf.svelte.test.ts19```20 21### Reading the output22 23- `mean` / `p95` / `max` - the synchronous window per token: prop write,24 `await tick()`, then a forced `offsetHeight` read so style and layout are25 included rather than deferred.26- `sync` - sum of those windows. This is the number to optimize.27- `wall` - the whole run including work `MarkdownContent` defers into its own28 `requestAnimationFrame`. It carries a ~16.7ms/token idle floor because the29 harness yields a frame each iteration, so compare `wall` **across fixtures**,30 never against `sync`.31 32### The knobs, and what each one discriminates33 34The point of the harness is the _scaling curve_, not any single number.35 36| Knob | Reads on |37| --------------------------- | ---------------------------------------------------------------------------------------------------- |38| `priorToolCalls` (0/1/5/20) | the reactive fan-out. Flat => no fan-out. Linear => confirmed. |39| `toolResultBytes` | whole-blob string scans (`extractSearchResults`, `parseToolResultWithImages`, `classifyToolResult`). |40| `editFileEdits` | `computeLineDiff`, the O(m\*n) LCS. |41| `openCodeFence` | `hljs.highlightAuto` on partial code. |42 43Deliberately no hard assertions: CI timing is noisy and the value here is the44before/after delta, not a gate.45 46### Caveat47 48This measures one message's subtree. In the real app `ChatMessages.svelte`49rebuilds its whole `displayMessages` list per token, so multiply by the number50of rendered messages to get the conversation-level cost.51 52## Tier 2 - `../unit/agentic-hotpath.bench.ts` (project: `unit`, node)53 54Per-call costs for the pure functions the curve implicates.55 56```57npx vitest bench --project=unit --run tests/unit/agentic-hotpath.bench.ts58```59 