commitcopilot/infer-001
chore(config): update llama server arguments
chore(config): switch default model to Gemma
perf(config): increase model batch size
docs(readme): remove default Spaces reference
chore(config): update default model artifact
feat(llama): add parallel server configuration
feat(generation): disable model reasoning by default
feat(config): enable prompt caching by default
chore(config): tune model validation and llama defaults
refactor(api): structure app config and lifecycle
feat(api): support Qwen tool calls
fix(api): use configurable readiness wait settings
chore(config): switch default model to Gemma E2B
chore(config): tune llama runtime defaults
feat(config): increase llama context window
fix(api): return requested model in generate response
fix(api): allow non-integer usage values
fix(model): validate downloaded GGUF files
feat(inference): run generation through llama-server
build(docker): add musl runtime dependencies
build(docker): use prebuilt llama.cpp wheels
build(docker): add pkg-config dependency
feat(inference): add cloud inference worker
initial commit
