models
Open weights, fine-tunes and adapters. Every listing here comes live from the Hugging Face Hub, attributed to it, and links back to the source.
TinyLlama-1.1B-Chat-v1.0-kvcache-fp8-tensorTinyLlama-1.1B-Chat-v1.0-kvcache-fp8-attn_headkv_cache_fp8-e2ekv_cache_gptq_tinyllama-e2eTinyLlama-1.1B-compressed-tensors-kv-cache-schemeopt-125m-gqa-ub-6-best-for-KV-cachefacebook-opt-125m-qcqa-ub-6-best-for-KV-cacheQwen3-30BA3B-GGUFpixtral-12b-FP8-dynamic-FP8-KV-cacheKimi-K2-Thinking-CPU-weightfacebook-opt-6.7b-qcqa-ub-16-best-for-KV-cachefacebook-opt-6.7b-gqa-ub-16-best-for-KV-cacheKimi-K2-Instruct-GGUFGLM-4.5-Air-NVFP4-KV-cache-FP8GLM-4.7-NVFP4-KV-cache-BF16Kimi-K2-Instruct-0905-GGUFPhi-3-mini-4k-instruct-kv_cache_default_phi3-e2eGLM-4.7-NVFP4-KV-cache-FP8TinyLlama-1.1B-Chat-v1.0-kv_cache_default_tinyllama-e2eGLM-4.5-Air-NVFP4-KV-cache-BF16sampling_with_kvcache_hf_helpersGLM-4.5-Air-NVFP4-KV-cache-NVFP4TinyLlama-1.1B-Chat-v1.0-kv_cache_default_gptq_tinyllama-e2eLlama-3.1-8B-Instruct-KV-Cache-FP8manga-ocr-kvcache-tflitesampling_with_kvcachesampling_with_kvcacheMeta-Llama-3-8B-Instruct-FP8-channel-output-activation-kv_cache-qkv_projqmd-query-expansion-1.7B-ONNX-kvcache-fp32Llama-3-70B-Instruct-awq-int8-kv-cache-trt-llm
