Model Gallery

86 models from 1 repositories

Filter by type:

Filter by tags:

ternary-bonsai-2-27b
Ternary Bonsai 2 27B (PrismML) is a 27B-class reasoning model with ternary transformer weights. This PTQ1_0 build packs the trits densely at 1.75 bits per weight (5.95 GB) and includes the Q8_0 vision projector. PTQ1_0 is a Prism-private GGUF type, so the entry uses the bonsai backend (PrismML's llama.cpp fork) instead of stock llama.cpp.

Repository: localaiLicense: apache-2.0

qwopus3.8-27b-flash-v2
Qwopus3.8-27B-Flash-V2 is a new post-training release for reasoning and agent workloads. This Q4_K_M GGUF includes the F32 vision projector and uses llama.cpp's embedded chat template with MTP speculative decoding.

Repository: localaiLicense: apache-2.0

qwopus3.8-27b-flash-v2-q8
Qwopus3.8-27B-Flash-V2 is a new post-training release for reasoning and agent workloads. This Q8_0 GGUF includes the F32 vision projector and uses llama.cpp's embedded chat template with MTP speculative decoding.

Repository: localaiLicense: apache-2.0

qwopus3.8-27b-flash
Qwopus3.8-27B-Flash is a Qwen3.8-27B fine-tune for reasoning and agent workloads. This Q4_K_M GGUF includes the F32 vision projector and uses llama.cpp's embedded chat template with MTP speculative decoding. The publisher reports a known Python code indentation issue.

Repository: localaiLicense: apache-2.0

qwopus3.8-27b-flash-q8
Qwopus3.8-27B-Flash is a Qwen3.8-27B fine-tune for reasoning and agent workloads. This Q8_0 GGUF includes the F32 vision projector and uses llama.cpp's embedded chat template with MTP speculative decoding. The publisher reports a known Python code indentation issue.

Repository: localaiLicense: apache-2.0

mimo-v2.6-distill-qwen-9b
MiMo-V2.6-Distill-Qwen-9B is Xiaomi MiMo's 9B Qwen3.5 fine-tune for coding, agent tasks, and visual coding. This Q4_K_M GGUF build uses llama.cpp with the model's embedded chat template and includes the F16 vision projector.

Repository: localaiLicense: mit

mimo-v2.6-distill-qwen-9b-q8
MiMo-V2.6-Distill-Qwen-9B is Xiaomi MiMo's 9B Qwen3.5 fine-tune for coding, agent tasks, and visual coding. This Q8_0 GGUF build uses llama.cpp with the model's embedded chat template and includes the F16 vision projector.

Repository: localaiLicense: mit

thinkingcap-qwen3.8-27b
ThinkingCap is a 27B Qwen3.8 fine-tune trained to reduce reasoning tokens, with text and image input. This Q4_K_M GGUF build uses llama.cpp, the embedded chat template, and the F16 vision projector. Licensed under PolyForm Small Business 1.0.0 with the publisher's personal-use grant; see the model license for permitted use.

Repository: localaiLicense: polyform-small-business-1.0.0

thinkingcap-qwen3.8-27b-q8
ThinkingCap is a 27B Qwen3.8 fine-tune trained to reduce reasoning tokens, with text and image input. This Q8_0 GGUF build uses llama.cpp, the embedded chat template, and the F16 vision projector. Licensed under PolyForm Small Business 1.0.0 with the publisher's personal-use grant; see the model license for permitted use.

Repository: localaiLicense: polyform-small-business-1.0.0

hemmingway-1
Hemmingway-1 is Altworld's English-first 27B text model, fine-tuned from Qwen3.8-27B for everyday messages and creative writing. This Q4_K_M GGUF build uses llama.cpp and the model's embedded chat template. Licensed under CC BY-NC 4.0; commercial use requires a separate agreement.

Repository: localaiLicense: cc-by-nc-4.0

hemmingway-1-q8
Hemmingway-1 is Altworld's English-first 27B text model, fine-tuned from Qwen3.8-27B for everyday messages and creative writing. This Q8_0 GGUF build uses llama.cpp and the model's embedded chat template. Licensed under CC BY-NC 4.0; commercial use requires a separate agreement.

Repository: localaiLicense: cc-by-nc-4.0

hy4-preview
# Hy4-preview GGUF Three GGUF builds of Hy4-Preview: https://huggingface.co/tencent/Hy4-preview **Language / 语言:** English · 中文 **Neither file runs on stock llama.cpp.** The `hyv4` architecture is not upstream. Apply the patches in `hy4-preview-patch/` ## English ### 1. What these are **`Hy4-preview-Q4_K_M.gguf`** — a conventional Q4_K_M. Most tensors are Q4_K; `ffn_down_exps` gets Q6_K on 37 layers via llama.cpp's own logic. Use this unless you are memory-constrained. **`Hy4-preview-UD-IQ1_M.gguf`** - mixed precision with UD-IQ1_M strategy at ~2.44 bpw, roughly **half the size** for the same model. The routed-expert `gate`/`up` projections run at 1.75 bpw (IQ1_M) and 2.0625 bpw (IQ2_XXS). **`Hy4-preview-STQ1_0.gguf`** — mixed precision with MIX-STQ1_0 strategy at ~2.38 bpw, roughly **half the size** for the same model. The routed-expert `gate`/`up` projections run at 1.3125 bpw (STQ1_0) on 29 layers and 2.0625 bpw (IQ2_XXS) on the other 48. See section 3. ### 2. Running them Build a patched llama.cpp ```bash git clone https://github.com/ggml-org/llama.cpp && cd llama.cpp git checkout 0cea36222 ...

Repository: localai

huihui-qwen3.8-flash-next-abliterated-q4
Huihui's abliterated Qwen3.8-Flash-Next is a vision-language mixture-of-experts model modified to reduce refusals. This entry uses the publisher's UD-Q4_K_XL GGUF and BF16 vision projector for text chat and image input through llama.cpp. The default context is 32,768 tokens. Model weights use the Qwen Community License 1.0.

Repository: localaiLicense: other

qwen3.8-flash-next-gsq-rco-q2-0
Qwen3.8 Flash Next with ISTA DASLab's GSQ-RCO Q2_0 mixed quantization. Includes both GGUF shards (transformer weights and n-gram embeddings) and the BF16 vision projector for text chat and image input. Uses llama.cpp with the embedded chat template and a 32K context.

Repository: localaiLicense: other

qwen3.8-flash-next-gsq-rco-iq2-xs
Qwen3.8 Flash Next with ISTA DASLab's GSQ-RCO IQ2_XS mixed quantization. Includes both GGUF shards (transformer weights and n-gram embeddings) and the BF16 vision projector for text chat and image input. Uses llama.cpp with the embedded chat template and a 32K context.

Repository: localaiLicense: other

qwen3.8-flash-next-gsq-rco-iq3-xxs
Qwen3.8 Flash Next with ISTA DASLab's GSQ-RCO IQ3_XXS mixed quantization. Includes both GGUF shards (transformer weights and n-gram embeddings) and the BF16 vision projector for text chat and image input. Uses llama.cpp with the embedded chat template and a 32K context.

Repository: localaiLicense: other

qwen3.8-flash-next-atomic-iq4
Qwen3.8 Flash Next in AtomicChat's AD-3.84bpw IQ4_XS M64 GGUF build, with the F16 vision projector. The n-gram table occupies a separate shard. This entry enables memory mapping and disables llama.cpp automatic parameter fitting as required by the publisher.

Repository: localaiLicense: other

qwen3.8-flash-next-atomic-q4
Qwen3.8 Flash Next in AtomicChat's AD-4.27bpw Q4_K_M M64 GGUF build, with the F16 vision projector. The n-gram table occupies a separate shard. This entry enables memory mapping and disables llama.cpp automatic parameter fitting as required by the publisher.

Repository: localaiLicense: other

huihui-qwen3.8-27b-abliterated
Huihui Qwen3.8 27B is an abliterated vision-language model published by huihui-ai. This BF16 GGUF build includes the shared BF16 vision projector and enables MTP speculative decoding through llama.cpp. Q4_K and Q8_0 variants are available as smaller downloads.

Repository: localaiLicense: apache-2.0

huihui-qwen3.8-27b-abliterated-q4
Huihui Qwen3.8 27B in Q4_K GGUF format, with the shared BF16 vision projector and MTP speculative decoding through llama.cpp.

Repository: localaiLicense: apache-2.0

huihui-qwen3.8-27b-abliterated-q8
Huihui Qwen3.8 27B in Q8_0 GGUF format, with the shared BF16 vision projector and MTP speculative decoding through llama.cpp.

Repository: localaiLicense: apache-2.0

Page 1