
Qwopus3.8-27B-Flash-V2 is a new post-training release for reasoning and agent workloads. This Q8_0 GGUF includes the F32 vision projector and uses llama.cpp's embedded chat template with MTP speculative decoding.
Links
Tags
Repository: localaiLicense: apache-2.0

Qwopus3.8-27B-Flash is a Qwen3.8-27B fine-tune for reasoning and agent workloads. This Q8_0 GGUF includes the F32 vision projector and uses llama.cpp's embedded chat template with MTP speculative decoding. The publisher reports a known Python code indentation issue.
Links
Tags
MiMo-V2.6-Distill-Qwen-9B is Xiaomi MiMo's 9B Qwen3.5 fine-tune for coding, agent tasks, and visual coding. This Q8_0 GGUF build uses llama.cpp with the model's embedded chat template and includes the F16 vision projector.
Links
Tags
Repository: localaiLicense: polyform-small-business-1.0.0
ThinkingCap is a 27B Qwen3.8 fine-tune trained to reduce reasoning tokens, with text and image input. This Q8_0 GGUF build uses llama.cpp, the embedded chat template, and the F16 vision projector. Licensed under PolyForm Small Business 1.0.0 with the publisher's personal-use grant; see the model license for permitted use.
Links
Tags
Repository: localaiLicense: cc-by-nc-4.0
Hemmingway-1 is Altworld's English-first 27B text model, fine-tuned from Qwen3.8-27B for everyday messages and creative writing. This Q8_0 GGUF build uses llama.cpp and the model's embedded chat template. Licensed under CC BY-NC 4.0; commercial use requires a separate agreement.
Links
Tags
Apodex-1.1-mini in the higher-quality Q8_0 GGUF format, with the shared F16 vision projector for multimodal prompts.
Links
Tags
Repository: localaiLicense: mit
GLM-5.3-Flash is Z.ai's natively multimodal 320B-parameter mixture-of-experts model with 18B active parameters. It combines sparse and linear attention for coding, agentic work, tool use, vision, and long-context tasks. This entry uses the higher-quality Q8_0 GGUF quantization and enables the model's MTP speculative-decoding head.
Links
Tags

Qwen3.8-Flash-Next in the higher-quality Q8_0 GGUF format, with the shared BF16 vision projector. This build preserves more model quality but needs more memory than the default Q4 variant.
Links
Tags
Ling-3.0-tiny in the higher-quality Q8_0 GGUF format. This variant preserves more model fidelity for hosts with enough memory.
Links
Tags
IBM Granite 4.2 3B in the higher-fidelity Q8_0 GGUF format. It is a compact multilingual reasoning model for chat, coding, and tool use.
Links
Tags
IBM Granite 4.2 8B in the higher-fidelity Q8_0 GGUF format. It is a multilingual reasoning model for chat, coding, and tool use.
Links
Tags
IBM Granite 4.2 30B in the higher-fidelity Q8_0 GGUF format. It is the family's flagship multilingual reasoning model for chat, coding, and tool use.
Links
Tags
Dirk in the higher-quality Q8_K_XL GGUF format, with MTP speculative decoding and the shared F16 vision projector for multimodal prompts.
Links
Tags
Qwen3.8-27B-Uncensored in the higher-quality Q8_0 GGUF format, with its integrated MTP head and shared F16 vision projector.
Links
Tags

Huihui Qwen3.8 27B in Q8_0 GGUF format, with the shared BF16 vision projector and MTP speculative decoding through llama.cpp.
Links
Tags
Repository: localaiLicense: apache-2.0
Hy-MT2-7B is Tencent's 7B multilingual translation model. It supports translation instructions across 33 languages, including terminology control and style transfer. This Q8_0 GGUF uses the embedded chat template with an 8K context window. Include the target language in your prompt.
Links
Tags
Hy-MT2-1.8B in the higher-quality 1.9 GB Q8_0 GGUF format. This variant preserves more model fidelity for hosts with enough memory.
Links
Tags
Repository: localaiLicense: apache-2.0
Qwen3.8-27B Uncensored Cyber is philbert440's security-focused derivative of Qwen3.8-27B. This build uses the original publisher's Q8_0 weights. Includes the BF16 vision projector, embedded Jinja chat template, and a 32K-token default context. Speculative decoding is not enabled.
Links
Tags
Carbon-3B in the higher-quality Q8_0 GGUF format for genomic sequence generation and analysis.
Links
Tags
Carbon-8B in the higher-quality Q8_0 GGUF format for genomic sequence generation and analysis.
Links
Tags
Ornith-1.0-9B in the higher-quality Q8_0 GGUF format, with the shared F16 vision projector for multimodal prompts.
Links
Tags