Model Gallery

109 models from 1 repositories

Filter by type:

Filter by tags:

ternary-bonsai-2-27b
Ternary Bonsai 2 27B (PrismML) is a 27B-class reasoning model with ternary transformer weights. This PTQ1_0 build packs the trits densely at 1.75 bits per weight (5.95 GB) and includes the Q8_0 vision projector. PTQ1_0 is a Prism-private GGUF type, so the entry uses the bonsai backend (PrismML's llama.cpp fork) instead of stock llama.cpp.

Repository: localaiLicense: apache-2.0

swift-qwen3.8-27b
Swift-Qwen3.8-27B is UkisAI's reasoning-efficient fine-tune of Qwen3.8-27B. The publisher reports 58.3% fewer thinking tokens with less than 1% quality loss. This Q4_K_M GGUF includes the F16 vision projector and enables MTP speculative decoding. The weights use the Swift Open License v1.0.

Repository: localaiLicense: swift-open-license-1.0

qwopus3.8-27b-flash-v2
Qwopus3.8-27B-Flash-V2 is a new post-training release for reasoning and agent workloads. This Q4_K_M GGUF includes the F32 vision projector and uses llama.cpp's embedded chat template with MTP speculative decoding.

Repository: localaiLicense: apache-2.0

qwopus3.8-27b-flash-v2-q8
Qwopus3.8-27B-Flash-V2 is a new post-training release for reasoning and agent workloads. This Q8_0 GGUF includes the F32 vision projector and uses llama.cpp's embedded chat template with MTP speculative decoding.

Repository: localaiLicense: apache-2.0

qwopus3.8-27b-flash
Qwopus3.8-27B-Flash is a Qwen3.8-27B fine-tune for reasoning and agent workloads. This Q4_K_M GGUF includes the F32 vision projector and uses llama.cpp's embedded chat template with MTP speculative decoding. The publisher reports a known Python code indentation issue.

Repository: localaiLicense: apache-2.0

qwopus3.8-27b-flash-q8
Qwopus3.8-27B-Flash is a Qwen3.8-27B fine-tune for reasoning and agent workloads. This Q8_0 GGUF includes the F32 vision projector and uses llama.cpp's embedded chat template with MTP speculative decoding. The publisher reports a known Python code indentation issue.

Repository: localaiLicense: apache-2.0

thinkingcap-qwen3.8-27b
ThinkingCap is a 27B Qwen3.8 fine-tune trained to reduce reasoning tokens, with text and image input. This Q4_K_M GGUF build uses llama.cpp, the embedded chat template, and the F16 vision projector. Licensed under PolyForm Small Business 1.0.0 with the publisher's personal-use grant; see the model license for permitted use.

Repository: localaiLicense: polyform-small-business-1.0.0

thinkingcap-qwen3.8-27b-q8
ThinkingCap is a 27B Qwen3.8 fine-tune trained to reduce reasoning tokens, with text and image input. This Q8_0 GGUF build uses llama.cpp, the embedded chat template, and the F16 vision projector. Licensed under PolyForm Small Business 1.0.0 with the publisher's personal-use grant; see the model license for permitted use.

Repository: localaiLicense: polyform-small-business-1.0.0

qwen3.8-27b
# Qwen3.8-27B > [!Note] > This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. > > These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, TokenSpeed, etc. > [!Tip] > For users seeking managed, scalable inference without infrastructure maintenance, the official Qwen API service is provided by Qwen Cloud. > In particular, **Qwen3.8-27B** will be available as a hosted version with more production features, e.g., 1M context length by default, official built-in tools. For more information, please refer to the Qwen3.8-27B Overview. The service is coming soon. Stay tuned for updates. Following the widespread community adoption of the Qwen3.5 and Qwen3.6 series, we are pleased to introduce Qwen3.8, the most capable generation in the Qwen open-model family to date. ...

Repository: localaiLicense: apache-2.0

qwen3.8-27b-uncensored-hauhaucs-aggressive-mtp
# Qwen3.8-27B > [!Note] > This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. > > These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, TokenSpeed, etc. > [!Tip] > For users seeking managed, scalable inference without infrastructure maintenance, the official Qwen API service is provided by Qwen Cloud. > In particular, **Qwen3.8-27B** will be available as a hosted version with more production features, e.g., 1M context length by default, official built-in tools. For more information, please refer to the Qwen3.8-27B Overview. The service is coming soon. Stay tuned for updates. Following the widespread community adoption of the Qwen3.5 and Qwen3.6 series, we are pleased to introduce Qwen3.8, the most capable generation in the Qwen open-model family to date. ...

Repository: localaiLicense: apache-2.0

qwen3.8-27b-uncensored
# Qwen3.8-27B > [!Note] > This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. > > These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, TokenSpeed, etc. > [!Tip] > For users seeking managed, scalable inference without infrastructure maintenance, the official Qwen API service is provided by Qwen Cloud. > In particular, **Qwen3.8-27B** will be available as a hosted version with more production features, e.g., 1M context length by default, official built-in tools. For more information, please refer to the Qwen3.8-27B Overview. The service is coming soon. Stay tuned for updates. Following the widespread community adoption of the Qwen3.5 and Qwen3.6 series, we are pleased to introduce Qwen3.8, the most capable generation in the Qwen open-model family to date. ...

Repository: localaiLicense: apache-2.0

qwen3.8-27b-turbo-fable-cold-fusion-735-882-heretic-uncensored-neo-coder-max-mtp
RELEASE #1 GGUFS [including detailed notes, how to use, benches and much more]: https://huggingface.co/DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NEO-CODER-MAX-MTP-GGUF Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU (release #1, others pending...) ( repo has 10+ other versions (and 3 branches) noted below that EXCEED the performance of all QWEN 27B models, including fine tunes. ) First, special thanks to Nightmedia for working on the first three stages prior to heretic'ing/post staging and benching everything (3 sections below). A number of my finetunes - both released and non-released - were used here as well as some third parties. Full details will be disclosed upon final release as the project shores up. THREE example generations [snippets] from STAGE1-PART2, STAGE1b-PART2 and STAGE2-rplus2 at the bottom of the page. Release(s) will be GGUFS first (linked here directly) then source code shortly thereafter [now released/open]. Some additional work and/ spawning of new branches from branch(es) below is still going on. NEW: Branch 3 added, see below. COMPLETED AND PENDING RELEASES: ...

Repository: localaiLicense: apache-2.0

qwen3.8-flash-next-atomic-q4
Qwen3.8 Flash Next in AtomicChat's AD-4.27bpw Q4_K_M M64 GGUF build, with the F16 vision projector. The n-gram table occupies a separate shard. This entry enables memory mapping and disables llama.cpp automatic parameter fitting as required by the publisher.

Repository: localaiLicense: other

byteshape-qwen3.8-27b
ByteShape's ShapeLearn quantization of Qwen3.8 27B for chat, reasoning, coding, and image input. This IQ4_XS-3.84bpw build uses mixed tensor precisions and includes a BF16 vision projector. MTP decoding is enabled.

Repository: localaiLicense: apache-2.0

byteshape-qwen3.8-27b-iq3-s
ByteShape's ShapeLearn quantization of Qwen3.8 27B for chat, reasoning, coding, and image input. This IQ3_S-3.23bpw build uses mixed tensor precisions and includes a BF16 vision projector. MTP decoding is enabled.

Repository: localaiLicense: apache-2.0

byteshape-qwen3.8-27b-iq3-xs
ByteShape's ShapeLearn quantization of Qwen3.8 27B for chat, reasoning, coding, and image input. This IQ3_XS-3.01bpw build uses mixed tensor precisions and includes a BF16 vision projector. MTP decoding is enabled.

Repository: localaiLicense: apache-2.0

byteshape-qwen3.8-27b-iq3-xxs
ByteShape's ShapeLearn quantization of Qwen3.8 27B for chat, reasoning, coding, and image input. This IQ3_XXS-2.88bpw build uses mixed tensor precisions and includes a BF16 vision projector. MTP decoding is enabled.

Repository: localaiLicense: apache-2.0

byteshape-qwen3.8-27b-iq2-xxs
ByteShape's ShapeLearn quantization of Qwen3.8 27B for chat, reasoning, coding, and image input. This IQ2_XXS-2.56bpw build uses mixed tensor precisions and includes a BF16 vision projector. MTP decoding is enabled.

Repository: localaiLicense: apache-2.0

dirk-qwen3.8-27b-q4
Dirk is a Qwen3.8 27B vision-language model with a concise chat template for agentic coding, reasoning, tool use, and general knowledge tasks. It preserves the model's MTP head for speculative decoding and supports a 262K-token context window. This default entry uses the Q4_K_XL GGUF and F16 vision projector. A choice of Q5_K_XL, Q6_K_XL, and Q8_K_XL builds is available through variants.

Repository: localaiLicense: apache-2.0

dirk-qwen3.8-27b-q8
Dirk in the higher-quality Q8_K_XL GGUF format, with MTP speculative decoding and the shared F16 vision projector for multimodal prompts.

Repository: localaiLicense: apache-2.0

dirk-qwen3.8-27b-q5
Dirk in the higher-quality Q5_K_XL GGUF format, with MTP speculative decoding and the shared F16 vision projector for multimodal prompts.

Repository: localaiLicense: apache-2.0

Page 1