Model Gallery

73 models from 1 repositories

Filter by type:

Filter by tags:

swift-qwen3.8-27b
Swift-Qwen3.8-27B is UkisAI's reasoning-efficient fine-tune of Qwen3.8-27B. The publisher reports 58.3% fewer thinking tokens with less than 1% quality loss. This Q4_K_M GGUF includes the F16 vision projector and enables MTP speculative decoding. The weights use the Swift Open License v1.0.

Repository: localaiLicense: swift-open-license-1.0

qwen3.8-flash-next-uncensored
# Qwen3.8-Flash-Next > [!Note] > This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. > > These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, TokenSpeed, etc. > [!Tip] > For users seeking managed, scalable inference without infrastructure maintenance, the official Qwen API service is provided by Qwen Cloud. > > In particular, **Qwen3.8-Flash** is the official version based on Qwen3.8-Flash-Next with more production features, e.g., 1M context length by default, official built-in tools. For more information, please refer to the Qwen3.8-Flash Overview. As the frontier of foundation models pushes toward ever-larger parameter counts and ever-longer context windows, the question is no longer just how much we can scale, but how efficiently we can do so. Sustainable progress toward artificial general intelligence (AGI) that benefits everyone demands architectural innovation. Today, we are sharing a concrete step in that direction: Qwen3.8-Flash-Next. ...

Repository: localaiLicense: apache-2.0

thinkingcap-qwen3.8-27b
ThinkingCap is a 27B Qwen3.8 fine-tune trained to reduce reasoning tokens, with text and image input. This Q4_K_M GGUF build uses llama.cpp, the embedded chat template, and the F16 vision projector. Licensed under PolyForm Small Business 1.0.0 with the publisher's personal-use grant; see the model license for permitted use.

Repository: localaiLicense: polyform-small-business-1.0.0

thinkingcap-qwen3.8-27b-q8
ThinkingCap is a 27B Qwen3.8 fine-tune trained to reduce reasoning tokens, with text and image input. This Q8_0 GGUF build uses llama.cpp, the embedded chat template, and the F16 vision projector. Licensed under PolyForm Small Business 1.0.0 with the publisher's personal-use grant; see the model license for permitted use.

Repository: localaiLicense: polyform-small-business-1.0.0

qwen3.8-27b
# Qwen3.8-27B > [!Note] > This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. > > These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, TokenSpeed, etc. > [!Tip] > For users seeking managed, scalable inference without infrastructure maintenance, the official Qwen API service is provided by Qwen Cloud. > In particular, **Qwen3.8-27B** will be available as a hosted version with more production features, e.g., 1M context length by default, official built-in tools. For more information, please refer to the Qwen3.8-27B Overview. The service is coming soon. Stay tuned for updates. Following the widespread community adoption of the Qwen3.5 and Qwen3.6 series, we are pleased to introduce Qwen3.8, the most capable generation in the Qwen open-model family to date. ...

Repository: localaiLicense: apache-2.0

qwen3.8-27b-uncensored-hauhaucs-aggressive-mtp
# Qwen3.8-27B > [!Note] > This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. > > These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, TokenSpeed, etc. > [!Tip] > For users seeking managed, scalable inference without infrastructure maintenance, the official Qwen API service is provided by Qwen Cloud. > In particular, **Qwen3.8-27B** will be available as a hosted version with more production features, e.g., 1M context length by default, official built-in tools. For more information, please refer to the Qwen3.8-27B Overview. The service is coming soon. Stay tuned for updates. Following the widespread community adoption of the Qwen3.5 and Qwen3.6 series, we are pleased to introduce Qwen3.8, the most capable generation in the Qwen open-model family to date. ...

Repository: localaiLicense: apache-2.0

qwen3.8-27b-uncensored
# Qwen3.8-27B > [!Note] > This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. > > These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, TokenSpeed, etc. > [!Tip] > For users seeking managed, scalable inference without infrastructure maintenance, the official Qwen API service is provided by Qwen Cloud. > In particular, **Qwen3.8-27B** will be available as a hosted version with more production features, e.g., 1M context length by default, official built-in tools. For more information, please refer to the Qwen3.8-27B Overview. The service is coming soon. Stay tuned for updates. Following the widespread community adoption of the Qwen3.5 and Qwen3.6 series, we are pleased to introduce Qwen3.8, the most capable generation in the Qwen open-model family to date. ...

Repository: localaiLicense: apache-2.0

qwen3.8-27b-turbo-fable-cold-fusion-735-882-heretic-uncensored-neo-coder-max-mtp
RELEASE #1 GGUFS [including detailed notes, how to use, benches and much more]: https://huggingface.co/DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NEO-CODER-MAX-MTP-GGUF Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU (release #1, others pending...) ( repo has 10+ other versions (and 3 branches) noted below that EXCEED the performance of all QWEN 27B models, including fine tunes. ) First, special thanks to Nightmedia for working on the first three stages prior to heretic'ing/post staging and benching everything (3 sections below). A number of my finetunes - both released and non-released - were used here as well as some third parties. Full details will be disclosed upon final release as the project shores up. THREE example generations [snippets] from STAGE1-PART2, STAGE1b-PART2 and STAGE2-rplus2 at the bottom of the page. Release(s) will be GGUFS first (linked here directly) then source code shortly thereafter [now released/open]. Some additional work and/ spawning of new branches from branch(es) below is still going on. NEW: Branch 3 added, see below. COMPLETED AND PENDING RELEASES: ...

Repository: localaiLicense: apache-2.0

huihui-qwen3.8-flash-next-abliterated-q4
Huihui's abliterated Qwen3.8-Flash-Next is a vision-language mixture-of-experts model modified to reduce refusals. This entry uses the publisher's UD-Q4_K_XL GGUF and BF16 vision projector for text chat and image input through llama.cpp. The default context is 32,768 tokens. Model weights use the Qwen Community License 1.0.

Repository: localaiLicense: other

qwen3.8-flash-next-q4
Qwen3.8-Flash-Next is Qwen's 125B-parameter, 6B-active experimental vision-language mixture-of-experts model. It targets agentic coding, reasoning, tool use, and long-context workloads with a native 262K-token context window. This default entry uses Unsloth's UD-Q4_K_XL GGUF and BF16 vision projector. Linked variants offer Q8_0 and AtomicChat's smaller IQ4_XS and Q4_K_M builds with a separate n-gram table shard.

Repository: localaiLicense: other

qwen3.8-flash-next-q8
Qwen3.8-Flash-Next in the higher-quality Q8_0 GGUF format, with the shared BF16 vision projector. This build preserves more model quality but needs more memory than the default Q4 variant.

Repository: localaiLicense: other

qwen3.8-flash-next-gsq-rco-q2-0
Qwen3.8 Flash Next with ISTA DASLab's GSQ-RCO Q2_0 mixed quantization. Includes both GGUF shards (transformer weights and n-gram embeddings) and the BF16 vision projector for text chat and image input. Uses llama.cpp with the embedded chat template and a 32K context.

Repository: localaiLicense: other

qwen3.8-flash-next-gsq-rco-iq2-xs
Qwen3.8 Flash Next with ISTA DASLab's GSQ-RCO IQ2_XS mixed quantization. Includes both GGUF shards (transformer weights and n-gram embeddings) and the BF16 vision projector for text chat and image input. Uses llama.cpp with the embedded chat template and a 32K context.

Repository: localaiLicense: other

qwen3.8-flash-next-gsq-rco-iq3-xxs
Qwen3.8 Flash Next with ISTA DASLab's GSQ-RCO IQ3_XXS mixed quantization. Includes both GGUF shards (transformer weights and n-gram embeddings) and the BF16 vision projector for text chat and image input. Uses llama.cpp with the embedded chat template and a 32K context.

Repository: localaiLicense: other

qwen3.8-flash-next-atomic-iq4
Qwen3.8 Flash Next in AtomicChat's AD-3.84bpw IQ4_XS M64 GGUF build, with the F16 vision projector. The n-gram table occupies a separate shard. This entry enables memory mapping and disables llama.cpp automatic parameter fitting as required by the publisher.

Repository: localaiLicense: other

qwen3.8-flash-next-atomic-q4
Qwen3.8 Flash Next in AtomicChat's AD-4.27bpw Q4_K_M M64 GGUF build, with the F16 vision projector. The n-gram table occupies a separate shard. This entry enables memory mapping and disables llama.cpp automatic parameter fitting as required by the publisher.

Repository: localaiLicense: other

byteshape-qwen3.8-27b
ByteShape's ShapeLearn quantization of Qwen3.8 27B for chat, reasoning, coding, and image input. This IQ4_XS-3.84bpw build uses mixed tensor precisions and includes a BF16 vision projector. MTP decoding is enabled.

Repository: localaiLicense: apache-2.0

byteshape-qwen3.8-27b-iq3-s
ByteShape's ShapeLearn quantization of Qwen3.8 27B for chat, reasoning, coding, and image input. This IQ3_S-3.23bpw build uses mixed tensor precisions and includes a BF16 vision projector. MTP decoding is enabled.

Repository: localaiLicense: apache-2.0

byteshape-qwen3.8-27b-iq3-xs
ByteShape's ShapeLearn quantization of Qwen3.8 27B for chat, reasoning, coding, and image input. This IQ3_XS-3.01bpw build uses mixed tensor precisions and includes a BF16 vision projector. MTP decoding is enabled.

Repository: localaiLicense: apache-2.0

byteshape-qwen3.8-27b-iq3-xxs
ByteShape's ShapeLearn quantization of Qwen3.8 27B for chat, reasoning, coding, and image input. This IQ3_XXS-2.88bpw build uses mixed tensor precisions and includes a BF16 vision projector. MTP decoding is enabled.

Repository: localaiLicense: apache-2.0

byteshape-qwen3.8-27b-iq2-xxs
ByteShape's ShapeLearn quantization of Qwen3.8 27B for chat, reasoning, coding, and image input. This IQ2_XXS-2.56bpw build uses mixed tensor precisions and includes a BF16 vision projector. MTP decoding is enabled.

Repository: localaiLicense: apache-2.0

Page 1