Model Gallery

817 models from 1 repositories

Filter by type:

Filter by tags:

mimo-v2.6-distill-qwen-9b
MiMo-V2.6-Distill-Qwen-9B is Xiaomi MiMo's 9B Qwen3.5 fine-tune for coding, agent tasks, and visual coding. This Q4_K_M GGUF build uses llama.cpp with the model's embedded chat template and includes the F16 vision projector.

Repository: localaiLicense: mit

mimo-v2.6-distill-qwen-9b-q8
MiMo-V2.6-Distill-Qwen-9B is Xiaomi MiMo's 9B Qwen3.5 fine-tune for coding, agent tasks, and visual coding. This Q8_0 GGUF build uses llama.cpp with the model's embedded chat template and includes the F16 vision projector.

Repository: localaiLicense: mit

thinkingcap-qwen3.8-27b
ThinkingCap is a 27B Qwen3.8 fine-tune trained to reduce reasoning tokens, with text and image input. This Q4_K_M GGUF build uses llama.cpp, the embedded chat template, and the F16 vision projector. Licensed under PolyForm Small Business 1.0.0 with the publisher's personal-use grant; see the model license for permitted use.

Repository: localaiLicense: polyform-small-business-1.0.0

thinkingcap-qwen3.8-27b-q8
ThinkingCap is a 27B Qwen3.8 fine-tune trained to reduce reasoning tokens, with text and image input. This Q8_0 GGUF build uses llama.cpp, the embedded chat template, and the F16 vision projector. Licensed under PolyForm Small Business 1.0.0 with the publisher's personal-use grant; see the model license for permitted use.

Repository: localaiLicense: polyform-small-business-1.0.0

hemmingway-1
Hemmingway-1 is Altworld's English-first 27B text model, fine-tuned from Qwen3.8-27B for everyday messages and creative writing. This Q4_K_M GGUF build uses llama.cpp and the model's embedded chat template. Licensed under CC BY-NC 4.0; commercial use requires a separate agreement.

Repository: localaiLicense: cc-by-nc-4.0

hemmingway-1-q8
Hemmingway-1 is Altworld's English-first 27B text model, fine-tuned from Qwen3.8-27B for everyday messages and creative writing. This Q8_0 GGUF build uses llama.cpp and the model's embedded chat template. Licensed under CC BY-NC 4.0; commercial use requires a separate agreement.

Repository: localaiLicense: cc-by-nc-4.0

apodex-1.1-mini-q4
Apodex-1.1-mini is an Apache-2.0 Qwen3.5 mixture-of-experts model for long-horizon research, data analysis, coding, file work, and tool use. It activates about 3B of its 35.95B parameters per token and supports text and image input with a context window of 262K tokens. This default entry uses the recommended Q4_K_M GGUF and F16 vision projector. An MTP-enabled build and a higher-quality Q8_0 model are available as variants.

Repository: localaiLicense: apache-2.0

apodex-1.1-mini-q4-mtp
Apodex-1.1-mini with MTP speculative decoding enabled on the recommended Q4_K_M GGUF. The model carries its native MTP head, so it needs no separate draft model. The F16 vision projector supports multimodal prompts.

Repository: localaiLicense: apache-2.0

apodex-1.1-mini-q8
Apodex-1.1-mini in the higher-quality Q8_0 GGUF format, with the shared F16 vision projector for multimodal prompts.

Repository: localaiLicense: apache-2.0

glm-5.3-flash-q4
GLM-5.3-Flash is Z.ai's natively multimodal 320B-parameter mixture-of-experts model with 18B active parameters. It combines sparse and linear attention for coding, agentic work, tool use, vision, and long-context tasks. This entry uses the UD-Q4_K_XL GGUF quantization and enables the model's MTP speculative-decoding head.

Repository: localaiLicense: mit

glm-5.3-flash-q8
GLM-5.3-Flash is Z.ai's natively multimodal 320B-parameter mixture-of-experts model with 18B active parameters. It combines sparse and linear attention for coding, agentic work, tool use, vision, and long-context tasks. This entry uses the higher-quality Q8_0 GGUF quantization and enables the model's MTP speculative-decoding head.

Repository: localaiLicense: mit

nl2sh-1.5b-q4
nl2sh-1.5b is a 1.5B Qwen2.5-Coder fine-tune that converts plain-English requests into single POSIX or Bash commands. This Q4_K_M GGUF is 941 MB and is designed for fast CPU inference. Use the system prompt from the model card and review every generated command before execution. The model can produce destructive commands and cannot inspect the local filesystem.

Repository: localaiLicense: apache-2.0

s1-mini-q4
S1-mini by Superwhisper is a 0.6B English text normalizer for raw speech transcripts. It removes fillers and false starts, restores punctuation and capitalization, and formats spoken numbers, dates, currency, and email addresses as written text. This default entry uses the publisher's 462 MB Q4_K_M GGUF and greedy decoding. A higher-fidelity F16 model is available as a variant. Prefix the transcript with the styling, structure, and context control line documented on the model page.

Repository: localaiLicense: s1-mini-license

s1-mini-f16
S1-mini by Superwhisper in the publisher's 1.4 GB F16 GGUF format. This variant preserves full model fidelity for hosts with enough memory.

Repository: localaiLicense: s1-mini-license

supra2-100m-instruct
Supra2-100M-Instruct is a compact English chat model trained from scratch by SupraLabs on the Qwen3 architecture. It has 100 million parameters, a 2,048-token context window, and is intended for lightweight experiments and constrained edge deployments. This entry uses the publisher's official F16 GGUF build.

Repository: localaiLicense: apache-2.0

llm-jp-4-33b-thinking-q4
LLM-jp-4-33B-thinking is an Apache-2.0 Japanese and English reasoning model from Japan's National Institute of Informatics. Its dense Llama architecture has 33 billion parameters and a 65K-token context window. The model was aligned with supervised fine-tuning and DPO for multi-turn conversation and instruction following. This default entry uses the 20.2 GB Q4_K_M GGUF. The official 66.4 GB BF16 weights are available as a higher-fidelity variant.

Repository: localaiLicense: apache-2.0

llm-jp-4-33b-thinking-bf16
LLM-jp-4-33B-thinking in the official 66.4 GB BF16 GGUF format. This variant preserves the original model precision for hosts with enough memory.

Repository: localaiLicense: apache-2.0

huihui-qwen3.8-flash-next-abliterated-q4
Huihui's abliterated Qwen3.8-Flash-Next is a vision-language mixture-of-experts model modified to reduce refusals. This entry uses the publisher's UD-Q4_K_XL GGUF and BF16 vision projector for text chat and image input through llama.cpp. The default context is 32,768 tokens. Model weights use the Qwen Community License 1.0.

Repository: localaiLicense: other

qwen3.8-flash-next-q4
Qwen3.8-Flash-Next is Qwen's 125B-parameter, 6B-active experimental vision-language mixture-of-experts model. It targets agentic coding, reasoning, tool use, and long-context workloads with a native 262K-token context window. This default entry uses Unsloth's UD-Q4_K_XL GGUF and BF16 vision projector. Linked variants offer Q8_0 and AtomicChat's smaller IQ4_XS and Q4_K_M builds with a separate n-gram table shard.

Repository: localaiLicense: other

qwen3.8-flash-next-q8
Qwen3.8-Flash-Next in the higher-quality Q8_0 GGUF format, with the shared BF16 vision projector. This build preserves more model quality but needs more memory than the default Q4 variant.

Repository: localaiLicense: other

qwen3.8-flash-next-gsq-rco-q2-0
Qwen3.8 Flash Next with ISTA DASLab's GSQ-RCO Q2_0 mixed quantization. Includes both GGUF shards (transformer weights and n-gram embeddings) and the BF16 vision projector for text chat and image input. Uses llama.cpp with the embedded chat template and a 32K context.

Repository: localaiLicense: other

Page 1