Model Gallery

561 models from 1 repositories

Filter by type:

Filter by tags:

pocket-35b
POCKET-35B is an Apache-2.0 Qwen3.5-family mixture-of-experts model from FINAL-Bench/VIDRAFT, derived from Darwin-36B-Opus and packaged for stock llama.cpp. This entry uses the quality-oriented Q4_K_M GGUF quantization.

Repository: localaiLicense: apache-2.0

pocket-35b-q2
POCKET-35B is an Apache-2.0 Qwen3.5-family mixture-of-experts model from FINAL-Bench/VIDRAFT, derived from Darwin-36B-Opus and packaged for stock llama.cpp. This entry uses the smaller Q2_K GGUF quantization.

Repository: localaiLicense: apache-2.0

pocket-35b-iq1
POCKET-35B is an Apache-2.0 Qwen3.5-family mixture-of-experts model from FINAL-Bench/VIDRAFT, derived from Darwin-36B-Opus and packaged for stock llama.cpp. This entry uses the most compact IQ1_M GGUF quantization.

Repository: localaiLicense: apache-2.0

pocket-26b
POCKET-26B is an Apache-2.0 Gemma 4 26B-A4B mixture-of-experts model from FINAL-Bench/VIDRAFT, tuned for Korean and packaged for stock llama.cpp. This entry uses the quality-oriented Q4_K_M GGUF quantization.

Repository: localaiLicense: apache-2.0

pocket-26b-q2
POCKET-26B is an Apache-2.0 Gemma 4 26B-A4B mixture-of-experts model from FINAL-Bench/VIDRAFT, tuned for Korean and packaged for stock llama.cpp. This entry uses the smaller Q2_K GGUF quantization.

Repository: localaiLicense: apache-2.0

mellum2-12b-a2.5b-instruct
Mellum2-12B-A2.5B-Instruct is an Apache-2.0 mixture-of-experts model from JetBrains with 12 billion total parameters, 2.5 billion activated per token, and a 131,072-token context window. This entry uses the Q4_K_M GGUF quantization.

Repository: localaiLicense: apache-2.0

mellum2-12b-a2.5b-instruct-q8
Mellum2-12B-A2.5B-Instruct is an Apache-2.0 mixture-of-experts model from JetBrains with 12 billion total parameters, 2.5 billion activated per token, and a 131,072-token context window. This entry uses the higher-quality Q8_0 GGUF quantization.

Repository: localaiLicense: apache-2.0

bonsai-8b-1bit
Bonsai 8B (PrismML) is an end-to-end 1-bit language model built on the Qwen3-8B dense architecture (GQA, SwiGLU, RoPE, RMSNorm, 36 layers, 65,536 context). Every weight is a single sign bit (`-scale` / `+scale`) with one FP16 scale per group of 128 weights, for an effective 1.125 bits/weight and a ~1.15 GB footprint (14.2x smaller than FP16) while matching full-precision 8B instruct models at ~70.5 average across 6 benchmark categories. The Q1_0 quantization is only decodable by the PrismML llama.cpp fork, so this entry runs on LocalAI's `bonsai` backend (that fork), not the stock `llama-cpp` backend. License: Apache 2.0.

Repository: localaiLicense: apache-2.0

ternary-bonsai-8b
Ternary Bonsai 8B (PrismML) is a 1.58-bit ternary language model on the Qwen3-8B dense architecture. Each weight takes a value from {-1, 0, +1} with one shared FP16 scale per group of 128 weights (GGUF Q2_0, ~2.18 GB deployed, 7.5x smaller than FP16). The extra zero state recovers more of the full-precision model than the 1-bit build: it ranks 2nd among compared 6-9B models at 75.5 average despite being ~1/8th their size. Q2_0 is the recommended, ternary-lossless variant. The Q2_0 kernels are only in the PrismML llama.cpp fork, so this runs on LocalAI's `bonsai` backend. License: Apache 2.0.

Repository: localaiLicense: apache-2.0

ternary-bonsai-8b-q2-g64
Ternary Bonsai 8B (PrismML), GGUF Q2_0 with group-64 packing (each FP16 scale shared across 64 weights instead of 128). Slightly larger (~2.31 GB) but matches llama.cpp's native 64-value Q2_0 block layout. Runs on LocalAI's `bonsai` backend. License: Apache 2.0.

Repository: localaiLicense: apache-2.0

ternary-bonsai-8b-pq2
Ternary Bonsai 8B (PrismML), GGUF PQ2_0 (packed Q2_0) ternary variant (~2.18 GB). Same {-1, 0, +1} weight alphabet as Q2_0. Runs on LocalAI's `bonsai` backend. License: Apache 2.0.

Repository: localaiLicense: apache-2.0

zero-gemma4-e4b-openzero-q5-k-m
Zero Gemma4-E4B OpenZero is a Gemma 4 E4B instruction fine-tune for local coding, research, chat, and agentic workflows. It was trained on 2,033 curated OpenZero examples and is distributed as one merged Q5_K_M GGUF file. License: OpenZero Community Source v1; see the model repository for terms.

Repository: localaiLicense: other

gemma-4-26b-a4b-it
Google Gemma 4 26B-A4B-IT is an open-source multimodal Mixture-of-Experts model with 26B total parameters and 4B active parameters. It handles text and image input, generating text output, with a 256K context window and support for 140+ languages. The MoE architecture provides strong performance with efficient inference. Well-suited for question answering, summarization, reasoning, and image understanding tasks.

Repository: localaiLicense: apache-2.0

gemma-4-e2b-it
Google Gemma 4 E2B-IT is a lightweight open-source multimodal model with 5B total parameters and 2B effective parameters using selective parameter activation. It handles text and image input, generating text output, with a 256K context window and support for 140+ languages. Optimized for efficient execution on low-resource devices including mobile and laptops.

Repository: localaiLicense: apache-2.0

gemma-4-e4b-it
Google Gemma 4 E4B-IT is an open-source multimodal model with 8B total parameters and 4B effective parameters using selective parameter activation. It handles text and image input, generating text output, with a 256K context window and support for 140+ languages. Offers a good balance of performance and efficiency for deployment on consumer hardware.

Repository: localaiLicense: apache-2.0

gemma-4-31b-it
Google Gemma 4 31B-IT is the largest dense model in the Gemma 4 family with 31B parameters. It handles text and image input, generating text output, with a 256K context window and support for 140+ languages. Provides the highest quality outputs in the Gemma 4 lineup, well-suited for complex reasoning, summarization, and image understanding tasks.

Repository: localaiLicense: apache-2.0

pocket-35b-q4-k-m
POCKET-35B is an Apache-2.0 Qwen3.5-family sparse MoE model derived from Darwin-36B-Opus. Its official GGUF ladder spans Q4_K_M, Q3_K_M, Q2_K, and IQ1_M so LocalAI can select a build for the available memory. All builds run with the stock llama.cpp backend.

Repository: localaiLicense: apache-2.0

pocket-35b-q3-k-m
POCKET-35B is an Apache-2.0 Qwen3.5-family sparse MoE model derived from Darwin-36B-Opus. Its official GGUF ladder spans Q4_K_M, Q3_K_M, Q2_K, and IQ1_M so LocalAI can select a build for the available memory. All builds run with the stock llama.cpp backend.

Repository: localaiLicense: apache-2.0

pocket-35b-q2-k
POCKET-35B is an Apache-2.0 Qwen3.5-family sparse MoE model derived from Darwin-36B-Opus. Its official GGUF ladder spans Q4_K_M, Q3_K_M, Q2_K, and IQ1_M so LocalAI can select a build for the available memory. All builds run with the stock llama.cpp backend.

Repository: localaiLicense: apache-2.0

pocket-35b-iq1-m
POCKET-35B is an Apache-2.0 Qwen3.5-family sparse MoE model derived from Darwin-36B-Opus. Its official GGUF ladder spans Q4_K_M, Q3_K_M, Q2_K, and IQ1_M so LocalAI can select a build for the available memory. All builds run with the stock llama.cpp backend.

Repository: localaiLicense: apache-2.0

fara1.5-27b
Fara1.5-27B is Microsoft's 27B-parameter multimodal computer-use agent for web browsers, fine-tuned from Qwen3.5-27B. It accepts screenshots and text, emits structured browser actions, supports a 262K-token context, and should be deployed with appropriate sandboxing and user-confirmation controls. This entry uses the Q4_K_M GGUF quantization.

Repository: localaiLicense: mit

Page 1