Model Gallery

193 models from 1 repositories

Filter by type:

Filter by tags:

nl2sh-1.5b-q4
nl2sh-1.5b is a 1.5B Qwen2.5-Coder fine-tune that converts plain-English requests into single POSIX or Bash commands. This Q4_K_M GGUF is 941 MB and is designed for fast CPU inference. Use the system prompt from the model card and review every generated command before execution. The model can produce destructive commands and cannot inspect the local filesystem.

Repository: localaiLicense: apache-2.0

ornith-1.5-35b-a3b-apex
Ornith-1.5-35B-A3B is an MIT-licensed Qwen3.5 mixture-of-experts model from Ornith AI for agentic coding, reasoning, repository-level software tasks, and tool use. It supports text and image input with a context window of 262K tokens. This default entry uses the APEX Balanced GGUF and BF16 vision projector. Compact APEX and MTP-enabled APEX builds are available as variants.

Repository: localaiLicense: mit

ornith-1.5-35b-a3b-apex-compact
Ornith-1.5-35B-A3B in the smaller APEX Compact GGUF format, with the shared BF16 vision projector for multimodal prompts.

Repository: localaiLicense: mit

ornith-1.5-35b-a3b-mtp-apex
Ornith-1.5-35B-A3B in the APEX Balanced GGUF format with native multi-token prediction enabled for speculative decoding, plus the shared BF16 vision projector.

Repository: localaiLicense: mit

ornith-1.5-35b-a3b-mtp-apex-compact
Ornith-1.5-35B-A3B in the APEX Compact GGUF format with native multi-token prediction enabled for speculative decoding, plus the shared BF16 vision projector.

Repository: localaiLicense: mit

ornith-1.5-35b-a3b-q4
Ornith-1.5-35B-A3B is an MIT-licensed Qwen3.5 mixture-of-experts model from Ornith AI for agentic coding, reasoning, repository-level software tasks, and tool use. It activates about 3B parameters per token and supports text and image input with a context window of 262K tokens. This default entry uses the Q4_K_M GGUF and BF16 vision projector. Q5_K_M, Q6_K, and Q8_0 builds are available as variants.

Repository: localaiLicense: mit

ornith-1.5-35b-a3b-q8
Ornith-1.5-35B-A3B in the higher-quality Q8_0 GGUF format, with the shared BF16 vision projector for multimodal prompts.

Repository: localaiLicense: mit

ornith-1.5-35b-a3b-q5
Ornith-1.5-35B-A3B in the Q5_K_M GGUF format, with the shared BF16 vision projector for multimodal prompts.

Repository: localaiLicense: mit

ornith-1.5-35b-a3b-q6
Ornith-1.5-35B-A3B in the Q6_K GGUF format, with the shared BF16 vision projector for multimodal prompts.

Repository: localaiLicense: mit

cyber-tiel-coder-35b-a3b-q4-mtp
Cyber-Tiel-Coder is a 35B mixture-of-experts coding model with 3B active parameters, based on Huihui's abliterated Ornith-1.5. This UD-Q4_K_XL build includes MTP speculative decoding, the embedded Sharp chat template, and a BF16 vision projector.

Repository: localaiLicense: mit

cyber-tiel-coder-35b-a3b-q8-mtp
Cyber-Tiel-Coder is a 35B mixture-of-experts coding model with 3B active parameters, based on Huihui's abliterated Ornith-1.5. This UD-Q8_K_XL build includes MTP speculative decoding, the embedded Sharp chat template, and a BF16 vision projector.

Repository: localaiLicense: mit

tiel-coder-35b-a3b-q4
Tiel-Coder-35B-A3B is a 35B-parameter mixture-of-experts model for coding, reasoning, tool use, and vision tasks. This default entry uses the Q4_K_XL GGUF and BF16 vision projector.

Repository: localaiLicense: mit

tiel-coder-35b-a3b-q4-mtp
Tiel-Coder-35B-A3B in Q4_K_XL format with MTP speculative decoding and a BF16 vision projector.

Repository: localaiLicense: mit

tiel-coder-35b-a3b-q5-mtp
Tiel-Coder-35B-A3B in Q5_K_XL format with MTP speculative decoding and a BF16 vision projector.

Repository: localaiLicense: mit

tiel-coder-35b-a3b-q6-mtp
Tiel-Coder-35B-A3B in Q6_K_XL format with MTP speculative decoding and a BF16 vision projector.

Repository: localaiLicense: mit

tiel-coder-35b-a3b-q8-mtp
Tiel-Coder-35B-A3B in Q8_K_XL format with MTP speculative decoding and a BF16 vision projector.

Repository: localaiLicense: mit

tiel-coder-35b-a3b-q8
Tiel-Coder-35B-A3B in the higher-quality Q8_K_XL GGUF format, with the BF16 vision projector for multimodal prompts.

Repository: localaiLicense: mit

qwen3.8-35b-a3b-distill-q4
Empero's Qwen3.8 distillation into Qwen3.6-35B-A3B has 35B total parameters and about 3B active per token. This Q4_K_M GGUF build uses the embedded reasoning template and includes the F16 vision projector. Vision is inherited from the base and was not evaluated by the publisher.

Repository: localaiLicense: apache-2.0

qwen3.8-35b-a3b-distill-q5
Empero's Qwen3.8 distillation into Qwen3.6-35B-A3B has 35B total parameters and about 3B active per token. This Q5_K_M GGUF build uses the embedded reasoning template and includes the F16 vision projector. Vision is inherited from the base and was not evaluated by the publisher.

Repository: localaiLicense: apache-2.0

qwen3.8-35b-a3b-distill-q8
Empero's Qwen3.8 distillation into Qwen3.6-35B-A3B has 35B total parameters and about 3B active per token. This Q8_0 GGUF build uses the embedded reasoning template and includes the F16 vision projector. Vision is inherited from the base and was not evaluated by the publisher.

Repository: localaiLicense: apache-2.0

pocket-35b
POCKET-35B is an Apache-2.0 Qwen3.5-family mixture-of-experts model from FINAL-Bench/VIDRAFT, derived from Darwin-36B-Opus and packaged for stock llama.cpp. This entry uses the quality-oriented Q4_K_M GGUF quantization.

Repository: localaiLicense: apache-2.0

Page 1