Model Gallery

641 models from 1 repositories

Filter by type:

Filter by tags:

spark-x2.5-4b
# Spark-X2.5 [](https://join.slack.com/t/tokenspark/shared_invite/zt-432qf8l2f-5~dLyXv8uETr0P0UuC07nw) [](https://discord.gg/kTDE2Hg8aw) [](https://www.youtube.com/@SparkLLM) [](https://dev.to/sparkllm) [](https://bsky.app/profile/sparkllm.bsky.social) [](https://x.com/sparkllm) [](https://www.zhihu.com/people/zhiikz7qh7m) [](images/xhtoken-wechat.jpg) > [!Note] > This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. ## Introduction We are introducing Spark-X2.5-4B and Spark-X2.5-1.7B, two compact, general-purpose language models designed to make capable AI more practical, efficient, and accessible. The models deliver strong performance across a broad range of everyday tasks—including conversation, writing, translation, reasoning, coding, tool use, and agentic workflows—achieving leading results among open-source models of comparable size. Spark-X2.5 combines an efficiency-oriented architecture with native context windows of up to 1M tokens, and support for more than 200 languages. ...

Repository: localaiLicense: apache-2.0

qwopus3.8-27b-flash-v2
Qwopus3.8-27B-Flash-V2 is a new post-training release for reasoning and agent workloads. This Q4_K_M GGUF includes the F32 vision projector and uses llama.cpp's embedded chat template with MTP speculative decoding.

Repository: localaiLicense: apache-2.0

qwopus3.8-27b-flash-v2-q8
Qwopus3.8-27B-Flash-V2 is a new post-training release for reasoning and agent workloads. This Q8_0 GGUF includes the F32 vision projector and uses llama.cpp's embedded chat template with MTP speculative decoding.

Repository: localaiLicense: apache-2.0

qwopus3.8-27b-flash
Qwopus3.8-27B-Flash is a Qwen3.8-27B fine-tune for reasoning and agent workloads. This Q4_K_M GGUF includes the F32 vision projector and uses llama.cpp's embedded chat template with MTP speculative decoding. The publisher reports a known Python code indentation issue.

Repository: localaiLicense: apache-2.0

qwopus3.8-27b-flash-q8
Qwopus3.8-27B-Flash is a Qwen3.8-27B fine-tune for reasoning and agent workloads. This Q8_0 GGUF includes the F32 vision projector and uses llama.cpp's embedded chat template with MTP speculative decoding. The publisher reports a known Python code indentation issue.

Repository: localaiLicense: apache-2.0

mimo-v2.6-distill-qwen-9b
MiMo-V2.6-Distill-Qwen-9B is Xiaomi MiMo's 9B Qwen3.5 fine-tune for coding, agent tasks, and visual coding. This Q4_K_M GGUF build uses llama.cpp with the model's embedded chat template and includes the F16 vision projector.

Repository: localaiLicense: mit

mimo-v2.6-distill-qwen-9b-q8
MiMo-V2.6-Distill-Qwen-9B is Xiaomi MiMo's 9B Qwen3.5 fine-tune for coding, agent tasks, and visual coding. This Q8_0 GGUF build uses llama.cpp with the model's embedded chat template and includes the F16 vision projector.

Repository: localaiLicense: mit

thinkingcap-qwen3.8-27b
ThinkingCap is a 27B Qwen3.8 fine-tune trained to reduce reasoning tokens, with text and image input. This Q4_K_M GGUF build uses llama.cpp, the embedded chat template, and the F16 vision projector. Licensed under PolyForm Small Business 1.0.0 with the publisher's personal-use grant; see the model license for permitted use.

Repository: localaiLicense: polyform-small-business-1.0.0

thinkingcap-qwen3.8-27b-q8
ThinkingCap is a 27B Qwen3.8 fine-tune trained to reduce reasoning tokens, with text and image input. This Q8_0 GGUF build uses llama.cpp, the embedded chat template, and the F16 vision projector. Licensed under PolyForm Small Business 1.0.0 with the publisher's personal-use grant; see the model license for permitted use.

Repository: localaiLicense: polyform-small-business-1.0.0

hemmingway-1
Hemmingway-1 is Altworld's English-first 27B text model, fine-tuned from Qwen3.8-27B for everyday messages and creative writing. This Q4_K_M GGUF build uses llama.cpp and the model's embedded chat template. Licensed under CC BY-NC 4.0; commercial use requires a separate agreement.

Repository: localaiLicense: cc-by-nc-4.0

hemmingway-1-q8
Hemmingway-1 is Altworld's English-first 27B text model, fine-tuned from Qwen3.8-27B for everyday messages and creative writing. This Q8_0 GGUF build uses llama.cpp and the model's embedded chat template. Licensed under CC BY-NC 4.0; commercial use requires a separate agreement.

Repository: localaiLicense: cc-by-nc-4.0

supra2-100m-instruct
Supra2-100M-Instruct is a compact English chat model trained from scratch by SupraLabs on the Qwen3 architecture. It has 100 million parameters, a 2,048-token context window, and is intended for lightweight experiments and constrained edge deployments. This entry uses the publisher's official F16 GGUF build.

Repository: localaiLicense: apache-2.0

huihui-qwen3.8-flash-next-abliterated-q4
Huihui's abliterated Qwen3.8-Flash-Next is a vision-language mixture-of-experts model modified to reduce refusals. This entry uses the publisher's UD-Q4_K_XL GGUF and BF16 vision projector for text chat and image input through llama.cpp. The default context is 32,768 tokens. Model weights use the Qwen Community License 1.0.

Repository: localaiLicense: other

qwen3.8-flash-next-gsq-rco-q2-0
Qwen3.8 Flash Next with ISTA DASLab's GSQ-RCO Q2_0 mixed quantization. Includes both GGUF shards (transformer weights and n-gram embeddings) and the BF16 vision projector for text chat and image input. Uses llama.cpp with the embedded chat template and a 32K context.

Repository: localaiLicense: other

qwen3.8-flash-next-gsq-rco-iq2-xs
Qwen3.8 Flash Next with ISTA DASLab's GSQ-RCO IQ2_XS mixed quantization. Includes both GGUF shards (transformer weights and n-gram embeddings) and the BF16 vision projector for text chat and image input. Uses llama.cpp with the embedded chat template and a 32K context.

Repository: localaiLicense: other

qwen3.8-flash-next-gsq-rco-iq3-xxs
Qwen3.8 Flash Next with ISTA DASLab's GSQ-RCO IQ3_XXS mixed quantization. Includes both GGUF shards (transformer weights and n-gram embeddings) and the BF16 vision projector for text chat and image input. Uses llama.cpp with the embedded chat template and a 32K context.

Repository: localaiLicense: other

dfm-mimir:vllm
DFM Mimir is an Apache-2.0, instruction-tuned HRM-Text model from Danish Foundation Models. It has about 1 billion parameters and a 4,096-token context window. The model focuses on Danish and English chat, reasoning, mathematics, and code generation, and uses only permissible post-training data. This entry serves the official BF16 safetensors checkpoint with vLLM.

Repository: localaiLicense: apache-2.0

granite-4.2-3b-q4
IBM Granite 4.2 3B is a compact multilingual reasoning model for chat, coding, long-context tasks, and tool use. This entry uses the Q4_K_M GGUF; a higher-fidelity Q8_0 build is available as a variant.

Repository: localaiLicense: apache-2.0

granite-4.2-3b-q8
IBM Granite 4.2 3B in the higher-fidelity Q8_0 GGUF format. It is a compact multilingual reasoning model for chat, coding, and tool use.

Repository: localaiLicense: apache-2.0

granite-4.2-8b-q4
IBM Granite 4.2 8B is a multilingual reasoning model for chat, coding, long-context tasks, and tool use. This entry uses the Q4_K_M GGUF; a higher-fidelity Q8_0 build is available as a variant.

Repository: localaiLicense: apache-2.0

granite-4.2-8b-q8
IBM Granite 4.2 8B in the higher-fidelity Q8_0 GGUF format. It is a multilingual reasoning model for chat, coding, and tool use.

Repository: localaiLicense: apache-2.0

Page 1