Model Gallery

258 models from 1 repositories

Filter by type:

Filter by tags:

llm-jp-4-33b-thinking-q4
LLM-jp-4-33B-thinking is an Apache-2.0 Japanese and English reasoning model from Japan's National Institute of Informatics. Its dense Llama architecture has 33 billion parameters and a 65K-token context window. The model was aligned with supervised fine-tuning and DPO for multi-turn conversation and instruction following. This default entry uses the 20.2 GB Q4_K_M GGUF. The official 66.4 GB BF16 weights are available as a higher-fidelity variant.

Repository: localaiLicense: apache-2.0

llm-jp-4-33b-thinking-bf16
LLM-jp-4-33B-thinking in the official 66.4 GB BF16 GGUF format. This variant preserves the original model precision for hosts with enough memory.

Repository: localaiLicense: apache-2.0

Attention: Trust Remote Code is required for this model
wemm-embedding-2b
WeMM-Embedding-2B is Tencent's Apache-2.0 multilingual embedding model built on Qwen3.5. This entry serves the original bfloat16 safetensors with LocalAI's Transformers backend and produces 2,048-dimensional normalized embeddings for text retrieval, semantic search, and RAG. The upstream model can also embed images and videos. LocalAI currently exposes text input through its embeddings API for this backend.

Repository: localaiLicense: apache-2.0

Attention: Trust Remote Code is required for this model
wemm-embedding-4b
WeMM-Embedding-4B is Tencent's mid-sized Apache-2.0 multilingual embedding model built on Qwen3.5. This entry serves the original bfloat16 safetensors with LocalAI's Transformers backend and produces 2,560-dimensional normalized embeddings for text retrieval, semantic search, and RAG. The upstream model can also embed images and videos. LocalAI currently exposes text input through its embeddings API for this backend.

Repository: localaiLicense: apache-2.0

Attention: Trust Remote Code is required for this model
wemm-embedding-9b
WeMM-Embedding-9B is Tencent's largest Apache-2.0 multilingual embedding model built on Qwen3.5. This entry serves the original bfloat16 safetensors with LocalAI's Transformers backend and produces 4,096-dimensional normalized embeddings for text retrieval, semantic search, and RAG. The upstream model can also embed images and videos. LocalAI currently exposes text input through its embeddings API for this backend.

Repository: localaiLicense: apache-2.0

granite-4.2-3b-q4
IBM Granite 4.2 3B is a compact multilingual reasoning model for chat, coding, long-context tasks, and tool use. This entry uses the Q4_K_M GGUF; a higher-fidelity Q8_0 build is available as a variant.

Repository: localaiLicense: apache-2.0

granite-4.2-3b-q8
IBM Granite 4.2 3B in the higher-fidelity Q8_0 GGUF format. It is a compact multilingual reasoning model for chat, coding, and tool use.

Repository: localaiLicense: apache-2.0

granite-4.2-8b-q4
IBM Granite 4.2 8B is a multilingual reasoning model for chat, coding, long-context tasks, and tool use. This entry uses the Q4_K_M GGUF; a higher-fidelity Q8_0 build is available as a variant.

Repository: localaiLicense: apache-2.0

granite-4.2-8b-q8
IBM Granite 4.2 8B in the higher-fidelity Q8_0 GGUF format. It is a multilingual reasoning model for chat, coding, and tool use.

Repository: localaiLicense: apache-2.0

granite-4.2-30b-q4
IBM Granite 4.2 30B is the family's flagship multilingual reasoning model for chat, coding, long-context tasks, and tool use. This entry uses the Q4_K_M GGUF; a higher-fidelity Q8_0 build is available as a variant.

Repository: localaiLicense: apache-2.0

granite-4.2-30b-q8
IBM Granite 4.2 30B in the higher-fidelity Q8_0 GGUF format. It is the family's flagship multilingual reasoning model for chat, coding, and tool use.

Repository: localaiLicense: apache-2.0

hy-mt2-7b-q4
Hy-MT2-7B is Tencent's 7B multilingual translation model. It supports translation instructions across 33 languages, including terminology control and style transfer. This Q4_K_M GGUF uses the embedded chat template with an 8K context window. Include the target language in your prompt.

Repository: localaiLicense: apache-2.0

hy-mt2-7b-q6
Hy-MT2-7B is Tencent's 7B multilingual translation model. It supports translation instructions across 33 languages, including terminology control and style transfer. This Q6_K GGUF uses the embedded chat template with an 8K context window. Include the target language in your prompt.

Repository: localaiLicense: apache-2.0

hy-mt2-7b-q8
Hy-MT2-7B is Tencent's 7B multilingual translation model. It supports translation instructions across 33 languages, including terminology control and style transfer. This Q8_0 GGUF uses the embedded chat template with an 8K context window. Include the target language in your prompt.

Repository: localaiLicense: apache-2.0

hy-mt2-1.8b-q4
Hy-MT2-1.8B is Tencent's compact multilingual translation model. It follows translation instructions across 33 languages and supports tasks such as terminology control, style transfer, and structure-preserving translation. This default entry uses the 1.1 GB Q4_K_M GGUF. A higher-quality Q8_0 model is available as a variant.

Repository: localaiLicense: apache-2.0

hy-mt2-1.8b-q8
Hy-MT2-1.8B in the higher-quality 1.9 GB Q8_0 GGUF format. This variant preserves more model fidelity for hosts with enough memory.

Repository: localaiLicense: apache-2.0

spark-x2.5-1.7b-q4
Spark-X2.5-1.7B is XHToken's 1.7B text model for conversation, reasoning, coding, and multilingual tasks. This build uses Q4_K_M GGUF weights, the embedded Jinja chat template, and a 32K-token default context.

Repository: localaiLicense: apache-2.0

spark-x2.5-1.7b-q8
Spark-X2.5-1.7B is XHToken's 1.7B text model for conversation, reasoning, coding, and multilingual tasks. This build uses Q8_0 GGUF weights, the embedded Jinja chat template, and a 32K-token default context.

Repository: localaiLicense: apache-2.0

spark-x2.5-4b-q4
Spark-X2.5-4B is XHToken's 4B text model for conversation, reasoning, coding, and multilingual tasks. This entry uses Q4_K_M GGUF weights; Q6_K and Q8_0 builds are available as variants. All builds use the embedded Jinja chat template and a 32K-token default context.

Repository: localaiLicense: apache-2.0

spark-x2.5-4b-q6
Spark-X2.5-4B in Q6_K GGUF format, with the embedded Jinja chat template and a 32K-token default context.

Repository: localaiLicense: apache-2.0

spark-x2.5-4b-q8
Spark-X2.5-4B in Q8_0 GGUF format, with the embedded Jinja chat template and a 32K-token default context.

Repository: localaiLicense: apache-2.0

Page 1