Model Gallery

36 models from 1 repositories

Filter by type:

Filter by tags:

ornith-1.5-9b-uncensored
# Ornith-1.5-9B-uncensored An **abliterated** (refusal-direction-ablated) build of `ornith-ai/Ornith-1.5-9B`, produced with ZeroFuse and published by junafinity. This is the **9B control checkpoint** (bf16). Mac users should start from the MLX-8bit or GGUF-8bit siblings. The official 9B base has **no `mtp.*` tensors**; nothing was grafted. **Vision tower and MTP heads are preserved** — see Vision & MTP preservation for the before/after audit. ## Intended use: red teaming and defensive cybersecurity research These uncensored (abliterated) weights are built as a **research instrument** for red teaming and defensive cybersecurity work. Safety training suppresses the *display* of capability, not capability itself. A refusal tells you the model declined. It does not tell you whether the weights could have complied. That conflation underestimates the true ceiling and hides holes in *your* filters, classifiers, and policy layer. Use each uncensored checkpoint as the **treatment half of a controlled pair** against its original base model: ...

Repository: localaiLicense: apache-2.0

s1-mini-f16
S1-mini by Superwhisper in the publisher's 1.4 GB F16 GGUF format. This variant preserves full model fidelity for hosts with enough memory.

Repository: localaiLicense: s1-mini-license

llm-jp-4-33b-thinking-bf16
LLM-jp-4-33B-thinking in the official 66.4 GB BF16 GGUF format. This variant preserves the original model precision for hosts with enough memory.

Repository: localaiLicense: apache-2.0

mxbai-embed-large-v1-f16
Mixedbread's mxbai-embed-large-v1 in the official full-precision F16 GGUF format. This 335M-parameter English BERT model produces 1,024-dimensional embeddings for retrieval, semantic search, and RAG.

Repository: localaiLicense: apache-2.0

maple-preview-tq1-0-head-f16
Maple-Preview is DeepGrove's 20B mixture-of-experts reasoning model with about 1B active parameters. This build uses TQ1_0 ternary GGUF weights with a F16 output head. It uses the publisher's CPU configuration and embedded Jinja chat template with an 8K default context. The native context is 128K tokens.

Repository: localaiLicense: mit

maple-preview-tq2-0-head-f16
Maple-Preview is DeepGrove's 20B mixture-of-experts reasoning model with about 1B active parameters. This build uses TQ2_0 ternary GGUF weights with a F16 output head. It uses the publisher's CPU configuration and embedded Jinja chat template with an 8K default context. The native context is 128K tokens.

Repository: localaiLicense: mit

neohorse-1-4b-bf16
NeoHorse-1-4B is TokenRhythm's text-only Qwen3.5-4B fine-tune for coding, reasoning, and agentic tasks. This official BF16 GGUF build uses the embedded Jinja chat template and a 32K-token default context.

Repository: localaiLicense: apache-2.0

ced-base-f16
CED (Consistent Ensemble Distillation, Xiaomi) is a sound-event classifier that tags everyday sounds (baby cry, footsteps, glass breaking, alarms, dog bark, ...) into the 527-class AudioSet ontology. This is the f16 GGUF for the ced backend (a standalone C++/ggml port). Recommended default: fastest on CPU and near-lossless. Use POST /v1/audio/classification, or the realtime websocket API for live recognition.

Repository: localaiLicense: apache-2.0

ced-tiny-f16
CED-tiny (5.5M params, Pi-class / edge) sound-event classifier over the 527-class AudioSet ontology (baby cry, footsteps, glass breaking, alarms, dog bark, ...). f16 GGUF for the ced backend (recommended (fastest on CPU)). Use POST /v1/audio/classification, or the realtime websocket API for live recognition.

Repository: localaiLicense: apache-2.0

ced-mini-f16
CED-mini (9.6M params, low-power) sound-event classifier over the 527-class AudioSet ontology (baby cry, footsteps, glass breaking, alarms, dog bark, ...). f16 GGUF for the ced backend (recommended (fastest on CPU)). Use POST /v1/audio/classification, or the realtime websocket API for live recognition.

Repository: localaiLicense: apache-2.0

ced-small-f16
CED-small (22M params, balanced size/accuracy) sound-event classifier over the 527-class AudioSet ontology (baby cry, footsteps, glass breaking, alarms, dog bark, ...). f16 GGUF for the ced backend (recommended (fastest on CPU)). Use POST /v1/audio/classification, or the realtime websocket API for live recognition.

Repository: localaiLicense: apache-2.0

qwen3-coder-30b-a3b-vllm-cpp
Qwen3-Coder-30B-A3B on vllm.cpp: a coding and agentic-tool-use model, 30B total parameters with about 3B active per token, gated token-exact against vLLM on this engine. The tool-call parser is named explicitly rather than auto-detected, and that matters here. Qwen3-Coder's tool dialect is byte-identical on the wire to another family's, so template sniffing cannot separate the two and would fall back to the wrong parser. With qwen3_coder named, tool calls arrive as real tool_calls on the OpenAI response. This is the bf16 checkpoint, roughly 57 GB of weights, which is what the engine was gated on. Being bf16 rather than NVFP4 it does not need Blackwell on its own account, but LocalAI's CUDA images for this backend are currently built for Blackwell-family GPUs only, so on an older card use the CPU build.

Repository: localaiLicense: apache-2.0

qwen3-4b-vllm-cpp
Qwen3-4B on vllm.cpp, in bf16. The small end of the engine's gated dense family, which reaches parity with vLLM on every axis at concurrency 1. bf16 rather than NVFP4 on purpose: this is the entry that runs where the flagship NVFP4 checkpoints cannot, including Apple Silicon via Metal, Vulkan and plain CPU. Roughly 8 GB of weights, plus about 4.5 GB of KV cache at the context configured here. Tool calling and the thinking split are parsed inside the engine.

Repository: localaiLicense: apache-2.0

qwen3-0.6b-vllm-cpp
Qwen3-0.6B on vllm.cpp, in bf16. Roughly 1.4 GB of weights, which makes it the cheapest way to confirm a vllm-cpp install actually serves before committing disk and memory to one of the large checkpoints. It runs anywhere the backend does, CPU included, and it is a real chat model rather than a stub, so tool calling and the thinking split can be exercised on it too.

Repository: localaiLicense: apache-2.0

moss-tts-cpp-v1_5-f16
MOSS-TTS-Local v1.5 (C++/ggml, moss-tts.cpp), F16 (~9.4 GB), code-exact versus the reference. 48kHz stereo text-to-speech with reference-audio voice cloning.

Repository: localaiLicense: apache-2.0

magpie-tts-cpp-357m-f16
Magpie TTS Multilingual 357M (C++/ggml, magpie-tts.cpp), F16 (~784 MB). 22.05kHz mono text-to-speech with 5 baked voices and 9+ languages.

Repository: localaiLicense: other

allenai_olmo-3.1-32b-think
The **Olmo-3.1-32B-Think** model is a large language model (LLM) optimized for efficient inference using quantized versions. It is a quantized version of the original **allenai/Olmo-3.1-32B-Think** model, developed by **bartowski** using the **imatrix** quantization method. ### Key Features: - **Base Model**: `allenai/Olmo-3.1-32B-Think` (unquantized version). - **Quantized Versions**: Available in multiple formats (e.g., `Q6_K_L`, `Q4_1`, `bf16`) with varying precision (e.g., Q8_0, Q6_K_L, Q5_K_M). These are derived from the original model using the **imatrix calibration dataset**. - **Performance**: Optimized for low-memory usage and efficient inference on GPUs/CPUs. Recommended quantization types include `Q6_K_L` (near-perfect quality) or `Q4_K_M` (default, balanced performance). - **Downloads**: Available via Hugging Face CLI. Split into multiple files if needed for large models. - **License**: Apache-2.0. ### Recommended Quantization: - Use `Q6_K_L` for highest quality (near-perfect performance). - Use `Q4_K_M` for balanced performance and size. - Avoid lower-quality options (e.g., `Q3_K_S`) unless specific hardware constraints apply. This model is ideal for deploying on GPUs/CPUs with limited memory, leveraging efficient quantization for practical use cases.

Repository: localaiLicense: apache-2.0

lightonocr-2-1b-f16
LightOnOCR-2-1B F16 is the full-precision GGUF build for optical character recognition and multilingual document understanding. It pairs the F16 language model with the matching F16 vision projector.

Repository: localaiLicense: apache-2.0

hunyuan-ocr-bf16
HunyuanOCR in BF16 GGUF format for maximum model and vision-projector fidelity. It runs on llama.cpp and supports document parsing, text spotting, information extraction, and text-image translation.

Repository: localaiLicense: tencent-hunyuan-community

huihui-ai_huihui-gpt-oss-20b-bf16-abliterated
This is an uncensored version of unsloth/gpt-oss-20b-BF16 created with abliteration (see remove-refusals-with-transformers to know more about it).

Repository: localaiLicense: apache-2.0

depth-anything-3-base-f16
Depth Anything 3 (base), f16 — half precision (~233 MB), no measurable accuracy loss vs f32. Depth + camera pose.

Repository: localaiLicense: apache-2.0

Page 1