Repository: localaiLicense: other

Grug 12B is kai-os's compact-reasoning fine-tune of Gemma 4 12B IT. It targets shorter, denser reasoning traces while preserving constraints, branching decisions, edge cases, and final-answer checks. This entry uses Bartowski's Q4_K_M quantization and includes the multimodal projector for Gemma 4 image inputs. The model is experimental and its reported evaluation is a small local math proxy rather than a broad benchmark. Review the upstream model card's dataset provenance and `other` license before commercial or sensitive use.
Links
Tags
Repository: localaiLicense: apache-2.0
**Orca-Agent-v0.1** is a 14-billion-parameter orchestration agent built on top of **Qwen3-14B**, designed to act as a smart decision-maker in multi-agent coding systems. Rather than writing code directly, it strategically breaks down complex tasks into subtasks, delegates to specialized agents (e.g., explorers and coders), verifies results, and maintains contextual knowledge throughout execution. Trained using GRPO and curriculum learning on 32 H100 GPUs, it achieves strong performance on TerminalBench (18.25% accuracy) when paired with a Qwen3-Coder-30B MoE subagent—nearly matching the performance of a 480B model. It's optimized for real-world coding workflows, especially in infrastructure automation and system recovery. **Key Features:** - Full fine-tuned Qwen3-14B base model - Designed for multi-agent collaboration (orchestrator + subagents) - Trained on real terminal tasks with structured feedback - Serves via vLLM or SGLang for high-throughput inference **Use Case:** Ideal for advanced autonomous coding systems, DevOps automation, and complex problem-solving in technical environments. 👉 **Original Training Repo:** [github.com/Danau5tin/Orca-Agent-RL](https://github.com/Danau5tin/Orca-Agent-RL) 👉 **Orchestration Code:** [github.com/Danau5tin/multi-agent-coding-system](https://github.com/Danau5tin/multi-agent-coding-system)
Links
Tags
Repository: localaiLicense: apache-2.0
**Model Name:** Spiral-Qwen3-4B-Multi-Env **Base Model:** Qwen3-4B (fine-tuned variant) **Repository:** [spiral-rl/Spiral-Qwen3-4B-Multi-Env](https://huggingface.co/spiral-rl/Spiral-Qwen3-4B-Multi-Env) **Quantized Version:** Available via GGUF (by mradermacher) --- ### 📌 Description: Spiral-Qwen3-4B-Multi-Env is a fine-tuned, instruction-optimized version of the Qwen3-4B language model, specifically enhanced for multi-environment reasoning and complex task execution. Built upon the foundational Qwen3-4B architecture, this model demonstrates strong performance in coding, logical reasoning, and domain-specific problem-solving across diverse environments. The model was developed by **spiral-rl**, with contributions from the community, and is designed to support advanced, real-world applications requiring robust reasoning, adaptability, and structured output generation. It is optimized for use in constrained environments, making it ideal for edge deployment and low-latency inference. --- ### 🔧 Key Features: - **Architecture:** Qwen3-4B (Decoder-only, Transformer-based) - **Fine-tuned For:** Multi-environment reasoning, instruction following, and complex task automation - **Language Support:** English (primary), with strong multilingual capability - **Model Size:** 4 billion parameters - **Training Data:** Proprietary and public datasets focused on reasoning, coding, and task planning - **Use Case:** Ideal for agent-based systems, automated workflows, and intelligent decision-making in dynamic environments --- ### 📦 Availability: While the original base model is hosted at `spiral-rl/Spiral-Qwen3-4B-Multi-Env`, a **quantized GGUF version** is available for efficient inference on consumer hardware: - **Repository:** [mradermacher/Spiral-Qwen3-4B-Multi-Env-GGUF](https://huggingface.co/mradermacher/Spiral-Qwen3-4B-Multi-Env-GGUF) - **Quantizations:** Q2_K to Q8_0 (including IQ4_XS), f16, and Q4_K_M recommended for balance of speed and quality --- ### 💡 Ideal For: - Local AI agents - Edge deployment - Code generation and debugging - Multi-step task planning - Research in low-resource reasoning systems --- > ✅ **Note:** The model card above reflects the *original, unquantized base model*. The quantized version (GGUF) is optimized for performance but may have minor quality trade-offs. For full fidelity, use the base model with full precision.
Links
Tags
Repository: localaiLicense: apache-2.0
Laya is a multilingual, non-autoregressive System 1 decision model. Given a state (text, email, ticket, or JSON) and typed questions, it returns typed answers with mathematically calibrated probabilities in a single forward pass. It never generates text, so there is nothing to parse and nothing to hallucinate. In LocalAI, serve via POST /v1/systemone with this model. The vllm.cpp engine runs the full decision pipeline (choice, noul, score question types) through the vllm_decide C ABI, returning the complete kev-compatible JSON response. ModernBERT-large backbone, 421M params, 512-token context. F16 weights, ~804 MB.
Links
Tags