| ID | Category | Owner | Size | Context | Created | Description | a u d i o |
g a t e d |
r e a s o n |
s t r e a m |
t o o l |
v i d e o |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| bge-reranker-v2-m3-Q8_0 | Rerank | gpustack | 636 MB | - | 2024-09-09 | ShowThe bge-reranker-v2-m3-GGUF is a quantized, lightweight cross-encoder model developed by the Beijing Academy of Artificial Intelligence (BAAI) for high-performance, multilingual re-ranking tasks. It is designed to take a query and a passage as input and directly output a similarity score (relevance) rather than generating embeddings, allowing for more accurate ranking in RAG (Retrieval-Augmented Generation) systems. |
||||||
| cerebras_Qwen3-Coder-REAP-25B-A3B-Q8_0 | Text-Generation | bartowski | 26.5 GB | - | 2025-10-21 | ShowThe Qwen3-Coder-30B-A3B-Instruct is a specialized, instruction-tuned Mixture-of-Experts (MoE) model for agentic coding tasks. Developed by Alibaba, it excels in code generation, understanding large codebases, and using external tools with a high degree of efficiency. |
✓ | ✓ | ||||
| embeddinggemma-300m-qat-Q8_0 | Embedding | ggml-org | 329 MB | - | 2025-09-01 | ShowThe embeddinggemma-300M-GGUF model is a quantized version of Google's EmbeddingGemma, a lightweight, open, and multilingual text embedding model designed for efficient performance on on-device hardware like laptops, phones, and desktops. |
||||||
| gemma-4-26B-A4B-it-UD-Q4_K_M | Image-Text-to-Text | unsloth | 18.1 GB | - | 2026-04-06 | ShowGemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on small models) and generating text output. Featuring both Dense and Mixture-of-Experts (MoE) architectures, Gemma 4 is well-suited for tasks like text generation, coding, and reasoning. |
✓ | ✓ | ||||
| gpt-oss-20b-Q8_0 | Text-Generation | unsloth | 12.1 GB | - | 2025-08-08 | ShowOpenAI's GPT-OSS (OpenAI Open-Source Series) are open-weight large language models designed for powerful reasoning, agentic tasks, and on-premises deployment, released under a permissive Apache 2.0 license. |
✓ | ✓ | ✓ | |||
| LFM2-700M-Q8_0 | Text-Generation | unsloth | 792 MB | - | 2025-07-11 | ShowLFM2 is a new generation of hybrid models developed by Liquid AI, specifically designed for edge AI and on-device deployment. It sets a new standard in terms of quality, speed, and memory efficiency. |
✓ | |||||
| LFM2.5-VL-1.6B-Q8_0 | Image-Text-to-Text | unsloth | 2.1 GB | - | 2025-07-11 | ShowLFM2.5-VL-1.6B is Liquid AI's refreshed version of the first vision-language model, LFM2-VL-1.6B, built on an updated backbone LFM2.5-1.2B-Base and tuned for stronger real-world performance. |
✓ | |||||
| Ministral-3-14B-Instruct-2512-Q4_0 | Image-Text-to-Text | unsloth | 8.7 GB | - | 2025-12-02 | ShowMinistral-3-14B-Instruct-2512-GGUF is the largest and most capable model in the Mistral AI Ministral 3 family, specifically optimized for edge deployment and high-performance local inference. |
✓ | ✓ | ||||
| Qwen2-Audio-7B.Q8_0 | Audio-Text-to-Text | mradermacher | 8.9 GB | - | 2025-06-04 | ShowQwen2-Audio is the new series of Qwen large audio-language models. Qwen2-Audio is capable of accepting various audio signal inputs and performing audio analysis or direct textual responses with regard to speech instructions. |
✓ | ✓ | ||||
| Qwen3-0.6B-Q8_0 | Text-Generation | Qwen | 639 MB | - | 2025-05-03 | ShowQwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. |
✓ | ✓ | ||||
| Qwen3-8B-Q8_0 | Text-Generation | Qwen | 8.7 GB | - | 2025-05-03 | ShowThe Qwen3-8B-Q8_0 is a highly efficient, 8.2 billion parameter causal language model from the Qwen3 series, heavily quantized using 8-bit precision (Q8_0) for improved performance on consumer hardware |
✓ | ✓ | ✓ | |||
| Qwen3-Embedding-0.6B-Q8_0 | Embedding | Qwen | 639 MB | - | 2025-07-14 | ShowThe Qwen3-Embedding-0.6B is a highly efficient, 8-bit quantized text embedding model from Alibaba's Qwen3 series. It is designed for high-speed retrieval, classification, and clustering tasks while maintaining a small memory footprint. |
||||||
| Qwen3-Omni-30B-A3B-Instruct-Q8_0 | Any-to-Any | ggml-org | 33.8 GB | - | 2026-04-12 | ShowQwen3-Omni-30B-A3B-Instruct-GGUF is a multi-modal "omni" foundation model developed by Alibaba Cloud. It is optimized for instruction-following and real-time interactive tasks. It handles text, images, audio, and video inputs. It delivers real-time streaming responses in text and natural speech. |
✓ | ✓ | ✓ | |||
| qwen3-reranker-0.6b-q8_0 | Rerank | ggml-org | 636 MB | - | 2025-10-03 | ShowThe Qwen3-Reranker-0.6B-Q8_0-GGUF is a compact, high-precision cross-encoder model designed for the second stage of information retrieval pipelines. Released in June 2025 as part of the Qwen3-Embedding series, it specializes in re-scoring and reordering initial search results to improve accuracy. |
||||||
| Qwen3-VL-30B-A3B-Instruct-Q4_K_M | Image-Text-to-Text | unsloth | 19.7 GB | - | 2026-01-01 | ShowThis is a quantized version of Alibaba's Qwen3-VL multimodal large language model, specifically designed for efficient inference on resource-constrained hardware while maintaining high accuracy. |
✓ | ✓ | ||||
| Qwen3.5-0.8B-Q8_0 | Image-Text-to-Text | unsloth | 1017 MB | - | 2026-03-02 | ShowThe Qwen3.5-0.8B-GGUF is an ultra-lightweight, quantized 0.8 billion parameter language model optimized by Unsloth for high-speed, local inference on CPU or low-end GPU devices. Based on Alibaba's Qwen3.5 architecture, it provides efficient, fast responses suitable for edge computing and basic text tasks. |
✓ | ✓ | ✓ | |||
| Qwen3.6-27B-Q4_K_M | Image-Text-to-Text | unsloth | 17.7 GB | - | 2026-04-16 | ShowThe Qwen3.6-27B-GGUF is a flagship-level 27 billion parameter dense model optimized for coding, multimodal reasoning, and vision tasks, featuring a hybrid architecture with Gated DeltaNet and Gated Attention layers. |
✓ | ✓ | ✓ | |||
| Qwen3.6-35B-A3B-UD-Q4_K_M | Image-Text-to-Text | unsloth | 23.0 GB | - | 2026-04-16 | ShowThe Qwen3.6-35B-A3B-GGUF is an efficient Mixture-of-Experts model that uses only 3 billion active parameters to provide high-performance coding and multimodal reasoning. |
✓ | ✓ | ✓ | |||
| rnj-1-instruct-Q6_K | Text-Generation | unsloth | 6.4 GB | - | 2025-12-16 | ShowThe rnj-1-instruct-Q6_K is a quantized version of the rnj-1-instruct 8-billion parameter model, specifically optimized for high-performance coding, STEM, and agentic workflows, developed by Essential AI. |
✓ | ✓ | ✓ |