Skip to content

Latest commit

 

History

History
23 lines (22 loc) · 8.58 KB

File metadata and controls

23 lines (22 loc) · 8.58 KB

Model Catalog

ID Category Owner Size Context Created Description a
u
d
i
o
g
a
t
e
d
r
e
a
s
o
n
s
t
r
e
a
m
t
o
o
l
v
i
d
e
o
bge-reranker-v2-m3-Q8_0 Rerank gpustack 636 MB - 2024-09-09
ShowThe bge-reranker-v2-m3-GGUF is a quantized, lightweight cross-encoder model developed by the Beijing Academy of Artificial Intelligence (BAAI) for high-performance, multilingual re-ranking tasks. It is designed to take a query and a passage as input and directly output a similarity score (relevance) rather than generating embeddings, allowing for more accurate ranking in RAG (Retrieval-Augmented Generation) systems.
cerebras_Qwen3-Coder-REAP-25B-A3B-Q8_0 Text-Generation bartowski 26.5 GB - 2025-10-21
ShowThe Qwen3-Coder-30B-A3B-Instruct is a specialized, instruction-tuned Mixture-of-Experts (MoE) model for agentic coding tasks. Developed by Alibaba, it excels in code generation, understanding large codebases, and using external tools with a high degree of efficiency.
embeddinggemma-300m-qat-Q8_0 Embedding ggml-org 329 MB - 2025-09-01
ShowThe embeddinggemma-300M-GGUF model is a quantized version of Google's EmbeddingGemma, a lightweight, open, and multilingual text embedding model designed for efficient performance on on-device hardware like laptops, phones, and desktops.
gemma-4-26B-A4B-it-UD-Q4_K_M Image-Text-to-Text unsloth 18.1 GB - 2026-04-06
ShowGemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on small models) and generating text output. Featuring both Dense and Mixture-of-Experts (MoE) architectures, Gemma 4 is well-suited for tasks like text generation, coding, and reasoning.
gpt-oss-20b-Q8_0 Text-Generation unsloth 12.1 GB - 2025-08-08
ShowOpenAI's GPT-OSS (OpenAI Open-Source Series) are open-weight large language models designed for powerful reasoning, agentic tasks, and on-premises deployment, released under a permissive Apache 2.0 license.
LFM2-700M-Q8_0 Text-Generation unsloth 792 MB - 2025-07-11
ShowLFM2 is a new generation of hybrid models developed by Liquid AI, specifically designed for edge AI and on-device deployment. It sets a new standard in terms of quality, speed, and memory efficiency.
LFM2.5-VL-1.6B-Q8_0 Image-Text-to-Text unsloth 2.1 GB - 2025-07-11
ShowLFM2.5-VL-1.6B is Liquid AI's refreshed version of the first vision-language model, LFM2-VL-1.6B, built on an updated backbone LFM2.5-1.2B-Base and tuned for stronger real-world performance.
Ministral-3-14B-Instruct-2512-Q4_0 Image-Text-to-Text unsloth 8.7 GB - 2025-12-02
ShowMinistral-3-14B-Instruct-2512-GGUF is the largest and most capable model in the Mistral AI Ministral 3 family, specifically optimized for edge deployment and high-performance local inference.
Qwen2-Audio-7B.Q8_0 Audio-Text-to-Text mradermacher 8.9 GB - 2025-06-04
ShowQwen2-Audio is the new series of Qwen large audio-language models. Qwen2-Audio is capable of accepting various audio signal inputs and performing audio analysis or direct textual responses with regard to speech instructions.
Qwen3-0.6B-Q8_0 Text-Generation Qwen 639 MB - 2025-05-03
ShowQwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models.
Qwen3-8B-Q8_0 Text-Generation Qwen 8.7 GB - 2025-05-03
ShowThe Qwen3-8B-Q8_0 is a highly efficient, 8.2 billion parameter causal language model from the Qwen3 series, heavily quantized using 8-bit precision (Q8_0) for improved performance on consumer hardware
Qwen3-Embedding-0.6B-Q8_0 Embedding Qwen 639 MB - 2025-07-14
ShowThe Qwen3-Embedding-0.6B is a highly efficient, 8-bit quantized text embedding model from Alibaba's Qwen3 series. It is designed for high-speed retrieval, classification, and clustering tasks while maintaining a small memory footprint.
Qwen3-Omni-30B-A3B-Instruct-Q8_0 Any-to-Any ggml-org 33.8 GB - 2026-04-12
ShowQwen3-Omni-30B-A3B-Instruct-GGUF is a multi-modal "omni" foundation model developed by Alibaba Cloud. It is optimized for instruction-following and real-time interactive tasks. It handles text, images, audio, and video inputs. It delivers real-time streaming responses in text and natural speech.
qwen3-reranker-0.6b-q8_0 Rerank ggml-org 636 MB - 2025-10-03
ShowThe Qwen3-Reranker-0.6B-Q8_0-GGUF is a compact, high-precision cross-encoder model designed for the second stage of information retrieval pipelines. Released in June 2025 as part of the Qwen3-Embedding series, it specializes in re-scoring and reordering initial search results to improve accuracy.
Qwen3-VL-30B-A3B-Instruct-Q4_K_M Image-Text-to-Text unsloth 19.7 GB - 2026-01-01
ShowThis is a quantized version of Alibaba's Qwen3-VL multimodal large language model, specifically designed for efficient inference on resource-constrained hardware while maintaining high accuracy.
Qwen3.5-0.8B-Q8_0 Image-Text-to-Text unsloth 1017 MB - 2026-03-02
ShowThe Qwen3.5-0.8B-GGUF is an ultra-lightweight, quantized 0.8 billion parameter language model optimized by Unsloth for high-speed, local inference on CPU or low-end GPU devices. Based on Alibaba's Qwen3.5 architecture, it provides efficient, fast responses suitable for edge computing and basic text tasks.
Qwen3.6-27B-Q4_K_M Image-Text-to-Text unsloth 17.7 GB - 2026-04-16
ShowThe Qwen3.6-27B-GGUF is a flagship-level 27 billion parameter dense model optimized for coding, multimodal reasoning, and vision tasks, featuring a hybrid architecture with Gated DeltaNet and Gated Attention layers.
Qwen3.6-35B-A3B-UD-Q4_K_M Image-Text-to-Text unsloth 23.0 GB - 2026-04-16
ShowThe Qwen3.6-35B-A3B-GGUF is an efficient Mixture-of-Experts model that uses only 3 billion active parameters to provide high-performance coding and multimodal reasoning.
rnj-1-instruct-Q6_K Text-Generation unsloth 6.4 GB - 2025-12-16
ShowThe rnj-1-instruct-Q6_K is a quantized version of the rnj-1-instruct 8-billion parameter model, specifically optimized for high-performance coding, STEM, and agentic workflows, developed by Essential AI.