Madrimed1.2 2B GGUF

Optimized GGUF Quantization Tiers for MadriMed 1.2 (2B Medical Mulitmodel)

GGUF Qwen3-VL GGUF llama.cpp


Overview

This repository provides official GGUF-quantized binaries of MadriMed 1.2 (Madrimed1.2-VL-2B), a compact 2B parameter medical vision-language model. These files are optimized for high-performance, low-latency local execution using llama.cpp, llama-server, llama-cpp-python, and Ollama.

Full Model


Model Architecture & Dual-File Requirement

GGUF multimodal pipelines require a dual-file architecture:

  1. The Language Backbone (medmax-2b-[quant].gguf): Autoregressive transformer decoder (28 layers) handling clinical text generation and reasoning.
  2. The Multimodal Projector (mmproj-[quant].gguf): Vision encoder and DeepStack patch merger (qwen3vl_merger) processing image inputs at a 768 resolution.

Crucial Note: To perform visual inference, you must pass both the quantized model file (-m) and the vision projector file (--mmproj).


Complete Quantization Inventory

File Name Format Size Description & Recommended Use
mmproj-f16.gguf F16 ~781 MiB Reference vision projector (Maximum visual fidelity).
mmproj-Q8_0.gguf Q8_0 421 MiB Recommended Vision Projector: ~46% smaller than F16 with subminimal perceptible visual loss.
medmax-2b-f16.gguf F16 4.06 GB Half precision reference (16.0 BPW).
medmax-2b-Q8_0.gguf Q8_0 2.01 GB High-precision (8.5 BPW). Ideal for clinical exam reasoning.
medmax-2b-Q6_K.gguf Q6_K 1.55 GB Recommended Sweet Spot: Very high quality (6.56 BPW) with minimal perplexity degradation.
medmax-2b-Q5_K_M.gguf Q5_K_M 1.37 GB Medium 5-bit quantization (5.77 BPW). Excellent quality-to-RAM balance.
medmax-2b-Q5_K_S.gguf Q5_K_S 1.31 GB Compact 5-bit quantization (5.66 BPW).
medmax-2b-Q4_K_M.gguf Q4_K_M 1.21 GB Standard Deployment: Optimal speed, low RAM, and high VQA grounding (5.03 BPW).
medmax-2b-Q4_K_S.gguf Q4_K_S 1.17 GB Smaller 4-bit variant (4.84 BPW).
medmax-2b-IQ4_XS.gguf IQ4_XS 1.12 GB Importance-matrix 4-bit quantization (4.63 BPW).
medmax-2b-Q3_K_L.gguf Q3_K_L 1.07 GB Large 3-bit variant (4.45 BPW).
medmax-2b-Q3_K_M.gguf Q3_K_M 0.99 GB Low Memory: Medium 3-bit quantization (~1 GB flat model size) (4.20 BPW).
medmax-2b-Q3_K_S.gguf Q3_K_S 948 MiB Low Memory: Small 3-bit variant (3.92 BPW).
medmax-2b-Q2_K.gguf Q2_K 833 MiB Extreme Compression: Ultra-low memory environments (3.44 BPW).

Deployment & Usage Guide

1. Local CLI Inference (llama-cli)

Run quick diagnostic queries directly in your terminal:

!./llama.cpp/build/bin/llama-cli \
    -m medmax-2b-Q6_K.gguf \
    --mmproj mmproj-Q8_0.gguf \
    --image /path/to/scan.jpg \
    -p "Question: What abnormality is visible in this scan? Answer:" \
    --image-min-tokens 1024 \
    -ngl 99

Important Flag: Always include --image-min-tokens 1024 for clinical grounding tasks. Qwen-VL architectures require adequate token allocation to maintain high-resolution perception on subtle lesions.

2. OpenAI-Compatible API Server (llama-server)

Host a local server on your machine for integration into applications:


./llama.cpp/build/bin/llama-server \
    -m medmax-2b-Q6_K.gguf \
    --mmproj mmproj-Q8_0.gguf \
    -c 8192 \
    -np 1 \
    -ngl 99 \
    --port 8080 \
    --image-min-tokens 1024

3. OpenAI SDK

import base64
from openai import OpenAI

client = OpenAI(base_url="[http://127.0.0.1:8080/v1](http://127.0.0.1:8080/v1)", api_key="none")

def encode_image(image_path):
    with open(image_path, "rb") as f:
        return base64.b64encode(f.read()).decode("utf-8")

image_b64 = encode_image("chest_xray.jpg")

response = client.chat.completions.create(
    model="madrimed1.2-2b",
    messages=[
        {
            "role": "user",
            "content": [
                {"type": "text", "text": "Is there evidence of consolidation?"},
                {"type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{image_b64}"}}
            ]
        }
    ],
    max_tokens=128,
    temperature=0.0
)

print(response.choices[0].message.content)

Limitations & Safety Guidelines

Research & Educational Use Only

MadriMed 1.2-GGUF is an experimental research model. It is not certified for autonomous clinical diagnostic decision-making. All model-generated interpretations and visual inferences should be reviewed and independently verified by a licensed medical professional before use in any clinical workflow.

Downloads last month
803
GGUF
Model size
2B params
Architecture
qwen3vl
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

5-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for madrisight/madrimed1.2-VL-2B-GGUF

Quantized
(2)
this model