Instructions to use giovannioliveira/Qwen3-VL-32B-vision-encoder with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use giovannioliveira/Qwen3-VL-32B-vision-encoder with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-feature-extraction", model="giovannioliveira/Qwen3-VL-32B-vision-encoder")# Load model directly from transformers import AutoTokenizer, AutoModel tokenizer = AutoTokenizer.from_pretrained("giovannioliveira/Qwen3-VL-32B-vision-encoder") model = AutoModel.from_pretrained("giovannioliveira/Qwen3-VL-32B-vision-encoder", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Qwen3-VL-32B vision encoder
Vision-only extraction of model.visual from
Qwen/Qwen3-VL-32B-Instruct, revision
0cfaf48183f594c314753d30a4c4974bc75f3ccb.
The checkpoint contains the complete Qwen3VLVisionModel: patch embedding,
27 vision-transformer blocks, the 5120-dimensional merger, and the three
DeepStack merger heads. It does not contain language-model or LM-head weights.
The state-dict prefix model.visual. was removed so the checkpoint is laid out
for a standalone Qwen3VLVisionModel. The extraction script operates directly
on safetensors headers and byte ranges; it never constructs or loads a model.
No tensor value was altered: the concatenated payload of the extracted file
hashes identically to the corresponding byte ranges of the source shards.
Usage
qwen3_vl_vision is a registered model_type, so the checkpoint loads through
the auto classes with no trust_remote_code.
import torch
from PIL import Image
from transformers import AutoImageProcessor, AutoModel
repo = "giovannioliveira/Qwen3-VL-32B-vision-encoder"
processor = AutoImageProcessor.from_pretrained(repo)
model = AutoModel.from_pretrained(repo, dtype=torch.bfloat16).eval().cuda()
image = Image.open("photo.jpg").convert("RGB")
inputs = processor(images=image, return_tensors="pt").to(model.device)
with torch.inference_mode():
output = model(
hidden_states=inputs["pixel_values"].to(torch.bfloat16),
grid_thw=inputs["image_grid_thw"],
)
output.pooler_output # [merged_tokens, 5120] -- the merger output
output.deepstack_features # 3 x [merged_tokens, 5120], blocks 8, 16, 24
output.last_hidden_state # [patches, 1152], pre-merge
A 448x448 image gives 28x28 patches and, after the 2x2 spatial merge,
196 merged tokens. To feed a Qwen3-VL language model, concatenate along the
feature axis in DeepStack order: [pooler_output, *deepstack_features].
preprocessor_config.json ships the upstream longest_edge of 16,777,216,
which smart_resize treats as a total-pixel budget rather than a side length --
one 16.7 MP image becomes 16,384 merged tokens. Attention inside an image is
quadratic in that image's own patch count, so lower
processor.size["longest_edge"] to whatever your token budget actually allows.
Contents
model.safetensors: 351 BF16 tensors and 595,266,800 parametersconfig.json: standaloneqwen3_vl_visionconfigurationpreprocessor_config.json: image preprocessing configurationvideo_preprocessor_config.json: video preprocessing configurationextraction_manifest.json: source revision, sizes, and checksumsextract_vision_safetensors.py,verify_extraction.py: the extraction and verification tooling
Verification
extraction_manifest.json records the checksums, so a download can be checked
against it without touching the source model:
sha256sum model.safetensors
# 8f079a0bbeaeebb9af02b9094cbf83a717349c850b8600601e6ec2ced32a80de
To reverify against the upstream weights instead, fetch the source shards and run the file-level verifier. It compares tensor names, dtypes, shapes, offsets, file size, and the source/destination payload checksums, and loads no tensor:
# ~64 GB of source shards
snapshot=$(hf download Qwen/Qwen3-VL-32B-Instruct \
--revision 0cfaf48183f594c314753d30a4c4974bc75f3ccb \
--include "model*.safetensors" "model.safetensors.index.json")
python3 verify_extraction.py "$snapshot" model.safetensors
extract_vision_safetensors.py reproduces the checkpoint from the same
snapshot:
python3 extract_vision_safetensors.py "$snapshot" model.safetensors
License and attribution
The original Qwen3-VL-32B-Instruct weights are copyright the Qwen team,
Alibaba Cloud, and are released under the Apache License, Version 2.0.
This repository redistributes a subset of those weights under the same license; see
LICENSE for the full text.
Modifications relative to the upstream checkpoint, in full:
- only tensors under the
model.visual.prefix are included; every language-model and LM-head tensor is dropped - the
model.visual.prefix is stripped from the remaining tensor names config.jsonis the parent config'svision_configverbatim, withmodel_typechanged fromqwen3_vltoqwen3_vl_visionandarchitectures,dtype, andtransformers_versionadded so the checkpoint resolves as a standalone model- tensor payload bytes are copied verbatim and are unchanged
- Downloads last month
- 25