WeMM-Embedding-2B ConvRot INT8
This is a Piper-native ConvRot INT8 conversion of Tencent's WeMM-Embedding-2B, based on upstream revision bbd6cd4bf52cfc6716f752a2df80b2706720bd95. The 285 linear weights use ConvRot INT8 with group size 256; other tensors retain their source storage dtypes. The checkpoint is one model.safetensors file, without shards.
Runtime availability: Piper Engine is currently private, with an open-source release planned. The weights are public, but the installation instructions below require access to the private Piper Engine repository. They are not yet an installation path for users without that access.
Install
The example below runs as a standalone Python script using the installed Piper loader and Transformers model class. It does not start an Engine server or import model code from Hugging Face.
For users with repository access and Git authentication configured: use Linux with an NVIDIA GPU and a driver compatible with CUDA 13.0. With uv installed, create an environment:
uv venv --python 3.14
source .venv/bin/activate
uv pip install --torch-backend=cu130 \
"piper-engine @ git+https://github.com/Boffee/piper-engine.git@b9eea403be8712da9d3b492e42f80eaf5dfcb21c" \
"torch==2.14.0+cu130" "transformers==5.17.0"
The Piper revision includes the quantized-directory loader. These instructions were tested in a fresh environment with repository access on an RTX 5090. They install the piper-engine package and its dependencies; only its loading API is used here.
Embed text
Save this complete example as embed.py, then run python embed.py in the activated environment. The first run downloads and caches the checkpoint.
from pathlib import Path
import torch
from huggingface_hub import snapshot_download
from transformers import AutoProcessor, Qwen3_5ForConditionalGeneration
from piper.loading.auto import load_auto_transformers_model
model_id = "piper-gg/WeMM-Embedding-2B-ConvRot-INT8"
path = Path(snapshot_download(model_id))
device = torch.device("cuda")
processor = AutoProcessor.from_pretrained(
path, trust_remote_code=False, local_files_only=True
)
processor.tokenizer.padding_side = "right"
model = load_auto_transformers_model(
Qwen3_5ForConditionalGeneration, path, device=device
).eval()
@torch.inference_mode()
def embed(value: str | list[dict]) -> torch.Tensor:
"""Embed one text string or one list of image/video/text content blocks."""
content = [{"type": "text", "text": value}] if isinstance(value, str) else value
inputs = processor.apply_chat_template(
[{"role": "user", "content": content}],
tokenize=True,
add_generation_prompt=False,
return_dict=True,
return_tensors="pt",
processor_kwargs={"videos_kwargs": {"cap_pixels_per_frame": True}},
).to(device)
model.model.rope_deltas = None
hidden = model.model(**inputs, use_cache=False).last_hidden_state
rows = torch.arange(hidden.shape[0], device=hidden.device)
last = (inputs["attention_mask"].sum(dim=1) - 1).clamp(min=0)
return torch.nn.functional.normalize(hidden[rows, last], dim=-1)
embedding = embed("A sunny day")
print(embedding.shape) # torch.Size([1, 2048])
embed(...) returns a normalized tensor on the GPU. For a NumPy array, use embedding.float().cpu().numpy().
To use an existing download, replace Path(snapshot_download(model_id)) with Path("/path/to/model-directory"). For reproducible downloads, pass a Hub commit hash as revision= to snapshot_download.
Embed images and video
The same helper accepts content blocks. Run these after the setup above, substituting your own local files.
For video files, also run uv pip install "torchcodec==0.16.0" and ensure FFmpeg with shared libraries is installed. Text and image inputs do not need this additional setup.
image_embedding = embed([{"type": "image", "path": "photo.jpg"}])
video_embedding = embed([{"type": "video", "path": "clip.mp4"}])
combined_embedding = embed([
{"type": "image", "path": "photo.jpg"},
{"type": "text", "text": "A sunny day"},
])
Each call encodes one item; a list combines content into one embedding. The installed processor handles media loading, preprocessing, and the repository's chat template. The helper resets rotary-position state and applies the upstream last-token L2 pooling. It is an example function, not an Engine model registration.
Checkpoint format and validation
The weights load through Piper's checkpoint loader, which reconstructs ConvRot tensors from the safetensors metadata. The processor uses Transformers' from_pretrained for its configuration and tokenizer. Plain Transformers AutoModel.from_pretrained does not reconstruct this checkpoint's ConvRot weights.
This repository contains no Python model files or auto_map. It stores one 3,238,013,576-byte model.safetensors file with SHA-256 2099b551f986a3bb921affb7cb9190ad8bc72db264592f0d3eb500a8caab73c9. All 285 quantized weights and their scales retain mmap backing when loaded on CPU; moving the model to CUDA copies the weights to GPU memory.
The example has been smoke-tested with text, image, video, and combined image/text inputs with remote model imports blocked. This is not an evaluation of embedding quality; the quantized model has not been benchmarked on the upstream evaluation suites. See the upstream model card for model capabilities and Matryoshka dimensions.
Provenance and license
The conversion used Piper's offline ConvRot INT8 exporter with no fine-tuning. Tencent's WeMM-Embedding-2B code, parameters, and weights are Apache-2.0 licensed, with third-party components under their original licenses. The included LICENSE preserves Tencent's notices and applicable third-party terms.
- Downloads last month
- 35