WeMM-Embedding-2B ConvRot INT8

This is a Piper-native ConvRot INT8 conversion of Tencent's WeMM-Embedding-2B, based on upstream revision bbd6cd4bf52cfc6716f752a2df80b2706720bd95. The 285 linear weights use ConvRot INT8 with group size 256; other tensors retain their source storage dtypes. The checkpoint is one model.safetensors file, without shards.

Runtime availability: Piper Engine is currently private, with an open-source release planned. The weights are public, but the installation instructions below require access to the private Piper Engine repository. They are not yet an installation path for users without that access.

Install

The example below runs as a standalone Python script using the installed Piper loader and Transformers model class. It does not start an Engine server or import model code from Hugging Face.

For users with repository access and Git authentication configured: use Linux with an NVIDIA GPU and a driver compatible with CUDA 13.0. With uv installed, create an environment:

uv venv --python 3.14
source .venv/bin/activate
uv pip install --torch-backend=cu130 \
  "piper-engine @ git+https://github.com/Boffee/piper-engine.git@b9eea403be8712da9d3b492e42f80eaf5dfcb21c" \
  "torch==2.14.0+cu130" "transformers==5.17.0"

The Piper revision includes the quantized-directory loader. These instructions were tested in a fresh environment with repository access on an RTX 5090. They install the piper-engine package and its dependencies; only its loading API is used here.

Embed text

Save this complete example as embed.py, then run python embed.py in the activated environment. The first run downloads and caches the checkpoint.

from pathlib import Path

import torch
from huggingface_hub import snapshot_download
from transformers import AutoProcessor, Qwen3_5ForConditionalGeneration

from piper.loading.auto import load_auto_transformers_model

model_id = "piper-gg/WeMM-Embedding-2B-ConvRot-INT8"
path = Path(snapshot_download(model_id))
device = torch.device("cuda")
processor = AutoProcessor.from_pretrained(
    path, trust_remote_code=False, local_files_only=True
)
processor.tokenizer.padding_side = "right"
model = load_auto_transformers_model(
    Qwen3_5ForConditionalGeneration, path, device=device
).eval()


@torch.inference_mode()
def embed(value: str | list[dict]) -> torch.Tensor:
    """Embed one text string or one list of image/video/text content blocks."""
    content = [{"type": "text", "text": value}] if isinstance(value, str) else value
    inputs = processor.apply_chat_template(
        [{"role": "user", "content": content}],
        tokenize=True,
        add_generation_prompt=False,
        return_dict=True,
        return_tensors="pt",
        processor_kwargs={"videos_kwargs": {"cap_pixels_per_frame": True}},
    ).to(device)
    model.model.rope_deltas = None
    hidden = model.model(**inputs, use_cache=False).last_hidden_state
    rows = torch.arange(hidden.shape[0], device=hidden.device)
    last = (inputs["attention_mask"].sum(dim=1) - 1).clamp(min=0)
    return torch.nn.functional.normalize(hidden[rows, last], dim=-1)


embedding = embed("A sunny day")
print(embedding.shape)  # torch.Size([1, 2048])

embed(...) returns a normalized tensor on the GPU. For a NumPy array, use embedding.float().cpu().numpy().

To use an existing download, replace Path(snapshot_download(model_id)) with Path("/path/to/model-directory"). For reproducible downloads, pass a Hub commit hash as revision= to snapshot_download.

Embed images and video

The same helper accepts content blocks. Run these after the setup above, substituting your own local files.

For video files, also run uv pip install "torchcodec==0.16.0" and ensure FFmpeg with shared libraries is installed. Text and image inputs do not need this additional setup.

image_embedding = embed([{"type": "image", "path": "photo.jpg"}])
video_embedding = embed([{"type": "video", "path": "clip.mp4"}])

combined_embedding = embed([
    {"type": "image", "path": "photo.jpg"},
    {"type": "text", "text": "A sunny day"},
])

Each call encodes one item; a list combines content into one embedding. The installed processor handles media loading, preprocessing, and the repository's chat template. The helper resets rotary-position state and applies the upstream last-token L2 pooling. It is an example function, not an Engine model registration.

Checkpoint format and validation

The weights load through Piper's checkpoint loader, which reconstructs ConvRot tensors from the safetensors metadata. The processor uses Transformers' from_pretrained for its configuration and tokenizer. Plain Transformers AutoModel.from_pretrained does not reconstruct this checkpoint's ConvRot weights.

This repository contains no Python model files or auto_map. It stores one 3,238,013,576-byte model.safetensors file with SHA-256 2099b551f986a3bb921affb7cb9190ad8bc72db264592f0d3eb500a8caab73c9. All 285 quantized weights and their scales retain mmap backing when loaded on CPU; moving the model to CUDA copies the weights to GPU memory.

The example has been smoke-tested with text, image, video, and combined image/text inputs with remote model imports blocked. This is not an evaluation of embedding quality; the quantized model has not been benchmarked on the upstream evaluation suites. See the upstream model card for model capabilities and Matryoshka dimensions.

Provenance and license

The conversion used Piper's offline ConvRot INT8 exporter with no fine-tuning. Tencent's WeMM-Embedding-2B code, parameters, and weights are Apache-2.0 licensed, with third-party components under their original licenses. The included LICENSE preserves Tencent's notices and applicable third-party terms.

Downloads last month
35
Safetensors
Model size
3B params
Tensor type
F32
路
BF16
路
I8
路
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for piper-gg/WeMM-Embedding-2B-ConvRot-INT8

Finetuned
Qwen/Qwen3.5-2B
Finetuned
(4)
this model