ELF-B-T5Gemma2
ELF-B, a continuous diffusion language model, trained on the embeddings of T5Gemma-2-270M. From the paper Scaling and Distilling Text Embeddings for Better Diffusibility.
Paper · Code · Project page · All models of the release
Model description
ELF (Embedded Language Flow) is a continuous diffusion language model that generates a sequence of text embeddings and decodes them to tokens. This model is the unchanged ELF-B trained on OpenWebText-1024 (sequences of 1024 tokens) for 5 epochs at a batch of 512, on the frozen embeddings of the T5Gemma-2 encoder. Sampling uses the SDE sampler with self-conditioning guidance; --sc sets the guidance scale and --nfe the number of steps.
How to use
With the code of the paper; the checkpoint and the tokenizer are downloaded from this repository on first use, no login needed:
git clone https://github.com/la0ka1/diffusing-scaled-text-embeddings
cd diffusing-scaled-text-embeddings
pip install -r requirements.txt
python sample.py --ckpt ELF-B-T5Gemma2 --out samples.json --n 1024 --sc 1 --nfe 64
python evaluate.py --samples samples.json
Evaluation
On OpenWebText-1024, sampling at --sc 1 --nfe 64 gives Gen. PPL 38.5 at entropy 5.41 (real text: 15.4 at 5.43; GPT-2-S: 34.1 at 5.45). This model does not reach the entropy of real text on our sampling grid, and this is its setting with the highest entropy. Gen. PPL is the perplexity of the samples under GPT-2-Large, and entropy is their unigram entropy. See the full sweep over --sc and --nfe in the paper.
Files
ELF-B-T5Gemma2.pt holds the averaged (EMA) weights and a small config (model size, embedding dimension, sequence
length, vocabulary, tokenizer, embedding mean and std). The T5Gemma-2 tokenizer is included unchanged.
Limitations
The model generates unconditional web-style English text, unfiltered; it can be false or biased. It is a research artifact for studying text embeddings as latent spaces, not for any downstream use.
License
The model is trained on the outputs of a Model Derivative of T5Gemma-2 and is released under the Gemma Terms of Use.
Gemma is provided under and subject to the Gemma Terms of Use found at
ai.google.dev/gemma/terms (see NOTICE). Use of this repository is
subject to those terms, including the Gemma Prohibited Use Policy.
The code is released under the MIT License. This is a research release by the authors of the paper; it is not
a Google product and is not endorsed by Google.
Citation
@article{zhang2026scaling,
title={Scaling and Distilling Text Embeddings for Better Diffusibility},
author={Zhang, Zekai and Tian, Yunjie and He, Yanjin and Zhang, Xiaoyan and Zhao, Dongdi and Qu, Qing and Fu, Di},
journal={arXiv preprint arXiv:2610.01016},
year={2026}
}