ELF-B-T5Gemma2

ELF-B, a continuous diffusion language model, trained on the embeddings of T5Gemma-2-270M. From the paper Scaling and Distilling Text Embeddings for Better Diffusibility.

Paper · Code · Project page · All models of the release

Model description

ELF (Embedded Language Flow) is a continuous diffusion language model that generates a sequence of text embeddings and decodes them to tokens. This model is the unchanged ELF-B trained on OpenWebText-1024 (sequences of 1024 tokens) for 5 epochs at a batch of 512, on the frozen embeddings of the T5Gemma-2 encoder. Sampling uses the SDE sampler with self-conditioning guidance; --sc sets the guidance scale and --nfe the number of steps.

How to use

With the code of the paper; the checkpoint and the tokenizer are downloaded from this repository on first use, no login needed:

git clone https://github.com/la0ka1/diffusing-scaled-text-embeddings
cd diffusing-scaled-text-embeddings
pip install -r requirements.txt

python sample.py --ckpt ELF-B-T5Gemma2 --out samples.json --n 1024 --sc 1 --nfe 64
python evaluate.py --samples samples.json

Evaluation

On OpenWebText-1024, sampling at --sc 1 --nfe 64 gives Gen. PPL 38.5 at entropy 5.41 (real text: 15.4 at 5.43; GPT-2-S: 34.1 at 5.45). This model does not reach the entropy of real text on our sampling grid, and this is its setting with the highest entropy. Gen. PPL is the perplexity of the samples under GPT-2-Large, and entropy is their unigram entropy. See the full sweep over --sc and --nfe in the paper.

Files

ELF-B-T5Gemma2.pt holds the averaged (EMA) weights and a small config (model size, embedding dimension, sequence length, vocabulary, tokenizer, embedding mean and std). The T5Gemma-2 tokenizer is included unchanged.

Limitations

The model generates unconditional web-style English text, unfiltered; it can be false or biased. It is a research artifact for studying text embeddings as latent spaces, not for any downstream use.

License

The model is trained on the outputs of a Model Derivative of T5Gemma-2 and is released under the Gemma Terms of Use. Gemma is provided under and subject to the Gemma Terms of Use found at ai.google.dev/gemma/terms (see NOTICE). Use of this repository is subject to those terms, including the Gemma Prohibited Use Policy. The code is released under the MIT License. This is a research release by the authors of the paper; it is not a Google product and is not endorsed by Google.

Citation

@article{zhang2026scaling,
  title={Scaling and Distilling Text Embeddings for Better Diffusibility},
  author={Zhang, Zekai and Tian, Yunjie and He, Yanjin and Zhang, Xiaoyan and Zhao, Dongdi and Qu, Qing and Fu, Di},
  journal={arXiv preprint arXiv:2610.01016},
  year={2026}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train la0ka1/ELF-B-T5Gemma2

Collection including la0ka1/ELF-B-T5Gemma2

Paper for la0ka1/ELF-B-T5Gemma2