On the Design Fundamentals of Pixel Text Representation Learning
Paper • 2609.01147 • Published • 29
Vision-only encoders for language representation directly in pixel space.
Note Stage 1: foundational multilingual pretraining.
Note Stage 1 + Stage 2: full curriculum.
Note Stage 2 only: paper headline English STS and ViDoRe model.