Home/Events/NeoMME: an efficient Multimodal-native and Multilingual Encoder by Hugging Face

NeoMME: an efficient Multimodal-native and Multilingual Encoder by Hugging Face

Confirmed
Confidence
90%
Impact: 80%
Updated 1h ago

Consensus Brief

Hugging Face has introduced NeoMME, a family of multilingual multimodal encoders available in 260M and 800M sizes. The model processes both text and images using a single bidirectional Transformer without relying on separate pretrained components, achieving efficient visual document retrieval.

Sourced from
Primary: Hugging Face

What Changed Since Last Update

1h ago

NeoMME represents a shift from traditional dual-tower architectures by integrating text and image processing into a single Transformer model.

Claim Ledger

4 claims tracked across sources

Confirmed Fact

Hierarchical token pooling and asymmetric quantization reduce late-interaction index storage from roughly 1.5 MB to 6 kB per page.

Confirmed Fact

NeoMME encodes about 51 pages per second on a matched 2048×2048 image input size.

Official Claim

NeoMME is available in Hugging Face Transformers and all model checkpoints are released under the Apache 2.0 license.

Confirmed Fact

NeoMME processes about 524 billion packed input tokens during training.

Role-Based Impact Analysis

Source Timeline

1 source corroborating