Published event
ArtificialIntelligence
Research
1 source(s)
NeoMME: an efficient Multimodal-native and Multilingual Encoder
Summary
NeoMME: an efficient Multimodal-native and Multilingual Encoder NeoMME: an efficient Multimodal-native and Multilingual Encoder Team Article Published September 3, 2026 Upvote 110 Tony Wu tonywu71 Hcompany Aurélien Lac h-aurelien-lac Hcompany TL;DR We introduce NeoMME , a family of 260M and 800M multilingual multimodal encoders. Unlike many generative visual language models, NeoMME does not use a separate pretrained vision tower or a causal language model.
Why it matters
This Research is relevant to the technology intelligence record because it involves NVIDIA, Hugging Face, GitHub, Hugging Face Transformers. The source article should remain the factual reference for follow-up coverage.
Key facts
- NeoMME: an efficient Multimodal-native and Multilingual Encoder Team Article Published September 3, 2026 Upvote 110 Tony Wu tonywu71 Hcompany Aurélien Lac h-aurelien-lac Hcompany TL;DR We introduce NeoMME , a family of 260M and 800M multilingual multimodal encoders.
- Unlike many generative visual language models, NeoMME does not use a separate pretrained vision tower or a causal language model.
- A single bidirectional Transformer processes both text tokens and raw image patches, and we train the entire model from scratch with a masked discrete-diffusion objective.
- We fine-tuned NeoMME for visual document retrieval using ColPali's page-image approach.
- NeoMME -Retriever returns dense and late-interaction embeddings in one forward pass.
- Both model sizes lie on the ViDoRe v3 Pareto frontier for nDCG@10 and model size.
Entities in this story
Related events