WeChat WeMM-Embedding 2B beats prior 8B on MMEB-v2, 9B hits 80.6
TL;DR
- The 9B WeMM-Embedding variant reports a new MMEB-v2 state-of-the-art overall score of 80.6.
- The 2B variant claims to already surpass the previously leading 8B open-source baseline on MMEB-v2.
- WeChat says the models are backed by a 26-task in-house benchmark and 14 online A/B tests.
WeChat's engineering team says the 9B variant of its new WeMM-Embedding family hits a new state-of-the-art overall score of 80.6 on MMEB-v2, according to a technical report on arXiv. The smaller 2B variant, the report claims, "already surpasses the previously leading 8B open-source baseline on MMEB-v2."
The family covers text, images, videos, visual documents, and arbitrarily interleaved multimodal inputs, in 2B, 4B, and 9B sizes with flexible output dimensions. Training runs in two stages: a large-scale multimodal alignment pass, then a refinement stage using curated data, fine-grained relevance supervision, and cross-scale knowledge transfer.
Alongside the public numbers, the authors point to internal evidence, reporting "substantial gains on a 26-task in-house benchmark and consistent improvements across 14 online A/B tests" across WeChat properties.
The abstract does not publish per-variant MMEB-v2 scores, per-task deltas, the identities of the in-house tasks, or the name of the 8B baseline that the 2B is said to beat.
Originally reported by paper
Read the original article →Original headline: WeChat's 2B WeMM-Embedding Beats Prior 8B SOTA, Sets MMEB-v2 Record at 80.6 in Production