huggingface.co via Reddit

Bilibili Open-Sources Index-Translate 35B MoE for 150 Languages

TL;DR

  • Bilibili's Index LLM Team released Index-Translate-35B-A3B-preview, a 35B-parameter mixture-of-experts with 3B active, covering 150 languages under Apache 2.0.
  • The preview scores 0.8794 COMET-22 on FLORES and reports a 2.4% off-target rate on low-resource pairs, described as lowest in its comparison set.
  • The family adds dense 2B/9B text models plus Echo speech translation, Homura syllable-controlled dubbing and NativeLong full-document checkpoints at 2B/9B.

Bilibili's Index LLM Team has put out an Apache 2.0 translation model family whose top preview is a 35-billion-parameter mixture-of-experts activating 3 billion at a time, covering 150 languages. The model card on Hugging Face lists Index-Translate-35B-A3B-preview as built on a Qwen3.5 backbone, alongside dense 2B and 9B siblings for text and separate 2B/9B checkpoints for the Echo (speech-to-text and speech-to-speech translation), Homura (syllable-controlled dubbing) and NativeLong (full-document) tasks.

The reported numbers are the pitch. On FLORES the preview lands a 0.8794 COMET-22, and on low-resource pairs the card puts its off-target rate at 2.4%, the lowest in its comparison set. Domain scores reach 0.8962 COMET-22 on MuST-Cinema subtitles and 0.8681 on biomedical text. The general-capability sheet shows 77.6% on C-Eval and 49.1% on GPQA-Diamond.

Training used 167.77 billion tokens of multilingual mid-training, split into a constant stage mixing general, parallel and monolingual data at 1:1:1 and a decay stage weighted 1:4:2 toward pivot data. Post-training runs through specialist SFT plus RL for general translation, instruction following and meme translation; a general-translation RL step judged by XCOMET-XXL with language-validity and adequacy checks; an instruction step the card describes as "Rubric-as-Reward with hard checks and graded constraints"; parameter interpolation of the specialists; and a final multi-teacher on-policy distillation.

The card qualifies its own claims: "Scores vary across languages and tasks. Combined constraints and low-resource directions remain harder than core-language translation, and successful individual examples do not guarantee that every instruction will be followed."

Serving is a `vllm serve` one-liner with a shipped 262,144-token context window and a recommended `--max-model-len 32768`. This is the second Chinese open-weights release we've logged this week, after InclusionAI's 560B Ling-3.1-flash.