Tag fidelity vs fluency: paper proposes three-level LLM fix
TL;DR
- The paper argues balancing translation fluency with tag fidelity requires interventions at data synthesis, capability building, and alignment simultaneously.
- Hy-LST combines two LLM-based synthesis methods to generate training data that is both tag-diverse and linguistically natural.
- The system was tested across six language directions and reportedly beats prior methods, though the abstract publishes no per-language numbers or baselines.
Large language models translating formatted web content run into a specific tension: keep the HTML-style tags in place, or let the prose read naturally. A new arXiv preprint from Zhanglin Wu and seven co-authors, posted September 24, argues the fix has to work at three levels at once.
"Internet texts are replete with format tags that carry structural, semantic, and functional meaning," the paper opens, calling out a "fundamental trade-off between structural tag diversity and translation naturalness in synthetic data generation." Their hybrid data-generation strategy, Hy-LST, combines two synthesis methods. On top of that, they split tag-aware translation into "four sub-tasks of increasing difficulty" for multi-task supervised fine-tuning, then apply three reward functions targeting fluency, tag fidelity, and tag-scoped translation quality under group relative policy optimization.
The evaluation covers six language directions: English into Chinese, Japanese, German, French, and Russian, plus German-to-French. The abstract reports each level contributes measurable gains and the full system "significantly outperforms existing methods." It publishes no per-language numbers, no named baselines, and nothing on whether the models will be released.
Two researchers we track shared the paper.
Shared on Bluesky by 2 AI experts
Originally reported by arxiv.org
Read the original article →Original headline: Tag-Aware Structured Text Translation: Towards a Systematic Understanding