Olmo Hybrid outperforms Olmo 3 7B via Gated DeltaNet layers
TL;DR
- The paper introduces Olmo Hybrid, a 7B-parameter model that keeps the Olmo 3 7B recipe but replaces its sliding window layers with Gated DeltaNet layers.
- Olmo Hybrid outperforms Olmo 3 across standard pretraining and mid-training evaluations, and the authors say it scales more efficiently than the pure transformer.
- Theoretically, the authors argue hybrids can express tasks beyond both pure transformers and linear RNNs, such as code execution.
Olmo Hybrid is a 7B-parameter model that keeps the Olmo 3 7B recipe but swaps its sliding window layers for Gated DeltaNet layers. The paper argues the swap wins on standard pretraining and mid-training evaluations, and that the hybrid scales more efficiently than the transformer baseline.
The authors open on an unresolved question: whether hybrid architectures that mix attention and linear recurrence are worth the risk of scaling up. Their answer is yes, argued on two levels. On the theory side, the paper reports that "hybrid models do not merely inherit the expressivity of transformers and linear RNNs, but can express tasks beyond both, such as code execution." On the empirical side: "Olmo Hybrid outperforms Olmo 3 across standard pretraining and mid-training evaluations."
The abstract then concedes the awkward bit. The authors admit it is "unclear why greater expressivity on specific formal problems should result in better scaling or superior performance on downstream tasks unrelated to those problems." Their closing move is to return to theory and argue why expressivity should translate to scaling efficiency, "completing the loop," as they put it. The framing they land on is that attention-plus-recurrence belongs "not merely to reduce memory during inference, but as a fundamental way to obtain more expressive models that scale better during pretraining."
The abstract publishes no per-benchmark numbers and no scaling curves; those live in the full paper. Two researchers we follow in the Who's Who tracker posted the link.
Shared on Bluesky by 2 AI experts
Originally reported by arxiv.org
Read the original article →Original headline: Olmo Hybrid: From Theory to Practice and Back