Hybrid linear-attention architecture
The model card describes a 32-layer decoder using Gated DeltaNet and GLA components.
The design aims to combine recurrent state updates with global anchoring while avoiding a fully quadratic attention path across the complete sequence.
