Hacker News
Recurrent Looped Transformer
The Recurrent Looped Transformer (RLT) pairs a causal encoder with a recurrent decoder that carries its final hidden state and a layerwise sliding-window attention cache across every token, creating a temporal path that traverses 48 decoder blocks per token (96 logical blocks total). The architecture reuses weights and memory, batches known-token encoder work, and checkpoints activations while preserving exact current-policy replay, enabling full back-propagation through recurrent outputs, decoder KV, and encoder memory.