Concepts

Looped transformers: how recurrent depth reuses the weights you already have.

Every token that passes through a standard transformer gets the same treatment: a fixed stack of layers, each with its own parameters, executed exactly once. Capability comes from adding more unique layers, more width, more data. There is another dial that gets less attention. Take a block of layers and run it repeatedly, feeding its output back in as input. The parameter count stays fixed while the executed depth grows, and the model gains a knob it can turn at test time, spending more sequential computation on harder inputs.

This idea has been rediscovered roughly every three years since 2016, under different names: adaptive computation time, the Universal Transformer, deep equilibrium models, deep thinking, looped transformers, recurrent depth. The most recent versions scale to billions of parameters and show something the early papers could only hint at: a language model that improves on reasoning benchmarks when you let it iterate longer, without generating a single extra token. This post walks through that line of work, what a shared block learns to compute when you loop it, and which training details decide whether extra loops at inference time help or hurt.

Two budgets you can scale separately

It helps to separate two budgets that the standard architecture fuses into one. The first is the number of unique parameters, which sets how much the model can store: facts, patterns, vocabulary statistics. The second is the number of layer executions per token, which sets how much sequential computation each prediction gets. A 32-layer transformer executes 32 layers because it has 32 layers. Sharing weights across layers cuts that link.

ALBERT made the efficiency case for sharing in 2019 (Lan et al., 2019). Tie all parameters across layers and a BERT-large-sized model drops from 334M to 18M parameters with a modest accuracy cost, and the ablation shows where that cost lives: sharing the attention matrices costs almost nothing, while sharing the feed-forward blocks accounts for most of the drop. Sharing also changes the representation geometry. The distance between a layer's input and output is far smoother across ALBERT's depth than across BERT's.

A recurrent-depth model pushes the same idea further. Huginn, a 3.5B language model trained by Geiping et al. (2025), has only 8 unique transformer layers: two in a prelude that embeds the input, four in a core block, two in a coda that decodes. Run the core block 32 times and one forward pass executes 132 layers, deeper than almost any fixed-depth model, from a small set of weights. Three things follow. Training and serving move fewer unique weights between devices, which helps when bandwidth is the constraint. The model occupies less memory than a dense model of the same executed depth. And the architecture carries a bias toward iterative procedures rather than stored lookups, which turns out to matter for which tasks improve.

The early line: halting, looping, and fixed points

The oldest member of this family is Adaptive Computation Time (Graves, 2016). An RNN gets a halting unit that scores, after each internal update, whether to keep processing the current input or move on, with a penalty in the loss charging for time spent. On synthetic tasks like parity and multi-digit addition, networks with ACT learn to scale their update count with problem difficulty, roughly one step per digit on addition. On character-level Wikipedia text the model allocates extra updates at word boundaries and punctuation, and none over unpredictable ID number strings. The weak point was the time penalty itself, a hand-set coefficient trading accuracy against speed, and the trained behavior was sensitive to its value.

The Universal Transformer (Dehghani et al., 2018) moved recurrence from the time axis to the depth axis. Instead of recurring over positions in the sequence, the model applies one self-attention block over and over, revising every position's representation in parallel, with optional per-position halting borrowed from ACT. Two results from that paper still frame the discussion. On algorithmic tasks like copying and addition, evaluated on strings ten times longer than the training strings, the Universal Transformer generalizes where the standard transformer and the LSTM fall apart. And because its depth can grow with the input, the model escapes the constant sequential step count of a fixed-depth transformer, which the authors use to argue Turing completeness under stated assumptions.

Deep equilibrium models (Bai et al., 2019) then asked what the loop converges to. If you tie weights and iterate, many networks approach a fixed point, so why not solve for that point directly with a root-finder and differentiate through it with the implicit function theorem? The payoff is constant training memory regardless of effective depth, up to an 88% reduction on WikiText-103 language modeling in their benchmarks. One honest complication: the ALBERT authors measured their layer embeddings oscillating rather than settling, so the equilibrium picture describes some weight-tied networks and not others. Whether a loop converges, cycles, or diverges is a property the training recipe has to establish, not one the architecture guarantees.

What a loop learns to compute

The strongest evidence that looping changes what a network can express comes from algorithmic tasks. Schwarzschild et al. (2021) trained recurrent ResNets on easy problem instances, 32-bit prefix sums, 9x9 mazes, low-rated chess puzzles, then tested on harder ones. Recurrent models trained with few iterations extrapolate when granted more iterations at test time, reaching above 90% accuracy on longer prefix sums where feed-forward networks of equal or greater effective depth stay near chance. Watching the per-iteration outputs shows something close to a known algorithm: on prefix sums the model resolves early bits first and marches down the string, on mazes it floods outward from the start and then prunes. Their framing is worth keeping: recurrence is both the test-time dial and a pressure during training toward parameters that make progress with every application.

On the theoretical side, Giannou et al. (2023) showed constructively what a looped transformer can be programmed to do. With 13 layers or fewer and an external loop feeding the output sequence back as input, they hand-built attention patterns for read, write, and conditional branch operations, then assembled a one-instruction-set computer, a calculator, matrix inversion, and gradient descent on a small network, all executed from a prompt formatted like a punchcard. This is not how trained models work, and the authors say so. What the construction establishes is a bound on expressivity: with a loop, network depth no longer needs to scale with program length. A fixed-depth model has to encode an iterative algorithm as one giant unrolled circuit. A looped model can encode the iteration itself.

Latent reasoning at scale

The recent wave applies looping to language model pretraining, in two designs. Coconut (Hao et al., 2024) loops horizontally through an existing fixed-depth model: the last hidden state, instead of being decoded into a token, is fed back as the next input embedding, so the model reasons in continuous space between special markers. Trained with a curriculum that progressively replaces language reasoning steps with these continuous thoughts, Coconut beats chain-of-thought on a synthetic planning task that requires search. The mechanism analysis is the interesting part: a continuous thought can hold several candidate next steps in superposition and prune them over later thoughts, something a sampled token cannot do, because sampling forces a commitment to one path.

Huginn takes the vertical design: a dedicated recurrent core block, looped r times per token, trained from scratch on 800B tokens. The recipe matters more than the diagram. The latent state starts from sampled noise, the embedded input is re-injected at every iteration (without this the recurrence destabilizes), the loop count is drawn per training batch from a heavy-tailed distribution so the model sees many depths, and gradients flow only through the final 8 iterations. The result is a model whose benchmark scores keep rising with r at test time, up to a compute load comparable to a 50B dense model, with the largest gains on math and code. The most diagnostic number is an OpenBookQA ablation: evaluated closed-book, Huginn trails models of similar size because it has fewer facts stored; given the relevant fact in context, it nearly closes the gap. The model trades storage for processing.

Mixture-of-Recursions (Bae et al., 2025) adds per-token routing. A lightweight router, trained from scratch rather than bolted on afterward, decides how many times each token passes through the shared block, and the KV cache stores entries only for tokens still active at a given recursion depth. Across models from 135M to 1.7B parameters, this beats both vanilla and fixed-recursion baselines at equal training compute, and the router's decisions are readable: predictable tokens like the second half of a word exit after one pass, while hard tokens iterate. Sharing and adaptivity, which earlier work treated as separate efficiency axes, end up composing well inside one architecture.

The training recipe decides whether extra loops help

In code the loop is four lines:

e = prelude(x)       # embed the input tokens
                    s = randn_like(e)    # random initial latent state
                    for i in range(r):   # r sampled per batch during training
                        s = core(e, s)   # one shared block, reused every iteration
                    p = coda(s)          # decode the final state to logits

Getting it to train is the work, and the published recipes read like a list of things that went wrong first.

  • Stability. A shared block iterated many times has failure patterns a plain stack never sees. Huginn's first large run collapsed: the correlation between token hidden states went to 1.0, meaning the model predicted the same state everywhere. A second run stayed stable but learned to ignore the incoming state, so one iteration scored the same as thirty-two. The fixes were a sandwich normalization layout, re-injecting the input at every step, careful initialization of the output projections, and a lower learning rate.
  • Gradient depth. Backpropagating through all r iterations costs memory linear in r, which defeats the purpose. The standard fix is truncated backpropagation through only the last k iterations, the same idea as truncated backprop through time in RNNs, applied along the depth axis. The prelude still receives gradients because its output enters every iteration.
  • Randomized depth. Training at one fixed loop count bakes that schedule into the learned computation. Read out early and the intermediate states can be far from any usable prediction; loop longer than trained and scores can degrade even though compute doubled. Sampling r per batch from a heavy-tailed distribution, mostly small values with occasional long runs, is what turns the loop count into a genuine dial at inference time.
  • Inference mechanics. A recurrent model gets several serving features for free. Early exit per token when successive iterates stop changing, KV-cache sharing across iterations, and self-speculative decoding all work without extra training, largely because shared projections keep cache entries compatible across depths.

When looping is the right tool

The gains concentrate where the answer requires sequential computation that scales with problem difficulty: algorithmic tasks, math, code, multi-hop reasoning, planning with search. The pattern has been stable for a decade, from ACT on parity through the Universal Transformer on length generalization to Huginn on GSM8K. Where the task is storage-bound, looping buys little. Recurrence adds no capacity for facts, and closed-book recall benchmarks show it. If your bottleneck is that the model does not know enough, wider and denser models address it; if the bottleneck is that the model cannot combine what it knows, recurrence addresses that.

The costs are real too. Iterations are sequential along the depth axis, so latency per token grows with r, and you are trading wall-clock for parameters. Serving techniques like continuous depth-wise batching, where tokens at different loop counts share the same weights in one batch, recover much of the throughput, but the hardware has to support that scheduling.

What I would watch next: scaling recurrent-depth pretraining past the 3.5B proof of concept, combining per-token loops with ordinary chain-of-thought rather than treating them as rivals, router designs like Mixture-of-Recursions that allocate depth per token, and the interpretability angle. A looped model exposes its intermediate state after every iteration, so you can watch the computation progress instead of inferring it from the output. That alone makes the architecture useful as a research instrument, whatever its final share of deployed models turns out to be.

Sources and further reading

Foundations: adaptive computation and weight sharing

What looped networks compute

Latent reasoning and recurrent depth at scale

Want to talk through what this week's research means for your own projects? I help teams turn state-of-the-art machine learning into working systems.

Get in touch
All posts