Article
Percepta's Spotlight Memory promises growing memory at linear cost
Percepta Research's Spotlight Memory gives language models a memory that grows with context at linear cost. Here is what its early results show.

Percepta Research has published Spotlight Memory, a new architecture that lets a language model's memory grow with its context while each token does a fixed amount of work. If the approach holds up, models could recall details from very long documents without the quadratic cost of attention or the hard limits of compressed recurrent memory.
The post is dated October 2, 2026, and the results in it are early and come from the company itself. It drew attention on r/LocalLLaMA, where it was framed as separating a model's intelligence from its memory. The real question is whether a model with fixed weights can learn to manage a memory that keeps growing.
The tradeoff Spotlight targets
Today's long-context designs force a choice. As Percepta explains, attention stores a key and value for every token and compares each new query against all of them. Processing T tokens therefore takes O(T²) comparisons.

Linear attention variants such as DeltaNet and Gated DeltaNet compress the past into a fixed-size state. That keeps the cost linear, but their memory capacity is finite. Hybrids like Qwen3-Next and Kimi Linear mix both kinds of layer, but Percepta argues they are still limited by fixed-length attention.
Percepta says it wants three properties at once: memory that grows with sequence length, sparse access where each token touches a constant number of cells, and addressing that the model learns end to end.
How Percepta compares memory and cost across architectures
Attention
- Memory grows with the history
- Cost over T steps is O(T²)
Linear attention / Gated DeltaNet
- Memory limited by state size and precision
- Cost over T steps is O(T)
Spotlight
- Memory grows with allocated cells
- Cost over T steps is O(T)
How it works
Spotlight maps keys and queries to addresses on a 2D grid of cells. Keys write to a small neighborhood around their address and queries read from a neighborhood around theirs. A cell is created the first time something writes to it, and later keys can update it.

To keep training possible, reads and writes are spread across nearby cells with a smooth bump function that approximates attention's kernel. In two dimensions, each key reads and writes exactly 3×3 cells. Percepta says this fixed region is what keeps access constant as context grows.
The 2D address only decides where to look. Each cell holds a full DeltaNet state per head, so in this design the richness of the stored content is separate from the size of the address. The model is effectively learning to read and write across a growing collection of DeltaNet states.
What the company's tests show
On synthetic recall tests, Percepta reports that Spotlight stays near-perfect. In a version of the associative recall task (MQAR) with 131,072 key-value pairs and a 524K-token context, where half the keys are overwritten, the company reports 0.998 accuracy for Spotlight. Attention scored 0.748 and Gated DeltaNet 0.002.
MQAR forgetting test at 131K pairs (company-reported)
| Item | accuracy |
|---|---|
| Spotlight | 0.998 |
| Attention | 0.748 |
| Gated DeltaNet | 0.002 |
In ordinary language modeling, the advantage is smaller. Percepta trained models at 140M, 280M and 670M parameters on FineWeb-Edu at an 8K sequence length. It says Spotlight's held-out loss stays close to the Gated DeltaNet models and below attention. On the short-context lm-eval suite, the 670M Spotlight model averaged 33.1, according to its benchmarks, compared with 34.1 for Gated DeltaNet and 33.3 for attention.
The biggest company-reported gap is in long-context recall. On RULER's single-needle task, models trained at 8K were tested at up to 128K tokens, 16 times their training length. Percepta says Spotlight reached 93 to 100% recall at 128K across all three sizes, while the fixed-state baselines reached at most 5.6% and attention scored 0% from 16K to 128K.
RULER single-needle recall at 128K, trained at 8K (company-reported)
- Spotlight, across three model sizes
- 93–100%
- Best fixed-state recurrent baseline
- 5.6%
- Beyond the training context length
- 16x
After a short extra training stage at 128K, Percepta reports that the 670M Spotlight model reached the lowest held-out loss, 1.440 nats per token at 128K. It also says accuracy on HELMET's TREC-coarse in-context classification task rose from 63% to 85% as prompts grew from 8K to 128K tokens.
Who is behind it
The post's citation lists Christos Tzamos, Guoqing Zheng and Athul Jacob as authors. It is part of a research series on Percepta's blog that includes "Can LLMs Be Computers?" and "Constructing an LLM-Computer". A companion post published the same day presents memory architectures as a path to AI that keeps learning.
Percepta is not a pure research lab. The same blog describes it as an AI transformation company announced by General Catalyst, with deployments at Summa Health and with the state of Maryland. That context matters: the company says its goal is decision-making systems that improve from experience without expensive retraining.
What is still unknown
Every number above comes from Percepta, and Percepta itself calls them early results. We could not find independent evaluations beyond the Reddit thread and a Korean aggregator post. We also could not confirm an arXiv paper, code, weights or a license.
The largest model tested has 670M parameters, far smaller than today's frontier models. The headline idea of adding knowledge and skills without changing weights rests mainly on the companion post's Python-interpreter demonstration, which we could not examine in detail.
So the answer to the opening question is a qualified yes. In Percepta's own tests, a model with fixed weights learned to manage a growing memory and recall from contexts far longer than it was trained on, at linear cost. Whether that holds at scale, and outside Percepta's lab, is still to be shown.