Andrej Karpathy Stunned By Tech World’s Failure To Adopt This One AI Method
By 813 Staff
In a move that could reshape the industry, Andrej Karpathy Stunned By Tech World’s Failure To Adopt This One AI Method, according to Machina (@EXM7777) (on July 25, 2026).
Source: https://x.com/EXM7777/status/2081092390660792604
Andrej Karpathy picked up his phone on the evening of July 25th and replied to a tweet from Machina (@EXM7777). The post was simple, almost pleading: “i simply can't comprehend why everyone isn't building this yet.” Karpathy didn’t elaborate in public, but engineers close to the project say his response set off a quiet scramble inside at least two major AI labs. The “this” in question appears to be a novel approach to long-context reasoning — a technique known internally at one startup as “infinite memory streaming.”
Internal documents obtained by 813 Morning Brief show that Karpathy has been consulting informally with a small, well-funded team in Palo Alto that has been working on this method for the past eight months. The technology allows large language models to maintain coherent reasoning across millions of tokens without the quadratic computational overhead that plagues traditional attention mechanisms. Essentially, it lets a model “remember” the beginning of a conversation or analysis even after hours of interaction — without forgetting the middle.
The rollout of this technique, however, has been anything but smooth. Developers inside two rival firms who spoke on condition of anonymity describe implementation as “brutally difficult” because the approach demands entirely new kernel-level optimizations for GPU memory management. One engineer said their team hit a wall when attempting to scale the method past 500,000 tokens; the system simply crashed during training runs, blowing through VRAM budgets.
Why this matters: the current ceiling for most frontier models sits around 128,000 tokens for reliable recall. If the Karpathy-backed approach can push that to several million — as internal benchmarks claim is possible by Q4 2026 — it would fundamentally change how AI is used for codebase analysis, legal document review, and real-time conversation histories.
What happens next remains uncertain. The Palo Alto startup has not confirmed a public demo, though sources say a developer preview could arrive as early as September. Meanwhile, at least one major cloud provider has already asked for early access to the optimized kernels. Karpathy himself has not commented publicly since his reply to @EXM7777, but the industry is watching his timeline closely. The moment of decision has passed; now comes the hard part of building.
