Elon Musk's retweet of Kimi's paper sparked a heated discussion in Silicon Valley, so what's the next battleground for Attention? On March 16, 2026, the Kimi team posted a paper called "Attention Residuals" on arXiv, and things quickly spiraled out of control. Musk retweeted it, and Karpathy commented, "We haven't really taken the title 'Attention is All You Need' to heart." Jerry Tworek, former co-founder of OpenAI, directly referred to it as "deep learning 2.0." A architecture paper from a Chinese team being able to spark such a level of discussion in Silicon Valley may not have been seen since DeepSeek-V3.
Musk's retweet of Kimi's paper sparked a heated discussion in Silicon Valley, so what's the next battleground for Attention? However, amidst all the excitement, most discussions remained at the level of "Kimi has come up with something new and the big shots are thrilled." Overlooked was the fact that on the same day, ByteDance's Seed team and Huazhong University of Science and Technology jointly released another paper, called Mixture-of-Depths Attention (MoDA), which addresses the same problem but takes a completely different approach. Within the same week, a third paper, "When Does Sparsity Mitigate the Curse of Depth in LLMs," co-authored by Dilxat Muhtar from Nanjing University and Shiwei Liu from MPI, among others, provided the most accurate pathological report from a theoretical perspective.
Musk's retweet of Kimi's paper sparked a heated discussion in Silicon Valley, so what's the next battleground for Attention? Three papers emerged in rapid succession, all targeting the same issue. This is no coincidence. A structural problem that has been neglected for nearly a decade has finally reached a critical point that can no longer be ignored.
Musk's retweet of Kimi's paper sparked a heated discussion in Silicon Valley, so what's the next battleground for Attention? The issue doesn't lie in the sequential dimension of attention. Attention has evolved significantly over the past few years, from multi-head attention to grouped query attention, to DeepSeek's MLA, to various sparse variants, with each generation optimizing how tokens interact with each other. This arms race is intriguing, but it obscures a fact - the way information is transmitted between layers has remained the same since the Transformer paper was published in 2017. The answer has always been the same: residual connections, h = h + f(h), a simple addition operation with no learnable parameters.
Musk's retweet of Kimi's paper sparked a heated discussion in Silicon Valley, what's the next battleground for Attention? The output of all historical layers is summed with equal weights. No selection, no forgetting, no learning. The contribution of each layer is treated equally and piled into the residual stream, regardless of whether it learns key features or noise.
Musk retweeted Kimi's paper, sparking a big discussion in Silicon Valley, what's the next battleground for Attention? Residual connections are the most successful "temporary solution" in the history of deep learning.
Musk Retweets Kimi's Paper, Sparking Heated Discussion in Silicon Valley, What's the Next Battlefield for Attention? The Most Successful Interim Solution
Musk's retweet of Kimi's paper sparked a heated discussion in Silicon Valley, so what's the next battleground for Attention? Residual connections were first proposed by He Kaiming in ResNet in 2015. The idea is extremely simple: when the network reaches around 20 layers, it can no longer be trained due to the vanishing gradient problem, which prevents the parameters in deeper layers from being updated. To solve this, a "highway" is added to each layer, allowing the input to bypass the layer and connect directly to the output. Even if the layer doesn't learn anything, the information and gradient can at least be transmitted through this shortcut. The effect was immediate, and ResNet was able to increase the network depth from around 20 layers to over 100 layers. Two years later, the Transformer was introduced, and the residual connection was adopted unchanged. Since then, this design has remained untouched.
Musk's retweet of Kimi's paper sparked a heated discussion in Silicon Valley, so what's the next battleground for Attention? It's not that no one has tried. ReZero, FixUp, and Highway Network have all made attempts at variants, making the residual weights learnable. But none of them have been incorporated into the mainstream architecture of large models, because residual connections are just too effective. They are simple, stable, and almost don't increase computational overhead, and at the scale of models at the time, the side effects had not yet been exposed.
Musk retweeted Kimi's paper, sparking a big discussion in Silicon Valley, what's the next battleground for Attention? 44% of the layers are idle
Musk's retweet of Kimi's paper sparked a heated discussion in Silicon Valley, what's the next battleground for Attention? What are the side effects? In early 2025, a team led by Shiwei Liu from Westlake University, Emory, and MPI published "The Curse of Depth". This March, a team including Dilxat Muhtar from Nanjing University further provided quantitative diagnosis in "When Does Sparsity Mitigate the Curse of Depth in LLMs", indicating that under the current mainstream large model architecture, the deeper transformations are increasingly close to identity mapping. Whatever is input, the same is output, making this layer equivalent to nothing.
Musk's retweet of Kimi's paper sparked a heated discussion in Silicon Valley, so what's the next battleground for Attention? The numbers are hard to look at. Researchers use a "usefulness score" to measure whether each layer is making meaningful transformations. In a 12-layer model, all layers are working. In a 16-layer model, three layers are idle. In a 24-layer model, nine layers are idle. In a 32-layer model, 14 layers are idle, with 44% of layers learning almost nothing. The number of parameters increased from 900 million to 2.3 billion, with a 156% increase in budget, but the number of effective layers only increased from 12 to 18.
Musk's retweet of Kimi's paper sparks heated discussion in Silicon Valley, what's the next battleground for Attention? Figure 2: Quantitative diagnosis of the depth curse - the law of diminishing returns of effective layers as the model scale grows. This image is AI-generated.
Musk's retweet of Kimi's paper sparked a heated discussion in Silicon Valley, so what's the next battleground for Attention? The reason is directly related to how residual connections work. The output of each layer is added to a "main highway" through residual connections. As the number of layers increases, the accumulated signal on the main highway becomes larger (can be understood as the "background volume" constantly increasing), but the amplitude of the new signal generated by each layer is limited. By the time it reaches the deeper layers, the new signal is drowned out by the background noise, and the input and output are almost the same, making the layer virtually useless.
Musk's retweet of Kimi's paper sparked a heated discussion in Silicon Valley, so what's the next battleground for Attention? Residual connections solved the problem of "letting gradients pass through", but created the problem of "making deep layers meaningful".
Musk's retweet of Kimi's paper sparked a heated discussion in Silicon Valley, so what's the next battleground for Attention? In the era of large models, this cost is real money. One layer is equivalent to tens of billions of floating-point operations. A 128-layer model with 44% of its layers idle would waste the computing power of nearly sixty layers. The community has been working on optimizing inference efficiency for years, with techniques such as quantization, distillation, pruning, sparse attention, and KV cache compression, all aimed at optimizing "useful computations".
Musk retweeted Kimi's paper, sparking a heated discussion in Silicon Valley: what's the next battleground for Attention? The biggest efficiency black hole is not in the quadratic complexity of attention, but in an addition operation that has remained unchanged since 2015.
Musk retweeted Kimi's paper, sparking a heated discussion in Silicon Valley, what's the next battleground for Attention? Not repairing the old road, but paving a new one
Musk's retweet of Kimi's paper sparked a heated discussion in Silicon Valley, so what's the next battleground for Attention? The starting point for ByteDance's Seed team and Huazhong University of Science and Technology is not "residual connections are broken, they need to be replaced." Their question is more direct - since the attention mechanism can already allow tokens to see each other, why can't it also see information in the depth direction at the same time?
Musk's retweet of Kimi's paper sparked a heated discussion in Silicon Valley, so what's the next battleground for Attention? Traditional attention only has one dimension - the sequence dimension. When performing attention calculations, a token at the 20th layer can only see information from other tokens within the same layer. It cannot see its own state at the 3rd layer or the 10th layer, even if those shallowly learned features are extremely useful for the current calculation. Although these shallow features are still present in the residual flow, they have been repeatedly updated and diluted by the residual connections of over a dozen layers. When deeper layers try to use shallow features, they can only use this "diluted fruit juice" that has been watered down over a dozen times.
Musk's retweet of Kimi's paper sparked a heated discussion in Silicon Valley, so what's the next battleground for Attention? MoDA's approach is to add a second dimension - a depth dimension - to attention. Each attention head performs normal sequential attention (token-by-token) while also performing depth attention (directly retrieving the original KV pairs from all previous layers). The two streams of information are jointly normalized under the same Softmax, allowing the model to decide whether to focus on the current layer's context or revisit the features learned in shallower layers. Residual connections are still used, but they are no longer the only way for deeper layers to access information from shallower layers.
Musk's retweet of Kimi's paper sparked a heated discussion in Silicon Valley, so what's the next battleground for Attention? The idea is not hard to understand, the hard part is how to implement it without compromising speed.
Musk's retweet of Kimi's paper sparks heated discussion in Silicon Valley, what's the next battleground for Attention? Figure 3 | MoDA's dual-dimensional attention mechanism - sequence dimension and depth dimension are jointly normalized under the same Softmax
Musk retweeted Kimi's paper, sparking a heated discussion in Silicon Valley. What's the next battleground for Attention? Moving scattered archives to workstations
Elon Musk's retweet of Kimi's paper sparked a heated discussion in Silicon Valley, so what's the next battleground for Attention? The problem lies in the memory access pattern of GPUs. In normal attention calculations, all key-value (KV) pairs come from the same layer and are stored continuously in video memory, allowing for high GPU read efficiency. However, MoDA requires retrieving KV pairs from all previous layers, which are scattered across different locations in video memory. GPUs are particularly averse to this kind of "scattered" random access, resulting in a drastic decline in speed. If all historical layers' KV pairs are naively concatenated, a 48-layer model would require each layer's attention calculation to access the "archives" of the previous 47 layers, making nearly all memory access random and rendering training speeds unusable.
Musk's retweet of Kimi's paper sparked a heated discussion in Silicon Valley, so what's the next battleground for Attention? MoDA's solution is called Grouped Rearrangement. The core idea is that since random access is slow, the data should be rearranged into a continuous format before computation.
Musk's retweet of Kimi's paper sparked a heated discussion in Silicon Valley, so what's the next battleground for Attention? The approach involves two steps. First, the current layer's query is divided into several groups of fixed size (e.g., 64 tokens per group). Second, for each group, the required depth KV (from previous layers) is moved from scattered memory locations to a continuous memory region, rearranged, and then attention calculations are performed in one go. This can be understood as having an assistant move all the necessary archives to the worker's desk, rather than having the worker run back and forth to retrieve them. The cost of moving the archives is much lower than the cost of repeated trips.
Musk's retweet of Kimi's paper sparked a heated discussion in Silicon Valley, so what's the next battleground for Attention? The key to this design lies in the granularity of the grouping. If the groups are too large, each group needs to handle too many deep KV pairs, making the handling itself a bottleneck. If the groups are too small, the GPU's parallel computing capabilities are underutilized. MoDA chose the same block size as FlashAttention (currently the industry-standard high-speed attention computing engine), allowing the deep attention calculation to directly reuse FlashAttention's underlying implementation, without the need to write a whole new set of GPU algorithms.
Musk's retweet of Kimi's paper sparked a heated discussion in Silicon Valley, what's the next battleground for Attention? At a sequence length of 64K, MoDA's operator efficiency reached 97.3% of FlashAttention-2. With the addition of the entire deep attention mechanism, the speed was only slowed by less than 3%.
Musk's retweet of Kimi's paper sparks heated discussion in Silicon Valley, what's the next battleground for Attention? Figure 4 | Group reorganization strategy - relocating scattered historical layer KV in video memory to continuous memory regions
Musk's retweet of Kimi's paper sparked a big discussion in Silicon Valley, what's the next battleground for Attention? This figure means that deep attention is not a lightweight plugin, it requires every layer to read all previous layers' KV cache. If the engineering is rough, this cross-layer data dependency can slow down the training speed by several times. MoDA compressed the extra overhead to a 3.7% FLOPs increment, indicating that the grouping and rearrangement strategy indeed solved the random access problem very cleanly.
Musk's retweet of Kimi's paper sparks heated discussion in Silicon Valley, what's the next battleground for Attention? With a cost of 3.7% and a return of 2.11%
Elon Musk's retweet of Kimi's paper sparked a heated discussion in Silicon Valley, so what's the next battleground for Attention? On a 1.5B parameter model (based on the OLMo2 training recipe), MoDA achieved an average performance boost of 2.11% on 10 downstream tasks, with additional computational overhead of only 3.7%. Although this may seem modest at first glance, this is an improvement at the architectural level, not achieved through more data or longer training, and will continue to take effect as the scale increases. Moreover, the differences between tasks are significant, with commonsense reasoning (WinoGrande) improving by 2.37% and scientific reasoning (ARC-Challenge) improving by 4.35%, with tasks requiring cross-layer feature integration benefiting more noticeably.
Musk's retweet of Kimi's paper sparks heated discussion in Silicon Valley, what's the next battleground for Attention? Figure 5 | MoDA's performance comparison on 10 downstream tasks
Musk's retweet of Kimi's paper sparks heated discussion in Silicon Valley, what's the next battleground for Attention? Pre-Norm's unpaid debt
Musk's retweet of Kimi's paper sparked a heated discussion in Silicon Valley, so what's the next battleground for Attention? The most valuable part of the MoDA paper may not be MoDA itself, but an experiment on normalization strategies.
Musk's retweet of Kimi's paper sparked a heated discussion in Silicon Valley, so what's the next battleground for Attention? To provide some background, after each layer of the Transformer completes its calculations, it undergoes a step called "normalization" (Normalization), which stabilizes the range of values and prevents numerical explosions or disappearances during training. There are two mainstream approaches to where this processing step is placed: before each layer's calculation, called Pre-Norm (also known as Pre-LN), and after the calculation, called Post-Norm (also known as Post-LN). Since 2020, almost all large models have used Pre-Norm, as it makes training more stable and less prone to collapse. However, the "deep layer idle rotation" problem mentioned earlier is actually a side effect of Pre-Norm. In order to stabilize training, Pre-Norm is constantly diluting the signal intensity of the deep layers.
Musk's retweet of Kimi's paper sparked a heated discussion in Silicon Valley, so what's the next battleground for Attention? MoDA's experiment conducted two comparative studies on a 48-layer model, using Pre-Norm and Post-Norm, and then added MoDA's deep attention to each group. Under the Post-Norm configuration, the addition of deep KV resulted in a 0.0409 reduction in validation loss; whereas Pre-Norm only saw a 0.0041 reduction, a difference of nearly tenfold.
Musk's retweet of Kimi's paper sparked a heated discussion in Silicon Valley, so what's the next battleground for Attention? This data reveals something bigger than MoDA itself, namely that Pre-Norm is not only used for "stable training" but also systematically suppresses the deep learning ability. In the past, people dared not use Post-Norm because the training was unstable and the gradient was prone to explosion. However, MoDA's deep attention provides a brand-new gradient path, allowing the gradient to be transmitted without completely relying on residual connections. With this new path, the original instability issue of Post-Norm is no longer a fatal flaw.
Musk's retweet of Kimi's paper sparked a heated discussion in Silicon Valley, so what's the next battleground for Attention? The combination of MoDA+Post-Norm opens up possibilities that the compromises made in the past for stable training (using Pre-Norm) may be reclaimed.
Musk's retweet of Kimi's paper sparks heated discussion in Silicon Valley, what's the next battleground for Attention? Figure 6 | Pre-Norm vs Post-Norm: Verification loss difference after adding deep KV
Musk's retweet of Kimi's paper sparked a heated discussion in Silicon Valley, what's the next battleground for Attention? Don't pave new roads, just repair the old ones.
Elon Musk's retweet of Kimi's paper sparked a heated discussion in Silicon Valley, so what's the next battleground for Attention? MoDA doesn't have residual connections, instead it chooses to take a different route outside of residuals. On the same day, Kimi's team released Attention Residuals (AttnRes), which took a more direct approach by directly modifying the residual connections themselves.
Musk's retweet of Kimi's paper sparked a heated discussion in Silicon Valley, so what's the next battleground for Attention? The standard residual connection does a simple thing, equally weighting and summing the outputs of all previous layers, and stacking them onto the main path. There's no selection, no forgetting. AttnRes replaces this fixed equal-weighted addition with an attention operation, where each layer uses its own state as a query, and the outputs of all previous layers as candidates. It uses attention to determine which previous layers' features are useful for the current layer, and what their respective weights are.
Musk's retweet of Kimi's paper sparked a heated discussion in Silicon Valley, so what's the next battleground for Attention? The residual connection has evolved from a fixed formula into a dynamically routable one that can be learned.
Musk's retweet of Kimi's paper sparks heated discussion in Silicon Valley, what's the next battleground for Attention? Figure 7: The core idea of AttnRes - using attention to replace equal-weighted residual addition
Musk's retweet of Kimi's paper sparked a heated discussion in Silicon Valley, so what's the next battleground for Attention? The cost is that each layer requires an additional run of deep attention computation, which is not cheap. The Kimi team uses a block-based strategy (Block AttnRes) to control costs, dividing layers into several blocks, performing complete deep attention within blocks, and only focusing on block-level aggregated representations between blocks.
Musk's retweet of Kimi's paper sparked a heated discussion in Silicon Valley, so what's the next battleground for Attention? AttnRes has already been integrated into Kimi Linear (480 billion total parameters / 300 million activated parameters), which was pre-trained on 14 trillion tokens, and its effectiveness has been consistently confirmed across different model scales. This paper has been widely reported on, so there's no need to delve into the technical details here. What's worth discussing is how it compares to MoDA's approach.
Musk retweeted Kimi's paper, sparking a heated discussion in Silicon Valley, what's the next battleground for Attention? Figure 8 | AttnRes's training curve and ablation experiment
Musk's retweet of Kimi's paper sparked a heated discussion in Silicon Valley, so what's the next battleground for Attention? Two diagnostic routes have completely consistent causes, namely, the shallow information obtained by the deep layer is repeatedly diluted by residual updates. However, the approaches differ in where they intervene. MoDA does not touch the residual connection, but instead adds a depth dimension to the attention mechanism, allowing the deep layer to bypass the residual flow and directly access the shallow layer's original features. AttnRes, on the other hand, directly targets the residual connection, replacing the equal-weighted addition with attention-weighted addition. One approach is "building a new road," while the other is "renovating the original road."
Musk's retweet of Kimi's paper sparked a heated discussion in Silicon Valley, so what's the next battleground for Attention? Two papers emerged on the same day, with different approaches but the same target. This is no coincidence. The depth of attention has already become a consensus in the research community, with the only difference being the direction of approach.
Musk retweeted Kimi's paper, sparking a heated discussion in Silicon Valley, what's the next battleground for Attention? Figure 9 | AttnRes's consistent performance across different model sizes
Musk Retweets Kimi's Paper, Sparking Heated Discussion in Silicon Valley, What's the Next Battlefield for Attention? Forgotten "Scaffolding" to be Dismantled
Musk's retweet of Kimi's paper sparked a heated discussion in Silicon Valley, what's the next battleground for Attention? Going back to the original question, why did the issue of deep learning overfitting only start to be taken seriously in 2026?
Musk's retweet of Kimi's paper sparked a heated discussion in Silicon Valley, so what's the next battleground for Attention? It's because residual connections are too useful. They solved a pressing problem at the time (vanishing gradients), with a controllable cost (degradation in deep models is not obvious in small models), and alternative solutions were not mature (ReZero and Highway Network had not been verified on a large scale). No one had the motivation to change it. It wasn't a deliberate design choice, but a temporary solution that was forgotten. The scaffolding that was set up initially was left standing after the building was completed, and over time, everyone thought it was a load-bearing wall.
Musk Retweets Kimi's Paper, Sparking Heated Discussion in Silicon Valley, What's the Next Battlefield for Attention? Figure 1|The Signal Dilution Effect of Residual Connections - The deeper the layers, the harder it is for new signals to be heard. This image is AI-generated.
Musk's retweet of Kimi's paper sparked a heated discussion in Silicon Valley, so what's the next battleground for Attention? However, what makes this issue difficult to discover is not the residual connection itself, but the fact that the attention mechanism has long been operating in only one dimension. Over the past eight years, all the evolution of attention - multi-head, grouped queries, sparse, linear - has been focused on the sequence dimension. How tokens interact with each other has been optimized countless times. But how do layers interact with each other? This question has never been asked. The depth dimension is the blind spot of attention.
Musk's retweet of Kimi's paper sparked a heated discussion in Silicon Valley, so what's the next battleground for Attention? MoDA and AttnRes have opened up this blind spot from different directions. MoDA added a second dimension to attention, allowing it to operate simultaneously in both sequence and depth directions. AttnRes turned inter-layer information transmission itself into an attention operation. Although the approaches differ, they both point to the same conclusion: attention should not only look horizontally, but also vertically.
Musk's retweet of Kimi's paper sparked a big discussion in Silicon Valley, what's the next battleground for Attention? The implications of this conclusion are greater than the two papers themselves. There are still many fixed mechanisms in Transformers that only work in a single dimension. Each layer must be executed in sequence and cannot be skipped. Each attention head calculates independently and then simply concatenates, without dynamic coordination between heads. Each token, regardless of difficulty, goes through the same calculation path. These designs were originally engineering compromises to make the model trainable and convergent.
Musk's retweet of Kimi's paper sparked a heated discussion in Silicon Valley, so what's the next battleground for Attention? The evolution of deep learning over the past decade, if abstracted to the highest level, can be summed up as one thing: transferring more and more structural decision-making from human designers to the models themselves. Hand-designed convolutional kernels have been replaced by learnable attention. Fixed positional encoding has been replaced by learnable rotational encoding. Fixed expert allocation has been replaced by learnable routing. Now, the way information flows in the depth dimension is also being determined by attention itself.
Musk's retweet of Kimi's paper sparked a heated discussion in Silicon Valley, so what's the next battleground for Attention? Karpathy said we haven't taken the literal meaning of "Attention is All You Need" to heart. He might be right, but not in the sense that "attention is enough", rather that "attention hasn't been utilized enough". It has evolved many generations in the sequence dimension, but has just begun in the depth dimension.
Musk's retweet of Kimi's paper sparked a heated discussion in Silicon Valley, what's the next battleground for Attention? Depth is the next battlefield for attention.
Musk's retweet of Kimi's paper sparks heated discussion in Silicon Valley, what's the next battleground for Attention? This article comes from "Tencent Technology", authored by Boyang, edited by Xu Qingyang, and published by 36Kr with authorization.