arXiv is now an independent nonprofit! Learn more
License: CC BY 4.0
arXiv:2609.00885v1 [cs.AI] 01 Sep 2026

Denoising Diffusion Generative Models Secretly Calculate Attentions

Farzan Haddadi ††thanks: School of Electrical Engineering, Iran University of Science & Technology, Tehran, IRAN    Leila Monfared ††thanks: School of Electrical Engineering, Iran University of Science & Technology, Tehran, IRAN    Ebrahim Rezaii ††thanks: School of Electrical Engineering, Iran University of Science & Technology, Tehran, IRAN    Mohammadreza Malek-Mohammadi*    Pejman Zakalvand ††thanks: School of Electrical Engineering, Iran University of Science & Technology, Tehran, IRAN    Narges Mokhtari ††thanks: * Independent researcher††thanks: School of Electrical Engineering, Iran University of Science & Technology, Tehran, IRAN
Abstract

Denoising diffusion models are the dominant architecture for image generation, whereas most natural language generation and modeling are primarily handled by well-known transformer architectures employing attention mechanism. Here, we show that diffusion models also inherently use an attention mechanism very similar to that of transformers. Therefore, attention emerges as a universal machine learning principle, based on a general training objective. We also show similarities in basic functional principle of auto-encoders and attention-based models. These equivalences allows us to interchange these designs based on practical requirements. As an example, we can reformulate the diffusion framework to reduce the lengthy training process and computation-intensive image generation. Using this approach, a simplified algorithm is proposed for image generation which is based on attention mechanism. Results show that the attention-based implementation achieves comparable performance with significantly less effort and computational resources.

Index Terms: 
Diffusion model, attention mechanism, autoencoder, data manifold, latent space, interpolation.

I Introduction

Artificial intelligence (AI) encompasses a multitude of proposed architectures for different tasks. Throughout the history of AI, it was a surprise that gradually efficient unified system architectures replaced this multitude of systems for various applications. An important turning point was the advent of the dot-product-based attention mechanism [1]. Transformers were first proposed for tasks of natural language processing (NLP) such as question-answering and chatbots [2]. Soon, it was realized that transformers also have stellar performance in vision applications with the advent of the vision transformer (ViT) model [3].

In image generation, diffusion systems are arguably the best option [4]. Yet, for the best performance they are used in tandem with transformer-based components. Specifically, diffusion principle is used to determine the loss function and high-level train and test schemes, while in network level, an auto-encoder is trained which in turn uses attention mechanism and transformer architecture as its main component [5].

Here, we show that the diffusion principle is equivalent to the attention mechanism. Diffusion denoising probabilistic generative model (DDPM) [4], trains a neural network to reconstruct the original image from its noisy copies. Then, starting from pure noise, the system can generate a realistic image based on the learned distribution of the dataset.

Detecting the presence of a signal from its noisy version is a classic statistical problem [6]. Specifically, when the noise is Gaussian, the solution is an inner-product detector. While in the context of machine learning theory, we train a neural network to solve this problem, we will show that this problem enjoys a closed-form solution which is basically the attention mechanism. This not only shows that the diffusion system is identical to the attention mechanism, but also provides a previously known statistical basis for the very attention mechanism as the closed form solution to the signal recovery problem in the Gaussian noise [7].

The closed form solution of the diffusion system was first derived in [8]. The authors calculate optimum score function of a diffusion system for a finite dataset. However, they did not notice the relation of their result to the well-known attention mechanism. Then [8] extends the result to include additional structures like CNN network which is used in the DDPM implementation and develops a creativity theory for CNN diffusion systems based on locality and positional invariance properties. Here we use a different approach based on the diffusion loss function and convex optimization to reach the same closed form solution while extending the analysis to include resemblance to transformer and autoencoders. The equivalence between diffusion and transformer principles have broad theoretical and practical consequences in machine learning theory.

It is not sufficient to understand the core diffusion principle in the forms of what we investigate as closed-form one-step solution or convergence trajectory. In practice, diffusion systems are implemented in the inner latent space of a Variational Auto-Encoder (VAE) [9]. Also, in practice the main neural network that performs denoising in the diffusion system is an AE. We argue that the same equivalence exists between attention and AE, in that both systems eventually project a random input on a learned data manifold [10, 11]. Therefore, we reach to an equivalence between three of the most important modern AI systems, namely: ‌attention, diffusion , and AE. This enables us to interchange these systems depending on the situation.

DDPM diffusion system suffers from heavy computations in both training and generation which requires running an encoder hundreds of times on a random input. Some efforts have been made to reduce computational complexity specially in generation. A trend in the literatur is to reduce the steps for image generation using non-Markovian statistical models as in Diffusion Denoising Implicit Model (DDIM) [12].

Regarding the taining, DDIM is equivalent to DDPM and very computationally complex. It requires adding noise for hundreds of times to each image in dataset to teach network to denoise images with varying degree of noise effectively. Therefore, reducing computational complexity in the training phase is highly important. Here, we calculate a closed form denoiser which reduces the need for training a network on the dataset.

As a result of the equivalence of diffusion and attention, we conclude that deliberate interpolation between neighbor real images in the latent space results is an approximately realistic image. This interpolation can be refined and projected on the real image manifold by a limited recursive application of the outer AE. Overall, the proposed system can generate a realistic image without training an inner AE network, and just by using the dataset images. The results show good quality with considerably lower computations.

The manuscript is organized as follows:‌ The main result and discussion is presented in Sec. II, where a background review, discussion on equivalence between diffusion and attention, and convergence topics are covered. AutoEncoder equivalence with attention and projection on data manifold are addressed in Sec. III. A data structure for fast attention calculation is discussed in Sec. IV. The proposed low complexity generative network is demonstrated in Sec. V, while simulation results are exhibited in Sec. VI.

II Diffusion Systems Calculate Attention

II-A Background

A diffusion system is trained to denoise noisy images in a recurssive manner. In the forward process, Gaussian noise is gradually added to the true image x0x_{0} in hundreds or thousands of steps [13]:

‌​xt=α​xt−1+‌​1−α​nt\displaystyle‌x_{t}=\sqrt{\alpha}\;x_{t-1}+‌\sqrt{1-\alpha}\;n_{t}
nt∼𝒩⁡(0,I)\displaystyle n_{t}\sim\mathcal{N}(0,I)

for α<‌​1\alpha<‌1. Therefore, each image in the forward process can be written based on the true initial image as:

xt=αt​x0+‌​1−αt​nt′x_{t}=\sqrt{\alpha^{t}}\;x_{0}+‌\sqrt{1-\alpha^{t}}\;n^{\prime}_{t} (1)

This eventually leads to a noise-only image at the final stage of the forward process since limt→+∞αt=0\lim_{t\to+\infty}\alpha^{t}=0. We train a neural network to recover the true image x0x_{0} from each noisy image xtx_{t}. Theoretically, this results in a neural network that can produce a realistic image from a pure noise and based on the true distribution of the data which is represented by the dataset used for training.

II-B Prior Art

In practice, we train the denoiser neural network on a finite dataset 𝒟={s1,…,sN}\mathcal{D}=\{s_{1},\ldots,s_{N}\}. This finite dataset training situation was first analyzed in [8]. Adding Gaussian noise incurs a mixture of Gaussian distribution on dataset points in signal space:

pt​(y)=1|𝒟|​∑s∈𝒟𝒩⁡(y|αt​s,(1−αt)​I)p_{t}(y)\;=\;\frac{1}{|\mathcal{D}|}\sum_{s\in\mathcal{D}}\mathcal{N}(y\;|\;\sqrt{\alpha^{t}}s\;,(1-\alpha^{t})I) (2)

Then [8] uses a flow matching continuous-time model which is solved backward in time for signal generation:

−y˙t=γt​(yt+st​(yt))-\dot{y}_{t}=\gamma_{t}\left(y_{t}+s_{t}(y_{t})\right) (3)

in which γt\gamma_{t} is a constant and st​(y):=∇y​log​pt​(y)s_{t}(y):=\nabla_{y}\log p_{t}(y) is the score function. The score function of the mixture distribution in (2) is then [8]:

st​(y)=11−αt​∑s∈𝒟(αt​s−y)​Wt​(y,s)s_{t}(y)=\frac{1}{1-\alpha^{t}}\sum_{s\in\mathcal{D}}(\sqrt{\alpha^{t}}s-y)W_{t}(y,s) (4)
Wt​(y,s)=𝒩⁡(y|αt​s,(1−αt)​I)∑s′∈𝒟𝒩⁡(y|αt​s′,(1−αt)​I)W_{t}(y,s)=\frac{\mathcal{N}(y\;|\;\sqrt{\alpha^{t}}s\;,(1-\alpha^{t})I)}{\sum_{s^{\prime}\in\mathcal{D}}\mathcal{N}(y\;|\;\sqrt{\alpha^{t}}s^{\prime}\;,(1-\alpha^{t})I)} (5)

Then [8] continues the analysis to develop a theory for combinatorial creativity of diffusion models when the denoiser is a CNN network.

II-C Diffusion is Attention

Our analysis starts from the same finite dataset training assumption and then proceeds in a different direction to find the relation of the diffusion model to transformers and autoencoders as current dominant models in machine learning research.

When we train the network on the finite dataset 𝒟\mathcal{D}, the network learns to recover each sis_{i} from its noisy versions. We guess that when sampling the system for a new signal, although we use a random noise as the input, the system still outputs just an interpolation of the signals it had seen in the dataset since it has not seen anything else.

For signal sis_{i} and Gaussian noise nn, DDPM minimizes Mean-Squared Error (MSE) loss function:

DDPM:minθ∥x^0(y,t;θ)−si∥2:‌y=si+‌n\text{DDPM}:\quad\min_{\theta}\|\hat{x}_{0}(y,t;\theta)-s_{i}\|^{2}\quad:‌\quad y=s_{i}+‌n (6)

in which x0x_{0} is the real signal, x^0\hat{x}_{0} its estimate, t∈{0,…,T}t\in\{0,\ldots,T\} time index, and θ\theta is the neural network’s parameters. If we infinitely train the network with fixed sis_{i} and yy, it learns to convert input yy to output sis_{i}.

But we do not train the network just for one sample image. Adding Gaussian noise to any signal sis_{i} for sufficiently many times, can reach any point yy in the signal space many times. A geometrical scheme of the problem is depicted in Fig. 1. Reaching any fixed point yy from various signals sis_{i} occurs in accordance with the Gaussian distribution of the noise term nn. Therefore, if we fix yy, the system is trained with an average loss function:

minθ∑i=1Np⁡(y|si)​‖x^0​(y,t,θ)−si‖2\displaystyle\min_{\theta}\quad\sum_{i=1}^{N}p(y|s_{i})\|\hat{x}_{0}(y,t;\theta)-s_{i}\|^{2} (7)
p⁡(y|si)=12​π​σ​exp⁡{−‖y−si‖22​σ2}\displaystyle p(y|s_{i})=\frac{1}{\sqrt{2\pi}\sigma}\exp\left\{-\frac{\|y-s_{i}\|^{2}}{2\sigma^{2}}\right\}
s1s_{1}s2s_{2}s3s_{3}yy
Fig. 1: Geometry of diffusion system. Three signals s1s_{1}, s2s_{2}, and s3s_{3} are used for system training. We are interested in calculating the output for random input yy. We show that the output is the attention combination of training signals.

For fixed yy, (7) is a Quadratic Programming (QP) convex optimization problem:

minz\displaystyle\min_{z}\quad ∑i=1Npi​‖z−si‖2\displaystyle\sum_{i=1}^{N}p_{i}\>\|z-s_{i}\|^{2} (8)
pi⩾0:‌i=1,⋯,N\displaystyle p_{i}\geqslant 0\quad:‌\quad i=1,\cdots,N (9)

Where (9) is a conic criterion for the strictly convex optimization in (8). The optimal solution for the above optimization problem can be derived using gradients [14]:

∇z∑ipi∥z−si∥2=2∑ipiz−2∑ipisi=0\nabla_{z}\sum_{i}p_{i}\|z-s_{i}\|^{2}=2\sum_{i}p_{i}z-2\sum_{i}p_{i}s_{i}=0 (10)
z∗=∑i=1Npi​si∑i=1Npi=∑i=1Nexp⁡{−‖y−si‖22​σ2}​si∑i=1Nexp⁡{−‖y−si‖22​σ2}z^{*}=\frac{\sum_{i=1}^{N}p_{i}s_{i}}{\sum_{i=1}^{N}p_{i}}=\frac{\sum_{i=1}^{N}\exp\left\{-\frac{\|y-s_{i}\|^{2}}{2\sigma^{2}}\right\}s_{i}}{\sum_{i=1}^{N}\exp\left\{-\frac{\|y-s_{i}\|^{2}}{2\sigma^{2}}\right\}} (11)

Assume that σ\sigma is sufficiently small. Then just those signals sis_{i} that are in the neighborhood of yy are effective in the convex combination in (11). Therefore, if NN is large enough and the space is sufficiently populated with data points, we can safely assume that ‖si‖\|s_{i}\|’s are approximately constant. Other situations also may hold the condition of approximately constant ‖si‖\|s_{i}\| like in image dataset when total energy of most of the images are in the same range. In this case we will have:

z∗≃∑i=1Nexp⁡{<y,si>σ2}​si∑i=1Nexp⁡{<‌​y,si>σ2}=Attention​{y,{si}}z^{*}\simeq\frac{\sum_{i=1}^{N}\exp\left\{\frac{<y,s_{i}>}{\sigma^{2}}\right\}s_{i}}{\sum_{i=1}^{N}\exp\left\{\frac{<‌y,s_{i}>}{\sigma^{2}}\right\}}=\text{Attention}\{y,\{s_{i}\}\} (12)

in which <⋅,⋅><\cdot\,,\cdot> denotes the inner vector product.

As we can see in (12), the optimum solution to the input signal yy in a diffusion system with fixed σ\sigma is the attention vector. But this is just one step of the overall diffusion system since each step in the forward diffusion process has a distinct increasing value of σ\sigma. We will discuss this convergence issue in the next section II-D.

The derivation resulting in (12) is also a mathematical basis for the very attention system. It shows that attention is the optimum solution for recovering signals in Gaussian noise as it was shown earlier in [7]. Note that in the conventional attention system [1], the noise variance σ2=p\sigma^{2}=\sqrt{p}, where pp is the dimension of the space.

This derivation also shows a basis for the softmax layer. The exponential form in (12) with σ=1\sigma=1, is the softmax output when sis_{i} is a standard one-hot label vector 𝕀i\mathbb{I}_{i}:

‌​zi∗=Softmaxi​{y}=eyi∑j=1Neyj\displaystyle‌z^{*}_{i}=\text{Softmax}_{i}\{y\}=\frac{e^{y_{i}}}{\sum_{j=1}^{N}e^{y_{j}}} (13)
z∗=∑i=1Ne<y,𝕀i>​𝕀i∑i=1Ne<y,𝕀i>=Attention​{y,{𝕀i}}\displaystyle z^{*}=\frac{\sum_{i=1}^{N}e^{<y,\mathbb{I}_{i}>}\mathbb{I}_{i}}{\sum_{i=1}^{N}e^{<y,\mathbb{I}_{i}>}}=\text{Attention}\{y,\{\mathbb{I}_{i}\}\} (14)

which shows the equivalence of softmax and attention layers:

Softmax​{y}=Attention​{y,{𝕀i}}\text{Softmax}\{y\}=\text{Attention}\{y,\{\mathbb{I}_{i}\}\} (15)

Note that the result in (11) can be deduced using the results in [8], replacing (4) and (5) in (3). However, the authors in [8] did not proceed to deduce the attention mechanism from their closed-form score function solution of the diffusion system in (4).

II-D Convergence

In Sec. II-C, it was shown that each step of a diffusion system calculates attention for a predefined σ\sigma value. The reverse generative process in DDPM starts with a random Gaussian yy and refines it with denoising network trained with noisy images from dataset. Theoretically, there should be hundreds of denoising networks each trained for denoising with a specific σ\sigma. But usually a shared network is trained and used for every values of σ\sigma to reduce complexity. Therefore, output of each step is then recursively fed in the network to generate another less noisy output image. This is repeated hundreds of times to reach a high quality image.

Refer to caption
Fig. 2: Convergence of a random point yy to a sparse solution through sequential attention calculation when σ\sigma is decreasing. MM is the mean of the sis_{i} signals. Starting point yy passes near MM when σ\sigma is large in the first stages.

As can be seen from (1), σt2=1−αt\sigma^{2}_{t}=1-\alpha^{t} is increasing with tt in the forward process. This will result in a decreasing σt\sigma_{t} in the backward generative process. We start with a Gaussian random yy. In the early stages when σ\sigma is large, differences in distances ‖y−si‖\|y-s_{i}\| in (11) make negligible effect. Therefore, the output sequence converges to the mean value of the nearby vectors.

When σ\sigma gets small, differences in distances are promoted in (11) and the results quickly converge to the nearest signal, with some contribution from nearby signals. This is a sparse representation, as the simulation of (11) for a simple setup shows in Fig. 2. In this simulation, σt=0.96​σt+1\sigma_{t}=0.96\,\sigma_{t+1} is a decreasing sequence which is decreased gradually from 11 to 0.370.37 in the reverse process. Although, as we will discuss later in Sec. II-E, it seems that σ\sigma does not converge to zero because of practical limitations in the denoising neural network design and training.

The diffusion process initially converges to the mean point of the distribution. Then in the long term, it converges to the vicinity of some dense region of the dataset. This is in accordance with the fact that the diffusion process simulates the real data distribution as reflected in the dataset.

Knowledge about the convergence of the diffusion process enables us to circumvent the whole convergence process by directly using a convex combination of some nearby dataset signals as output of the generation process. This basically reduces not only the lengthy generation process, but also the very lengthy training process of denoising neural network.

II-E Finite Training

In this manuscript, we have generally assumed infinite training of denoising network. This enables the training process to cover every yy point in the space. It also enables the network through many gradient descent updates, to achieve the ground-truth closed form solution of the training process optimization.

This is rarely the case in practice. Due to limited resources like silicon, time, and energy, we usually perform minimal training up to the point of acceptable network performance. Here we discuss the effect of non-asymptotic training on the network behavior.

First, the network does not experience every yy in the space during the training. It only sees finite adjacent inputs {y1,…,yk}\{y_{1},\ldots,y_{k}\}. It learns the correct direction of move from these inputs to outputs {o1,…,ok}\{o_{1},\ldots,o_{k}\}, respectively. Any new yy is sorrounded by multiple seen points. Since the network is a continouos system, we expect an average behavior in yy if the training set is sufficiently dense. This guides yy to regions of higher density which are experienced more in training and gradually we experience better refinements.

oi=𝔻⁡(𝔼⁡(yi))\displaystyle o_{i}=\mathbb{D}\left(\mathbb{E}\left(y_{i}\right)\right) (16)
y=∑i=1kαi​yi→𝔻⁡(𝔼⁡(y))=∑i=1kβi​oi\displaystyle y=\sum_{i=1}^{k}\alpha_{i}y_{i}\quad\to\quad\mathbb{D}\left(\mathbb{E}\left(y\right)\right)=\sum_{i=1}^{k}\beta_{i}o_{i} (17)

where {αi}\{\alpha_{i}\} and {βi}\{\beta_{i}\} are two convex coefficient sets, 𝔼\mathbb{E} is the encoder function and 𝔻\mathbb{D} is the decoder.

The second consequence of finite training is that the network does not learn the optimal solution of the optimization in (7) for each training input yiy_{i}. We can model this training error as an additive noise in the output of the network. The same is true for the unseen input error and we can model the error in 𝔻⁡(𝔼⁡(yi))\mathbb{D}\left(\mathbb{E}\left(y_{i}\right)\right) also as an additive noise.

When noise is added to the output of the denoiser network in a diffusion system, the output will no longer be a sparse attention combination even when σ\sigma of the diffusion system is very small. This is in fact a new additive noise in the backward process:

zt−1=zt+nt′′⇒xt=αt​x0+1−αt​nt′−nt′′z_{t-1}=z_{t}+n^{\prime\prime}_{t}\quad\Rightarrow\quad x_{t}=\sqrt{\alpha^{t}}x_{0}+\sqrt{1-\alpha^{t}}\,n^{\prime}_{t}-n^{\prime\prime}_{t} (18)

where nt′′n^{\prime\prime}_{t} is the new noise term indicating errors emanating from finite training process.

Therefore, the overall effect of finite training can be modeled as a higher value of σ\sigma which increases the number of intermittent signals. This means that although in the diffusion framework, theoretically σ\sigma gradually converges to zero, practical limitations leads to a stable higher limiting value for σ\sigma. This prevents the optimization process to achieve a very sparse solution and the final output will be a convex combination of many signals.

III Auto Encoders

In practice, diffusion systems use auto-encoders as both the outer encoder layer that the diffusion system operates inside [9], and the inner recurrent denoising network. Therefore, it is also crucial to understand working principles of AE to reduce computational complexity of the diffusion systems. In this section we use a different perspective to explain AE and attention system similarity.

III-A Auto-encoders are attentions

As we showed in Sec. II, there is a deep relationship between diffusion generative systems and attention mechanism. Here, we also show a strong resemblance between AEs, attention systems, and diffusion systems. This will present a great unity between three of the most important artificial intelligence systems currently in use.

First of all, notice that the derivations in Sec. II-C is not limited to a diffusion system. In fact, it is a widespread setting in unsupervised learning systems. Assume we have an AE network which as usual use MSE loss function to reconstruct input images or other signals in an unsupervised manner. Regardless of the AE’s inner structure, we assume perfect training of the system. Most of the times, an augmentation technique of adding Gaussian noise to limited dataset signals is used for better training and performance. Now, we can see that the mathematical setup is identical to the optimization in (7). Therefore, the same solution as in (11) applies to the problem. This certifies that such an AE system is equivalent to an attention system in (12) with fixed σ\sigma in the infinite training setup. Attention parameter σ\sigma in (11) can be tuned according to the structural parameters of the AE for a satisfactory equivalence.

The equivalence between AE and attention is double-sided. It means that we can in principle replace current attention systems with AEs. A good example for such a replacement is the linear attention [15], which is known to be a finite-state attention mechanism with compressed rather than full state. This is reminiscent of recursive auto-encoder structure [16], especially when we consider the causal implementation of the linear attention.

III-B Auto-encoders and Manifolds

m k m v v c inputcodeoutput E Encoder D Decoder
Fig. 3: An auto-encoder with mm-dimensional input/output vectors and kk-dimensional code layer.

Assume an auto-encoder with mm-dimensional input-output and kk-dimensional code layer as in Fig. 3. Manifold hypothesis assumes that data from the dataset is located on a high-dimensional manifold in input space [10]. It is well known that the autoencoder estimates the data manifold during the training phase [11]. We will show that AE projects the mm-dimensional input signal on to a kk-dimensional manifold which represents locus of data in the dataset and their meaningful augmentations in the space.

Assume that the AE is fully trained with the dataset and therefore, it can now regenerate signal ss in the output approximately. In the input space, we add a differential length vector v1v_{1} to ss. With high probability, no ReLU-like switch in the network is on its change position. Therefore, for any differential change in the input, the system is linear. Then, changing the input ss with v1v_{1} results in a differential c1c_{1} vector change in the code layer:

𝔼⁡(s+v1)=𝔼⁡(s)+c1\mathbb{E}\left(s+v_{1}\right)=\mathbb{E}(s)+c_{1} (19)

Continue these changes with orthogonal differential vectors viv_{i} that results each in code layer change of cic_{i}. In the k+1k+1’th step, the set of {c1,…,ck+1}\{c_{1},\ldots,c_{k+1}\} is a linearly dependent set since the code layer vector space is kk-dimensional. Therefore, a linear combination of these vectors sums to zero:

∑i=1k+1αi​ci=0\sum_{i=1}^{k+1}\alpha_{i}c_{i}=0 (20)

The same set of αi\alpha_{i} coefficients can be used to find a direction in the input space that keeps the code layer, and therefore the output constant:

𝔻⁡(𝔼⁡(s+∑i=1k+1αi​vi))=𝔻⁡(𝔼⁡(s)+‌​∑i=1k+1αi​ci)=s\mathbb{D}\left(\mathbb{E}\left(s+\sum_{i=1}^{k+1}\alpha_{i}v_{i}\right)\right)=\mathbb{D}\left(\mathbb{E}(s)+‌\sum_{i=1}^{k+1}\alpha_{i}c_{i}\right)=s (21)

Therefore, we showed that any k+1k+1 vectors viv_{i} in the input space can form a direction with null effect in the output. This direction is:

w=∑i=1k+1αi​viw=\sum_{i=1}^{k+1}\alpha_{i}v_{i} (22)

The local null space of the encoder at ss is defined as the local Affine vector space of all such ww vectors:

ℕ(s):={‌s+w|𝔼(s+w)=𝔼(s),∥w∥<‌δ}\mathbb{N}(s):=\left\{\>‌s+w\;|\;\mathbb{E}(s+w)=\mathbb{E}(s)\;,\;\|w\|<‌\delta\right\} (23)

where δ\delta is a small constant. The null space ℕ⁡(s)\mathbb{N}(s) is local to ss and is effective just in a δ\delta-neighborhood of ss. At most kk input vectors viv_{i} can be chosen that does not lie in this local null space. Therefore:

dim​(ℕ​(s))=m−k\textrm{dim}(\mathbb{N}(s))=m-k (24)

Suppose 𝕄⁡(s)\mathbb{M}(s) is the kk-dimensional local Affine space orthogonal to the local null space:

𝕄(s):={s+v|\displaystyle\mathbb{M}(s):=\{\>s+v\;| ‌⁡<v,w>=0:‌\displaystyle‌<v,w>=0\;:‌
∀w:‌s+w∈ℕ(s),∥v∥<δ}\displaystyle\forall w:‌s+w\in\mathbb{N}(s)\;,\;\|v\|<\delta\,\} (25)

In the vicinity of ss, any displacement along the null space ℕ⁡(s)\mathbb{N}(s) does not change the values of the code layer and therefore, the output. But displacement along manifold 𝕄⁡(s)\mathbb{M}(s) translates to changes in code layer and output. Therefore, the AE will tune its null space ℕ⁡(si)\mathbb{N}(s_{i}) such that it direct farthest from the neighbors of sis_{i}. This will set the estimated data subspace 𝕄⁡(si)\mathbb{M}(s_{i}) in the nearest position that best cover the neighbors of sis_{i} and orthogonal to the local null space. This is in accordance with autoencoder estimating a manifold with good fit to neighboring data points [11] with extensive training on sufficiently dense datasets.

𝔽\mathbb{F}sis_{i}yy×\times
Fig. 4: The dataset 𝒟\mathcal{D} is assumed dense enough that the data manifold 𝔽\mathbb{F} can be followed by the auto encoder after sufficient training. Input yy is projected onto the data manifold in a curved path following the local null spaces ℕ\mathbb{N} and eventually arrives at 𝔽\mathbb{F} in a orthogonal direction.

Here we assume that the dataset is dense enough that the continuity of the data manifold can be tracked by the AE as in Fig. 4 . We also assume that the AE does not have a specific structure (like convolutional or any other weight sharing scheme). This helps the AE to follow any direction in the data space without any bias. It is well-known that networks like convolutional networks better estimate the data manifold. This is because some of the data structure, like image content translation invariance is embedded in such networks. This is equivalent to more data points which helps the network better estimate the data manifold.

Any input yy outside the manifold defines a specific local null space ℕ⁡(y)\mathbb{N}(y) which aggregates directions of moves that does not change the code layer and output layers values. Therefore, AE moves yy in a curved path along local null spaces to finally reach the AE estimated data manifold 𝔽\mathbb{F} in hh. Since ℕ⁡(h)⟂𝕄⁡(h)\mathbb{N}(h)\perp\mathbb{M}(h) we reach to the AE manifold perpendicularly in the last touch as is shown in Fig. 4.

III-C Attention and manifold

Here, we show that the attention mechanism also projects data on a low dimensional manifold. This is in-line with our core insight that basically these three structures:‌ AE, attention, and diffusion systems are similar.

Assume that yy is near the data manifold. Dataset satisfies conditions that we can use (11) in place of the better-known attention formula in (12). Also, assume that σ\sigma is small enough that only kk nearest signals:

Ai:={sij}j=1kA_{i}:=\{s_{i_{j}}\}_{j=1}^{k} (26)

are effective in the attention sum in (11). Then, the attention coefficients are:

pi≃exp⁡{−‖y−si‖22​σ2}∑s∈Aiexp⁡{−‖y−s‖22​σ2}=:pi​(y,Ai)p_{i}\simeq\frac{\exp\left\{-\frac{\|y-s_{i}\|^{2}}{2\sigma^{2}}\right\}}{\sum_{s\in A_{i}}\exp\left\{-\frac{\|y-s\|^{2}}{2\sigma^{2}}\right\}}=:p_{i}(y,A_{i}) (27)

The adjacency set AiA_{i} forms a k−1k-1-dimensional Affine subspace. Assume that the orthogonal projection of yy onto this subspace is hh as in Fig. 5. Then:

‖y−si‖2=‖y−h‖2+‖h−si‖2\|y-s_{i}\|^{2}=\|y-h\|^{2}+\|h-s_{i}\|^{2} (28)

Then ‖y−h‖2\|y-h\|^{2} is shared in numerator and denominator of (27), eliminating it:

pi=exp⁡{−‖h−si‖22​σ2}∑s∈Aiexp⁡{−‖h−s‖22​σ2}=pi​(h,Ai)p_{i}=\frac{\exp\left\{-\frac{\|h-s_{i}\|^{2}}{2\sigma^{2}}\right\}}{\sum_{s\in A_{i}}\exp\left\{-\frac{\|h-s\|^{2}}{2\sigma^{2}}\right\}}=p_{i}(h,A_{i}) (29)

Output of hh from the attention mechanism is a convex combination of s∈Ais\in A_{i} and therefore is on their Affine subspace. Also, the fact that the output of yy is the same as the output for hh shows that there is a local Affine null space orthogonal to the local affine space of nearby signals in AiA_{i} that does not change the output of the attention.

𝔽\mathbb{F}si1s_{i_{1}}si2s_{i_{2}}si3s_{i_{3}}si4s_{i_{4}}yy×\timeshh∗*
Fig. 5: Geometry of attention mechanism. Point yy with a small offset from the data manifold 𝔽\mathbb{F} is projected orthogonally to the Affine subspace spanned by nearby signals on the data manifold by attention mechanism.

III-D Auto-encoders resemble attention

The attention mechanism when act on a point yy near the data manifold 𝔽\mathbb{F}, orthogonally projects it on the manifold at point hh, and then transform hh to another point in the convex hull of signals in AiA_{i} in the output. Overall, we can conclude that the attention mechanism projects any point yy near the data manifold to some point on the manifold too and this is what also auto encoders do.

The attention mechanism depends on the parameter σ\sigma, while the AE is characterized by its code layer dimension kk. The equivalence between attention and AE requires fine tuning these parameters. As stated before, setting a small σ\sigma is required in order that approximately k+1k+1 nearby signals are effective in the attention. This will result in the estimation of a kk-dimensional manifold for data similar to an AE with kk-dimensional code layer. This is not an exact equivalence because the required σ\sigma parameter may not be constant in different parts of the input space.

In a conventional diffusion system, we have an outer autoencoder which compresses data to a lower dimension. Then a recursive inner autoencoder denoises an input noise. We have shown that not only the high level diffusion framework is equivalent to the attention mechanism, but also both autoencoders in the network structure realization perform in a similar way to the attention mechanism. Therefore, we can alternate between these three structures based on the situation.

If attention and auto encoder structures are basically similar, what is their difference?‌ The difference lies in where the computational load is concentrated. In an auto encoder, we use vast amount of data in training phase while the test phase is as simple as one forward calculation of the network. The attention mechanism is reverse. Calculating attention does not require any training while all the computational load is concentrated in the test time where the forward pass of the network requires vast amount of product calculations with a long context of vectors. Therefore, we may be able to reduce the complex training phase of the auto encoder and diffusion systems by using some simpler form of the attention mechanism.

IV Tree Cross Attention

The problem with the attention mechanism is that it involves calculating the dot product of the input yy with every image in the dataset. These image datasets are usually very large in industry-level network training. Since usually σ\sigma parameter of attention is small, it suffices to calculate the attention just for nearby signals. But finding the nearbies requires calculating the distances to all images in the dataset which is equivalent to inner product computation.

Here, a data structure can lead to considerable reduction in the computational complexity. Calculating the dot product of yy with every images in the dataset is of 𝒪⁡(|𝒟|)\mathcal{O}(|\mathcal{D}|) complexity. But since there is a great deal of redundancy in this calculation, we can save some information for reuse.

Data can be order in a tree structure based on similarity [17]. Each node in the tree is an image and any new image is tested with every nodes in each tree level to see with which branch it is more similar. If there was a similar branch, the new image goes down to that branch’s next level children. If there were no similar branch, make a new branch in the current level.

In this way, a balanced tree can be constructed in which similar and nearby images are clustered in a branch. Then for each new input signal yy, we need a few inner product calculations of 𝒪⁡(log⁡|𝒟|)\mathcal{O}(\log|\mathcal{D}|) which is the tree depth to find its neighbors for efficient attention calculations. This method is called tree cross attention and is presented in the literature [18], for the situations where base attention vectors are constant. Therefore, attention calculation is not restrictive in an image generation scenario.

V Proposed Generation Algorithm

Refer to captionReal imageEncoder𝔼\mathbb{E}Select querry Find neighborsAttention Interpolationwi∝e−12​σ2​‖𝐪−𝐬i‖2w_{i}\propto e^{-\frac{1}{2\sigma^{2}}\|\mathbf{q}-\mathbf{s}_{i}\|^{2}}𝐳=∑i=1Kwi​𝐬i\mathbf{z}=\sum_{i=1}^{K}w_{i}\mathbf{s}_{i}Normalize𝐳‖𝐳‖2​r\frac{\mathbf{z}}{\|\mathbf{z}\|_{2}}r+\boldsymbol{+}Add noise ϵ\boldsymbol{\epsilon}Decoder𝔻\mathbb{D}Refer to captionGenerated image
Fig. 6: Proposed attention-based generative architecture.

Theoretical equivalence of diffusion and attention systems helps to reduce the computational complexity of image generation. We have seen in Sec. II-D that any noise input to a diffusion system experiences a long road to its final position in latent space.

Since the solution to diffusion is an approximately sparse one (Fig. 2), we can start image synthesize from one of dataset images. Since attention to latent signals is approximately constant in the neighborhood of the image latent, we can calculate the attention of the latent signal to other neighbor signals.

This attention calculation helps the resulting signal to reside on or very near the data manifold since it is equivalent to passing the signal once through the hypothetical diffusion trained AE denoiser. Therefore, we can pass the resulting latent through the decoder and reach to a high quality innovative image. Since there may be slight disposition from the data manifold, we can pass the resulting image a few times through the AE to better fit it to the data manifold.

Therefore, there will be no need to train a denoiser auto encoder as the working engine of the diffusion system. We just need to train an outer AE which projects the input images to the low-dimensional latent space. Then in the generation phase, we just rely on attention calculation near the signal latents. Therefore, we not only reduced the training phase calculations, but also in the test phase, we reduce the long journey of the input noise to the distribution mean value and then to nearby of one of the signals. We just perform the final step with calculating the attention of the signal to its neighbbors including itself. The block diagram of the proposed image generation algorithm is depicted in Fig. 6.

The generation process is done in the learned latent space of trained VAE. As described in Alg. 1, for a random latent query 𝐪\mathbf{q}, we find its KK nearest neighbors {𝐬k}k=1K\{\mathbf{s}_{k}\}_{k=1}^{K}. New latent representation is generated through attention interpolation between the query and its neighbors. The generated latent representation is then passed through the VAE decoder to obtain a new image. Since nearby latent representations correspond to semantically related data, their interpolation produce meaningful intermediate representations.

Although using plain attention between querry and its neighbor images produce meaningful images, we use a fine grain spatial attention mechanism to get better quality. Spatial attention is calculated based on the distance of each pixel depth vector in querry to its corresponding pixel depth vector in a neighbor image as presented in Alg. 2. This generates a spatial attention map that determines the contribution of each pixel to the generated sample. Example spatial attention maps in Fig. 7 illustrate how the proposed method selectively exploits spatial information from neighbors during latent-space sampling. Note that spatial attention resembles the patch similarity selection mechanism that is developed as the theoretical solution of a convolutional diffusion network in [8].

Algorithm 1 Attention-Based Image Generation
1: Input: Latents {𝐬i}i=1N\{\mathbf{s}_{i}\}_{i=1}^{N}, query 𝐪\mathbf{q}, number of neighbors KK, attention parameter σ\sigma, noise scale η\eta
2: Output: Generated image 𝐱\mathbf{x}
3: r←1N​∑i=1N‖𝐬i‖2r\leftarrow\frac{1}{N}\sum_{i=1}^{N}\|\mathbf{s}_{i}\|_{2}
4: di←‖𝐪−𝐬i‖2d_{i}\leftarrow\|\mathbf{q}-\mathbf{s}_{i}\|_{2}  ∀i\forall i
5: 𝒮←{𝐬i∣i∈TopK⁡(−di)}\mathcal{S}\leftarrow\{\mathbf{s}_{i}\mid i\in\operatorname{TopK}(-d_{i})\}
6: 𝐖←SpatialAttention⁡(𝐪,𝒮,σ)\mathbf{W}\leftarrow\operatorname{SpatialAttention}(\mathbf{q},\mathcal{S},\sigma)
7: 𝐳←∑𝐬i∈𝒮𝐖i⊙𝐬i\mathbf{z}\leftarrow\sum_{\mathbf{s}_{i}\in\mathcal{S}}\mathbf{W}_{i}\odot\mathbf{s}_{i}
8: 𝐳←r​𝐳‖𝐳‖2+ϵ\mathbf{z}\leftarrow r\frac{\mathbf{z}}{\|\mathbf{z}\|_{2}+\epsilon}
9: 𝐳←𝐳+η​ϵ,ϵ∼𝒩⁡(𝟎,𝐈)\mathbf{z}\leftarrow\mathbf{z}+\eta\boldsymbol{\epsilon},\quad\boldsymbol{\epsilon}\sim\mathcal{N}(\mathbf{0},\mathbf{I})
10: 𝐳q←Quantize⁡(𝐳)\mathbf{z}_{q}\leftarrow\operatorname{Quantize}(\mathbf{z})
11: 𝐱←𝔻⁡(𝐳q)\mathbf{x}\leftarrow\mathbb{D}(\mathbf{z}_{q})
12: for j=1,…,3j=1,\ldots,3 do
13:   𝐱←AE⁡(𝐱)\mathbf{x}\leftarrow\operatorname{AE}\left(\mathbf{x}\right)
14: end for
15: return 𝐱\mathbf{x}
Algorithm 2 Spatial Attention
1: Input: Query 𝐪∈ℝD×H×W\mathbf{q}\in\mathbb{R}^{D\times H\times W}, neighbor latents 𝒮={𝐬i}i=1K\mathcal{S}=\{\mathbf{s}_{i}\}_{i=1}^{K}, scale parameter σ\sigma
2: Output: Attention weights 𝐖∈ℝK×H×W\mathbf{W}\in\mathbb{R}^{K\times H\times W}
3: for i=1,…,Ki=1,\ldots,K do
4:   𝐬←𝐬i\mathbf{s}\leftarrow\mathbf{s}_{i}
5:   for all j,kj,k do
6:    di​j​k←‖𝐪:j​k−𝐬:j​k‖2d_{ijk}\leftarrow\left\|\mathbf{q}_{:jk}-\mathbf{s}_{:jk}\right\|_{2}
7:    pi​j​k←−di​j​k2σ2+ϵp_{ijk}\leftarrow-\dfrac{d_{ijk}^{2}}{\sigma^{2}+\epsilon}
8:    wi​j​k←softmaxi⁡(pi​j​k)w_{ijk}\leftarrow\operatorname{softmax}_{i}(p_{ijk})
9:   end for
10: end for
11: return 𝐖←{wi​j​k}\mathbf{W}\leftarrow\{w_{ijk}\}
NN1 NN2 NN3
Refer to caption Refer to caption Refer to caption
Fig. 7: Spatial attention weight maps obtained for the top-33 nearest neighbors of the query on the FFHQ dataset. Each map illustrates the spatial contribution of the corresponding neighbor region to the generated latent representation.

VI Simulation Results

We train the model using a P100 GPU on the MNIST, Fashion-MNIST, FFHQ, and CelebA datasets with an image size of 28×2828\times 28 for MNIST and Fashion-MNIST, 64×6464\times 64 for FFHQ and 128×128128\times 128 for CelebA. The Adam optimizer is used with β1=0.9\beta_{1}=0.9 and β2=0.99\beta_{2}=0.99, and a learning rate of 1×10−41\times 10^{-4}. For outer VAE training, we use a combination of L1L_{1} loss, perceptual loss based on VGG network and an adversarial (GAN) loss of StyleGAN discriminator with the following architecture:

Lrec=L1+Lperceptual+LadvL_{\mathrm{rec}}=L_{1}+L_{\mathrm{perceptual}}+L_{\mathrm{adv}} (30)

In our approach, generation is performed by local interpolation which is substantially faster than iterative reverse diffusion. The diffusion method requires 0.170.17 second versus the proposed method just 0.020.02 second to generate an image on the FFHQ dataset, hence a 8.5×8.5\times speedup.

Example generated images are presented in Fig. 8 for FFHQ, MNIST, and Fashion-MNIST datasets. The first column from left in each dataset is the chosen query, the second is the generated image, and the last three are nearest neighbors, respectively. Generated images are quite natural and innovative with overall similarity to the query image while they receive special features from neighbors.

 

FFHQ
 

Refer to caption Refer to caption Refer to caption Refer to caption Refer to caption
Refer to caption Refer to caption Refer to caption Refer to caption Refer to caption
Refer to caption Refer to caption Refer to caption Refer to caption Refer to caption
 
 

MNIST
 

Refer to caption Refer to caption Refer to caption Refer to caption Refer to caption
Refer to caption Refer to caption Refer to caption Refer to caption Refer to caption
Refer to caption Refer to caption Refer to caption Refer to caption Refer to caption
 
 

Fashion MNIST
 

Refer to caption Refer to caption Refer to caption Refer to caption Refer to caption
Refer to caption Refer to caption Refer to caption Refer to caption Refer to caption
Refer to caption Refer to caption Refer to caption Refer to caption Refer to caption
 
Fig. 8: Qualitative comparison on FFHQ, MNIST, and Fashion-MNIST. From left to right, each row presents the query image, the generated image, and the Top-3 nearest neighbors retrieved from the latent space.

More generated sample images is shown in Fig. 9. It can be seen that most of the images are natural and of good quality. The higher resolution of 128×\times128 with CelebA dataset is also examined in Fig. 10.

Refer to caption Refer to caption Refer to caption Refer to caption Refer to caption Refer to caption Refer to caption Refer to caption Refer to caption Refer to caption
Refer to caption Refer to caption Refer to caption Refer to caption Refer to caption Refer to caption Refer to caption Refer to caption Refer to caption Refer to caption
Refer to caption Refer to caption Refer to caption Refer to caption Refer to caption Refer to caption Refer to caption Refer to caption Refer to caption Refer to caption
Fig. 9: Generated samples on FFHQ 64×\times64
Refer to caption Refer to caption Refer to caption Refer to caption Refer to caption Refer to caption
Refer to caption Refer to caption Refer to caption Refer to caption Refer to caption Refer to caption
Fig. 10: Generated samples on CelebA 128×\times128

Baseline diffusion model configurations are listed in Table I. A different setting is used for each dataset used for comparison. Similar settings of the proposed method is also listed in Table II. Table III compares FID of the proposed method versus the baseline diffusion model. While our proposed model has much less parameters, it achieves better FID in MNIST and comparable FID in Fashio MNIST and CelebA, while the performance is lower in FFHQ dataset. Table IV shows an optimization study which shows that lower σ\sigma parameter of the attention mechanism will result in better FID metric.

TABLE I: Training configurations of the baseline diffusion models.
MNIST FashionMNIST FFHQ CelebA
ff 1 1 4 4
Diffusion steps 1000 1000 1000 1000
Noise schedule linear linear linear linear
Backbone CNN CNN Transformer CNN
#Params 15.7M 15.7M 130M 47M
Attention resolution 7 7 – 16, 8
#Heads 4 4 12 1
Batch size 256 256 256 32
Iterations 30k 36k 134k 51k
Learning rate 10−410^{-4} 10−410^{-4} 10−410^{-4} 10−410^{-4}
TABLE II: Architecture and training configurations of the proposed latent-space models.
MNIST FashionMNIST FFHQ CelebA
ff 4 4 8 8
zz-shape 3×7×73\times 7\times 7 3×7×73\times 7\times 7 16×8×816\times 8\times 8 4×16×164\times 16\times 16
Latent Signals 1024 1024 2048 -
Backbone CNN CNN CNN CNN
#Params 2.95M 2.95M 64.8M 53M
Attention resolution – – Encoder: 8 Decoder: 8, 16 –
#Heads – – 1 –
Batch size 256 256 256 32
Iterations 13.6k 31.3k 44k 51k
Learning rate 10−410^{-4} 10−410^{-4} 10−410^{-4} 10−410^{-4}
TABLE III: Quantitative comparison of the baseline and proposed methods in terms of FID.
Dataset Method FID
MNIST Diffusion 19.30
Ours 8.60
FMNIST Diffusion 13.00
Ours 14.60
FFHQ Diffusion 8.60
Ours 16.58
CelebA Diffusion 30.49
Ours 32.15
TABLE IV: Ablation study of the proposed sampling method under different σ\sigma and Top-KK settings on the FFHQ dataset.
σ\sigma Top-KK Noise scale FID
30 3 0.01 27.43
30 5 0.01 34.33
0.01 3 0.01 16.69
0.001 3 0.01 16.58

References

  • [1] A. Vaswani, et al., “Attention is all you need,” in Proc. Ann. Conf. Neur. Inf. Process. Syst., Long Beach, CA, USA, pp. 5998–6008, 2017.
  • [2] W. Sun, J. Hu, Y. Zhou, et al., “Speed always wins: A survey on efficient architectures for large language models,” arXiv:2508.09834, 2025.
  • [3] A. Dosovitskiy, et al. “An image is worth 16x16 words: Transformers for image recognition at scale,” arXiv preprint arXiv:2010.11929. 2020.
  • [4] J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli, “Deep unsupervised learning using nonequilibrium thermodynamics,” In International Conference on Machine Learning, pages 2256–2265, PMLR, 2015.
  • [5] W. Peebles, and S. Xie, “Scalable diffusion models with transformers,” In Proceedings of the IEEE/CVF international conference on computer vision, pp. 4195-4205. 2023.
  • [6] H. V. Poor, An introduction to signal detection and estimation, Springer, 2013.
  • [7] J. Henderson and F. Fehr, “A VAE for Transformers with Nonparametric Variational Information Bottleneck,” in Proceedings of International Conference on Machine Learning ICLR, 2023.
  • [8] M. Kamb and S. Ganguli, “An analytic theory of creativity in convolutional diffusion models,” In Proceedings of the 42nd International Conference on Machine Learning, Vancouver, Canada, 2025.
  • [9] R. Rombach, et al. “High-resolution image synthesis with latent diffusion models,” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022.
  • [10] J. Braunsmann, M. Rajkovic , M Rumpf, and B. Wirth, “Convergent autoencoder approximation of low bending and low distortion manifold embeddings,” Mathematical Modelling and Numerical Analysis, vol. 58, pp. 335-361, 2024.
  • [11] Y. Lee, “A geometric perspective on autoencoders,” available online:‌ https://arxiv.org/abs/2309.08247
  • [12] J. Song, C. Meng, and S. Ermon. “Denoising diffusion implicit models,” arXiv preprint arXiv:2010.02502, 2020
  • [13] C. Luo, “Understanding diffusion models: A unified perspective,” arXiv preprint arXiv:2208.11970, 2022
  • [14] S Boyd and L. Vandenberghe, Convex optimization, Cambridge Univ. Press, 2004.
  • [15] A. Katharopoulos, et al., “Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention,” in Proceedings of the 37th International Conference on Machine Learning, pp. 5156–5165, 2020.
  • [16] J. Zhang, R. Su, C. Liu, et al, “A Survey of Efficient Attention Methods: Hardware-efficient, Sparse, Compact, and Linear Attention,” 2025.
  • [17] R. Sedgewick and K. Wayne, Algorithms, Addison-Wesley, 4’th ed., 2011.
  • [18] L. Feng, et al. “Tree cross attention,” arXiv preprint, arXiv:2309.17388, 2023.