Denoising Diffusion Generative Models Secretly Calculate Attentions
Abstract
Denoising diffusion models are the dominant architecture for image generation, whereas most natural language generation and modeling are primarily handled by well-known transformer architectures employing attention mechanism. Here, we show that diffusion models also inherently use an attention mechanism very similar to that of transformers. Therefore, attention emerges as a universal machine learning principle, based on a general training objective. We also show similarities in basic functional principle of auto-encoders and attention-based models. These equivalences allows us to interchange these designs based on practical requirements. As an example, we can reformulate the diffusion framework to reduce the lengthy training process and computation-intensive image generation. Using this approach, a simplified algorithm is proposed for image generation which is based on attention mechanism. Results show that the attention-based implementation achieves comparable performance with significantly less effort and computational resources.
Index Terms:
Diffusion model, attention mechanism, autoencoder, data manifold, latent space, interpolation.I Introduction
Artificial intelligence (AI) encompasses a multitude of proposed architectures for different tasks. Throughout the history of AI, it was a surprise that gradually efficient unified system architectures replaced this multitude of systems for various applications. An important turning point was the advent of the dot-product-based attention mechanism [1]. Transformers were first proposed for tasks of natural language processing (NLP) such as question-answering and chatbots [2]. Soon, it was realized that transformers also have stellar performance in vision applications with the advent of the vision transformer (ViT) model [3].
In image generation, diffusion systems are arguably the best option [4]. Yet, for the best performance they are used in tandem with transformer-based components. Specifically, diffusion principle is used to determine the loss function and high-level train and test schemes, while in network level, an auto-encoder is trained which in turn uses attention mechanism and transformer architecture as its main component [5].
Here, we show that the diffusion principle is equivalent to the attention mechanism. Diffusion denoising probabilistic generative model (DDPM) [4], trains a neural network to reconstruct the original image from its noisy copies. Then, starting from pure noise, the system can generate a realistic image based on the learned distribution of the dataset.
Detecting the presence of a signal from its noisy version is a classic statistical problem [6]. Specifically, when the noise is Gaussian, the solution is an inner-product detector. While in the context of machine learning theory, we train a neural network to solve this problem, we will show that this problem enjoys a closed-form solution which is basically the attention mechanism. This not only shows that the diffusion system is identical to the attention mechanism, but also provides a previously known statistical basis for the very attention mechanism as the closed form solution to the signal recovery problem in the Gaussian noise [7].
The closed form solution of the diffusion system was first derived in [8]. The authors calculate optimum score function of a diffusion system for a finite dataset. However, they did not notice the relation of their result to the well-known attention mechanism. Then [8] extends the result to include additional structures like CNN network which is used in the DDPM implementation and develops a creativity theory for CNN diffusion systems based on locality and positional invariance properties. Here we use a different approach based on the diffusion loss function and convex optimization to reach the same closed form solution while extending the analysis to include resemblance to transformer and autoencoders. The equivalence between diffusion and transformer principles have broad theoretical and practical consequences in machine learning theory.
It is not sufficient to understand the core diffusion principle in the forms of what we investigate as closed-form one-step solution or convergence trajectory. In practice, diffusion systems are implemented in the inner latent space of a Variational Auto-Encoder (VAE) [9]. Also, in practice the main neural network that performs denoising in the diffusion system is an AE. We argue that the same equivalence exists between attention and AE, in that both systems eventually project a random input on a learned data manifold [10, 11]. Therefore, we reach to an equivalence between three of the most important modern AI systems, namely: attention, diffusion , and AE. This enables us to interchange these systems depending on the situation.
DDPM diffusion system suffers from heavy computations in both training and generation which requires running an encoder hundreds of times on a random input. Some efforts have been made to reduce computational complexity specially in generation. A trend in the literatur is to reduce the steps for image generation using non-Markovian statistical models as in Diffusion Denoising Implicit Model (DDIM) [12].
Regarding the taining, DDIM is equivalent to DDPM and very computationally complex. It requires adding noise for hundreds of times to each image in dataset to teach network to denoise images with varying degree of noise effectively. Therefore, reducing computational complexity in the training phase is highly important. Here, we calculate a closed form denoiser which reduces the need for training a network on the dataset.
As a result of the equivalence of diffusion and attention, we conclude that deliberate interpolation between neighbor real images in the latent space results is an approximately realistic image. This interpolation can be refined and projected on the real image manifold by a limited recursive application of the outer AE. Overall, the proposed system can generate a realistic image without training an inner AE network, and just by using the dataset images. The results show good quality with considerably lower computations.
The manuscript is organized as follows: The main result and discussion is presented in Sec. II, where a background review, discussion on equivalence between diffusion and attention, and convergence topics are covered. AutoEncoder equivalence with attention and projection on data manifold are addressed in Sec. III. A data structure for fast attention calculation is discussed in Sec. IV. The proposed low complexity generative network is demonstrated in Sec. V, while simulation results are exhibited in Sec. VI.
II Diffusion Systems Calculate Attention
II-A Background
A diffusion system is trained to denoise noisy images in a recurssive manner. In the forward process, Gaussian noise is gradually added to the true image in hundreds or thousands of steps [13]:
for . Therefore, each image in the forward process can be written based on the true initial image as:
| (1) |
This eventually leads to a noise-only image at the final stage of the forward process since . We train a neural network to recover the true image from each noisy image . Theoretically, this results in a neural network that can produce a realistic image from a pure noise and based on the true distribution of the data which is represented by the dataset used for training.
II-B Prior Art
In practice, we train the denoiser neural network on a finite dataset . This finite dataset training situation was first analyzed in [8]. Adding Gaussian noise incurs a mixture of Gaussian distribution on dataset points in signal space:
| (2) |
Then [8] uses a flow matching continuous-time model which is solved backward in time for signal generation:
| (3) |
in which is a constant and is the score function. The score function of the mixture distribution in (2) is then [8]:
| (4) |
| (5) |
Then [8] continues the analysis to develop a theory for combinatorial creativity of diffusion models when the denoiser is a CNN network.
II-C Diffusion is Attention
Our analysis starts from the same finite dataset training assumption and then proceeds in a different direction to find the relation of the diffusion model to transformers and autoencoders as current dominant models in machine learning research.
When we train the network on the finite dataset , the network learns to recover each from its noisy versions. We guess that when sampling the system for a new signal, although we use a random noise as the input, the system still outputs just an interpolation of the signals it had seen in the dataset since it has not seen anything else.
For signal and Gaussian noise , DDPM minimizes Mean-Squared Error (MSE) loss function:
| (6) |
in which is the real signal, its estimate, time index, and is the neural network’s parameters. If we infinitely train the network with fixed and , it learns to convert input to output .
But we do not train the network just for one sample image. Adding Gaussian noise to any signal for sufficiently many times, can reach any point in the signal space many times. A geometrical scheme of the problem is depicted in Fig. 1. Reaching any fixed point from various signals occurs in accordance with the Gaussian distribution of the noise term . Therefore, if we fix , the system is trained with an average loss function:
| (7) | |||
For fixed , (7) is a Quadratic Programming (QP) convex optimization problem:
| (8) | ||||
| (9) |
Where (9) is a conic criterion for the strictly convex optimization in (8). The optimal solution for the above optimization problem can be derived using gradients [14]:
| (10) |
| (11) |
Assume that is sufficiently small. Then just those signals that are in the neighborhood of are effective in the convex combination in (11). Therefore, if is large enough and the space is sufficiently populated with data points, we can safely assume that ’s are approximately constant. Other situations also may hold the condition of approximately constant like in image dataset when total energy of most of the images are in the same range. In this case we will have:
| (12) |
in which denotes the inner vector product.
As we can see in (12), the optimum solution to the input signal in a diffusion system with fixed is the attention vector. But this is just one step of the overall diffusion system since each step in the forward diffusion process has a distinct increasing value of . We will discuss this convergence issue in the next section II-D.
The derivation resulting in (12) is also a mathematical basis for the very attention system. It shows that attention is the optimum solution for recovering signals in Gaussian noise as it was shown earlier in [7]. Note that in the conventional attention system [1], the noise variance , where is the dimension of the space.
This derivation also shows a basis for the softmax layer. The exponential form in (12) with , is the softmax output when is a standard one-hot label vector :
| (13) | ||||
| (14) |
which shows the equivalence of softmax and attention layers:
| (15) |
II-D Convergence
In Sec. II-C, it was shown that each step of a diffusion system calculates attention for a predefined value. The reverse generative process in DDPM starts with a random Gaussian and refines it with denoising network trained with noisy images from dataset. Theoretically, there should be hundreds of denoising networks each trained for denoising with a specific . But usually a shared network is trained and used for every values of to reduce complexity. Therefore, output of each step is then recursively fed in the network to generate another less noisy output image. This is repeated hundreds of times to reach a high quality image.
As can be seen from (1), is increasing with in the forward process. This will result in a decreasing in the backward generative process. We start with a Gaussian random . In the early stages when is large, differences in distances in (11) make negligible effect. Therefore, the output sequence converges to the mean value of the nearby vectors.
When gets small, differences in distances are promoted in (11) and the results quickly converge to the nearest signal, with some contribution from nearby signals. This is a sparse representation, as the simulation of (11) for a simple setup shows in Fig. 2. In this simulation, is a decreasing sequence which is decreased gradually from to in the reverse process. Although, as we will discuss later in Sec. II-E, it seems that does not converge to zero because of practical limitations in the denoising neural network design and training.
The diffusion process initially converges to the mean point of the distribution. Then in the long term, it converges to the vicinity of some dense region of the dataset. This is in accordance with the fact that the diffusion process simulates the real data distribution as reflected in the dataset.
Knowledge about the convergence of the diffusion process enables us to circumvent the whole convergence process by directly using a convex combination of some nearby dataset signals as output of the generation process. This basically reduces not only the lengthy generation process, but also the very lengthy training process of denoising neural network.
II-E Finite Training
In this manuscript, we have generally assumed infinite training of denoising network. This enables the training process to cover every point in the space. It also enables the network through many gradient descent updates, to achieve the ground-truth closed form solution of the training process optimization.
This is rarely the case in practice. Due to limited resources like silicon, time, and energy, we usually perform minimal training up to the point of acceptable network performance. Here we discuss the effect of non-asymptotic training on the network behavior.
First, the network does not experience every in the space during the training. It only sees finite adjacent inputs . It learns the correct direction of move from these inputs to outputs , respectively. Any new is sorrounded by multiple seen points. Since the network is a continouos system, we expect an average behavior in if the training set is sufficiently dense. This guides to regions of higher density which are experienced more in training and gradually we experience better refinements.
| (16) | ||||
| (17) |
where and are two convex coefficient sets, is the encoder function and is the decoder.
The second consequence of finite training is that the network does not learn the optimal solution of the optimization in (7) for each training input . We can model this training error as an additive noise in the output of the network. The same is true for the unseen input error and we can model the error in also as an additive noise.
When noise is added to the output of the denoiser network in a diffusion system, the output will no longer be a sparse attention combination even when of the diffusion system is very small. This is in fact a new additive noise in the backward process:
| (18) |
where is the new noise term indicating errors emanating from finite training process.
Therefore, the overall effect of finite training can be modeled as a higher value of which increases the number of intermittent signals. This means that although in the diffusion framework, theoretically gradually converges to zero, practical limitations leads to a stable higher limiting value for . This prevents the optimization process to achieve a very sparse solution and the final output will be a convex combination of many signals.
III Auto Encoders
In practice, diffusion systems use auto-encoders as both the outer encoder layer that the diffusion system operates inside [9], and the inner recurrent denoising network. Therefore, it is also crucial to understand working principles of AE to reduce computational complexity of the diffusion systems. In this section we use a different perspective to explain AE and attention system similarity.
III-A Auto-encoders are attentions
As we showed in Sec. II, there is a deep relationship between diffusion generative systems and attention mechanism. Here, we also show a strong resemblance between AEs, attention systems, and diffusion systems. This will present a great unity between three of the most important artificial intelligence systems currently in use.
First of all, notice that the derivations in Sec. II-C is not limited to a diffusion system. In fact, it is a widespread setting in unsupervised learning systems. Assume we have an AE network which as usual use MSE loss function to reconstruct input images or other signals in an unsupervised manner. Regardless of the AE’s inner structure, we assume perfect training of the system. Most of the times, an augmentation technique of adding Gaussian noise to limited dataset signals is used for better training and performance. Now, we can see that the mathematical setup is identical to the optimization in (7). Therefore, the same solution as in (11) applies to the problem. This certifies that such an AE system is equivalent to an attention system in (12) with fixed in the infinite training setup. Attention parameter in (11) can be tuned according to the structural parameters of the AE for a satisfactory equivalence.
The equivalence between AE and attention is double-sided. It means that we can in principle replace current attention systems with AEs. A good example for such a replacement is the linear attention [15], which is known to be a finite-state attention mechanism with compressed rather than full state. This is reminiscent of recursive auto-encoder structure [16], especially when we consider the causal implementation of the linear attention.
III-B Auto-encoders and Manifolds
Assume an auto-encoder with -dimensional input-output and -dimensional code layer as in Fig. 3. Manifold hypothesis assumes that data from the dataset is located on a high-dimensional manifold in input space [10]. It is well known that the autoencoder estimates the data manifold during the training phase [11]. We will show that AE projects the -dimensional input signal on to a -dimensional manifold which represents locus of data in the dataset and their meaningful augmentations in the space.
Assume that the AE is fully trained with the dataset and therefore, it can now regenerate signal in the output approximately. In the input space, we add a differential length vector to . With high probability, no ReLU-like switch in the network is on its change position. Therefore, for any differential change in the input, the system is linear. Then, changing the input with results in a differential vector change in the code layer:
| (19) |
Continue these changes with orthogonal differential vectors that results each in code layer change of . In the ’th step, the set of is a linearly dependent set since the code layer vector space is -dimensional. Therefore, a linear combination of these vectors sums to zero:
| (20) |
The same set of coefficients can be used to find a direction in the input space that keeps the code layer, and therefore the output constant:
| (21) |
Therefore, we showed that any vectors in the input space can form a direction with null effect in the output. This direction is:
| (22) |
The local null space of the encoder at is defined as the local Affine vector space of all such vectors:
| (23) |
where is a small constant. The null space is local to and is effective just in a -neighborhood of . At most input vectors can be chosen that does not lie in this local null space. Therefore:
| (24) |
Suppose is the -dimensional local Affine space orthogonal to the local null space:
| (25) |
In the vicinity of , any displacement along the null space does not change the values of the code layer and therefore, the output. But displacement along manifold translates to changes in code layer and output. Therefore, the AE will tune its null space such that it direct farthest from the neighbors of . This will set the estimated data subspace in the nearest position that best cover the neighbors of and orthogonal to the local null space. This is in accordance with autoencoder estimating a manifold with good fit to neighboring data points [11] with extensive training on sufficiently dense datasets.
Here we assume that the dataset is dense enough that the continuity of the data manifold can be tracked by the AE as in Fig. 4 . We also assume that the AE does not have a specific structure (like convolutional or any other weight sharing scheme). This helps the AE to follow any direction in the data space without any bias. It is well-known that networks like convolutional networks better estimate the data manifold. This is because some of the data structure, like image content translation invariance is embedded in such networks. This is equivalent to more data points which helps the network better estimate the data manifold.
Any input outside the manifold defines a specific local null space which aggregates directions of moves that does not change the code layer and output layers values. Therefore, AE moves in a curved path along local null spaces to finally reach the AE estimated data manifold in . Since we reach to the AE manifold perpendicularly in the last touch as is shown in Fig. 4.
III-C Attention and manifold
Here, we show that the attention mechanism also projects data on a low dimensional manifold. This is in-line with our core insight that basically these three structures: AE, attention, and diffusion systems are similar.
Assume that is near the data manifold. Dataset satisfies conditions that we can use (11) in place of the better-known attention formula in (12). Also, assume that is small enough that only nearest signals:
| (26) |
are effective in the attention sum in (11). Then, the attention coefficients are:
| (27) |
The adjacency set forms a -dimensional Affine subspace. Assume that the orthogonal projection of onto this subspace is as in Fig. 5. Then:
| (28) |
Then is shared in numerator and denominator of (27), eliminating it:
| (29) |
Output of from the attention mechanism is a convex combination of and therefore is on their Affine subspace. Also, the fact that the output of is the same as the output for shows that there is a local Affine null space orthogonal to the local affine space of nearby signals in that does not change the output of the attention.
III-D Auto-encoders resemble attention
The attention mechanism when act on a point near the data manifold , orthogonally projects it on the manifold at point , and then transform to another point in the convex hull of signals in in the output. Overall, we can conclude that the attention mechanism projects any point near the data manifold to some point on the manifold too and this is what also auto encoders do.
The attention mechanism depends on the parameter , while the AE is characterized by its code layer dimension . The equivalence between attention and AE requires fine tuning these parameters. As stated before, setting a small is required in order that approximately nearby signals are effective in the attention. This will result in the estimation of a -dimensional manifold for data similar to an AE with -dimensional code layer. This is not an exact equivalence because the required parameter may not be constant in different parts of the input space.
In a conventional diffusion system, we have an outer autoencoder which compresses data to a lower dimension. Then a recursive inner autoencoder denoises an input noise. We have shown that not only the high level diffusion framework is equivalent to the attention mechanism, but also both autoencoders in the network structure realization perform in a similar way to the attention mechanism. Therefore, we can alternate between these three structures based on the situation.
If attention and auto encoder structures are basically similar, what is their difference? The difference lies in where the computational load is concentrated. In an auto encoder, we use vast amount of data in training phase while the test phase is as simple as one forward calculation of the network. The attention mechanism is reverse. Calculating attention does not require any training while all the computational load is concentrated in the test time where the forward pass of the network requires vast amount of product calculations with a long context of vectors. Therefore, we may be able to reduce the complex training phase of the auto encoder and diffusion systems by using some simpler form of the attention mechanism.
IV Tree Cross Attention
The problem with the attention mechanism is that it involves calculating the dot product of the input with every image in the dataset. These image datasets are usually very large in industry-level network training. Since usually parameter of attention is small, it suffices to calculate the attention just for nearby signals. But finding the nearbies requires calculating the distances to all images in the dataset which is equivalent to inner product computation.
Here, a data structure can lead to considerable reduction in the computational complexity. Calculating the dot product of with every images in the dataset is of complexity. But since there is a great deal of redundancy in this calculation, we can save some information for reuse.
Data can be order in a tree structure based on similarity [17]. Each node in the tree is an image and any new image is tested with every nodes in each tree level to see with which branch it is more similar. If there was a similar branch, the new image goes down to that branch’s next level children. If there were no similar branch, make a new branch in the current level.
In this way, a balanced tree can be constructed in which similar and nearby images are clustered in a branch. Then for each new input signal , we need a few inner product calculations of which is the tree depth to find its neighbors for efficient attention calculations. This method is called tree cross attention and is presented in the literature [18], for the situations where base attention vectors are constant. Therefore, attention calculation is not restrictive in an image generation scenario.
V Proposed Generation Algorithm
Theoretical equivalence of diffusion and attention systems helps to reduce the computational complexity of image generation. We have seen in Sec. II-D that any noise input to a diffusion system experiences a long road to its final position in latent space.
Since the solution to diffusion is an approximately sparse one (Fig. 2), we can start image synthesize from one of dataset images. Since attention to latent signals is approximately constant in the neighborhood of the image latent, we can calculate the attention of the latent signal to other neighbor signals.
This attention calculation helps the resulting signal to reside on or very near the data manifold since it is equivalent to passing the signal once through the hypothetical diffusion trained AE denoiser. Therefore, we can pass the resulting latent through the decoder and reach to a high quality innovative image. Since there may be slight disposition from the data manifold, we can pass the resulting image a few times through the AE to better fit it to the data manifold.
Therefore, there will be no need to train a denoiser auto encoder as the working engine of the diffusion system. We just need to train an outer AE which projects the input images to the low-dimensional latent space. Then in the generation phase, we just rely on attention calculation near the signal latents. Therefore, we not only reduced the training phase calculations, but also in the test phase, we reduce the long journey of the input noise to the distribution mean value and then to nearby of one of the signals. We just perform the final step with calculating the attention of the signal to its neighbbors including itself. The block diagram of the proposed image generation algorithm is depicted in Fig. 6.
The generation process is done in the learned latent space of trained VAE. As described in Alg. 1, for a random latent query , we find its nearest neighbors . New latent representation is generated through attention interpolation between the query and its neighbors. The generated latent representation is then passed through the VAE decoder to obtain a new image. Since nearby latent representations correspond to semantically related data, their interpolation produce meaningful intermediate representations.
Although using plain attention between querry and its neighbor images produce meaningful images, we use a fine grain spatial attention mechanism to get better quality. Spatial attention is calculated based on the distance of each pixel depth vector in querry to its corresponding pixel depth vector in a neighbor image as presented in Alg. 2. This generates a spatial attention map that determines the contribution of each pixel to the generated sample. Example spatial attention maps in Fig. 7 illustrate how the proposed method selectively exploits spatial information from neighbors during latent-space sampling. Note that spatial attention resembles the patch similarity selection mechanism that is developed as the theoretical solution of a convolutional diffusion network in [8].
| NN1 | NN2 | NN3 |
|
|
|
|
VI Simulation Results
We train the model using a P100 GPU on the MNIST, Fashion-MNIST, FFHQ, and CelebA datasets with an image size of for MNIST and Fashion-MNIST, for FFHQ and for CelebA. The Adam optimizer is used with and , and a learning rate of . For outer VAE training, we use a combination of loss, perceptual loss based on VGG network and an adversarial (GAN) loss of StyleGAN discriminator with the following architecture:
| (30) |
In our approach, generation is performed by local interpolation which is substantially faster than iterative reverse diffusion. The diffusion method requires second versus the proposed method just second to generate an image on the FFHQ dataset, hence a speedup.
Example generated images are presented in Fig. 8 for FFHQ, MNIST, and Fashion-MNIST datasets. The first column from left in each dataset is the chosen query, the second is the generated image, and the last three are nearest neighbors, respectively. Generated images are quite natural and innovative with overall similarity to the query image while they receive special features from neighbors.
FFHQ
MNIST
Fashion MNIST
More generated sample images is shown in Fig. 9. It can be seen that most of the images are natural and of good quality. The higher resolution of 128128 with CelebA dataset is also examined in Fig. 10.
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
Baseline diffusion model configurations are listed in Table I. A different setting is used for each dataset used for comparison. Similar settings of the proposed method is also listed in Table II. Table III compares FID of the proposed method versus the baseline diffusion model. While our proposed model has much less parameters, it achieves better FID in MNIST and comparable FID in Fashio MNIST and CelebA, while the performance is lower in FFHQ dataset. Table IV shows an optimization study which shows that lower parameter of the attention mechanism will result in better FID metric.
| MNIST | FashionMNIST | FFHQ | CelebA | |
| 1 | 1 | 4 | 4 | |
| Diffusion steps | 1000 | 1000 | 1000 | 1000 |
| Noise schedule | linear | linear | linear | linear |
| Backbone | CNN | CNN | Transformer | CNN |
| #Params | 15.7M | 15.7M | 130M | 47M |
| Attention resolution | 7 | 7 | – | 16, 8 |
| #Heads | 4 | 4 | 12 | 1 |
| Batch size | 256 | 256 | 256 | 32 |
| Iterations | 30k | 36k | 134k | 51k |
| Learning rate |
| MNIST | FashionMNIST | FFHQ | CelebA | |
| 4 | 4 | 8 | 8 | |
| -shape | ||||
| Latent Signals | 1024 | 1024 | 2048 | - |
| Backbone | CNN | CNN | CNN | CNN |
| #Params | 2.95M | 2.95M | 64.8M | 53M |
| Attention resolution | – | – | Encoder: 8 Decoder: 8, 16 | – |
| #Heads | – | – | 1 | – |
| Batch size | 256 | 256 | 256 | 32 |
| Iterations | 13.6k | 31.3k | 44k | 51k |
| Learning rate |
| Dataset | Method | FID |
| MNIST | Diffusion | 19.30 |
| Ours | 8.60 | |
| FMNIST | Diffusion | 13.00 |
| Ours | 14.60 | |
| FFHQ | Diffusion | 8.60 |
| Ours | 16.58 | |
| CelebA | Diffusion | 30.49 |
| Ours | 32.15 |
| Top- | Noise scale | FID | |
| 30 | 3 | 0.01 | 27.43 |
| 30 | 5 | 0.01 | 34.33 |
| 0.01 | 3 | 0.01 | 16.69 |
| 0.001 | 3 | 0.01 | 16.58 |
References
- [1] A. Vaswani, et al., “Attention is all you need,” in Proc. Ann. Conf. Neur. Inf. Process. Syst., Long Beach, CA, USA, pp. 5998–6008, 2017.
- [2] W. Sun, J. Hu, Y. Zhou, et al., “Speed always wins: A survey on efficient architectures for large language models,” arXiv:2508.09834, 2025.
- [3] A. Dosovitskiy, et al. “An image is worth 16x16 words: Transformers for image recognition at scale,” arXiv preprint arXiv:2010.11929. 2020.
- [4] J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli, “Deep unsupervised learning using nonequilibrium thermodynamics,” In International Conference on Machine Learning, pages 2256–2265, PMLR, 2015.
- [5] W. Peebles, and S. Xie, “Scalable diffusion models with transformers,” In Proceedings of the IEEE/CVF international conference on computer vision, pp. 4195-4205. 2023.
- [6] H. V. Poor, An introduction to signal detection and estimation, Springer, 2013.
- [7] J. Henderson and F. Fehr, “A VAE for Transformers with Nonparametric Variational Information Bottleneck,” in Proceedings of International Conference on Machine Learning ICLR, 2023.
- [8] M. Kamb and S. Ganguli, “An analytic theory of creativity in convolutional diffusion models,” In Proceedings of the 42nd International Conference on Machine Learning, Vancouver, Canada, 2025.
- [9] R. Rombach, et al. “High-resolution image synthesis with latent diffusion models,” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022.
- [10] J. Braunsmann, M. Rajkovic , M Rumpf, and B. Wirth, “Convergent autoencoder approximation of low bending and low distortion manifold embeddings,” Mathematical Modelling and Numerical Analysis, vol. 58, pp. 335-361, 2024.
- [11] Y. Lee, “A geometric perspective on autoencoders,” available online: https://arxiv.org/abs/2309.08247
- [12] J. Song, C. Meng, and S. Ermon. “Denoising diffusion implicit models,” arXiv preprint arXiv:2010.02502, 2020
- [13] C. Luo, “Understanding diffusion models: A unified perspective,” arXiv preprint arXiv:2208.11970, 2022
- [14] S Boyd and L. Vandenberghe, Convex optimization, Cambridge Univ. Press, 2004.
- [15] A. Katharopoulos, et al., “Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention,” in Proceedings of the 37th International Conference on Machine Learning, pp. 5156–5165, 2020.
- [16] J. Zhang, R. Su, C. Liu, et al, “A Survey of Efficient Attention Methods: Hardware-efficient, Sparse, Compact, and Linear Attention,” 2025.
- [17] R. Sedgewick and K. Wayne, Algorithms, Addison-Wesley, 4’th ed., 2011.
- [18] L. Feng, et al. “Tree cross attention,” arXiv preprint, arXiv:2309.17388, 2023.











