arXiv is now an independent nonprofit! Learn more
License: CC BY 4.0
arXiv:2609.40347v1 [cs.CV] 30 Sep 2026

VideoMSN

     Image Classifiers are Efficient Self-Supervised Video Representation Learners

Owais Iqbal    Sudipta Sarkar    Shyam Marjit    Omprakash Chakraborty    Anirban Chakraborty    Abir Das
Abstract

We introduce VideoMSN\xspace, a Masked Siamese Network framework for efficient self-supervised spatio-temporal representation learning in videos. Instead of relying on heavy 3D architectures or reconstruction-based autoencoders for learning with unlabeled data, we repurpose standard image Vision Transformers by representing videos as super images which are grids composed of frames sampled from videos. From each super image, we construct two views: one with spatial patch masking and the other with temporal frame masking, ensuring no information leakage across frames. A shared Vision Transformer (ViT) encoder aligns their embeddings using a masked Siamese loss, capturing both motion and appearance cues without reconstruction. Our decoder-free formulation leverages an image foundation model towards efficient video representation learning. Starting from pretrained DINO-v3 and DeiT-v3 image encoders, VideoMSN\xspaceachieves state-of-the-art performance on Kinetics-400, UCF101, and HMDB51 while requiring up to 32×32\times fewer and 160×160\times fewer video pretraining epochs compared to prior video self-supervised learning methods. Our proposed approach also shows strong performance in low-shot classification, confirming the transferability of the learned representations in a label-scarce scenario. Project Page: https://cvir.github.io/projects/videomsn.

††email: owais.iqbal@kgpian.iitkgp.ac.in††email: sudipta25t@kgpian.iitkgp.ac.in††email: shyammarjit@iisc.ac.in††email: omprakash.chakraborty@livia.etsmtl.ca††email: anirban@iisc.ac.in††email: abir@cse.iitkgp.ac.in††affiliation: Indian Institute of Technology
Kharagpur, India
††affiliation: Indian Institute of Science
Bangalore, India
††affiliation: École de technologie supérieure
Montreal, Canada

1 Introduction

Self-supervised learning (SSL) has emerged as a strong paradigm for visual representation learning without the need for meticulously labeled data. Instead, it intelligently leverages patterns naturally present in images and learns through pretext tasks. Pretext tasks for images can vary from exploiting their spatial structure [13, 49, 15], solving jigsaw puzzles [48], colorizing images [73, 38], predicting rotations in artificially rotated images [24], etc. One of the common and effective pretext tasks is Masked Image Modeling (MIM). It involves masking a portion of the input and either predicting the masked regions [28, 5, 4] or producing similar embeddings for the masked and the unmasked inputs [2, 60]. While reconstructing masked patches through an encoder-decoder framework is effective, this approach emphasizes unnecessary low-level details, often at the expense of longer training and higher compute costs. In contrast, a Masked Siamese Network (MSN) [2] uses an encoder-only framework to align features of masked and unmasked views of the same image. As masked image portions do not go through the encoder, it exhibits good computational scaling with fewer training epochs compared to the reconstruction-based approaches.

Figure 1: Comparison of top-1 accuracy on Kinetics-400 across state-of-the-art self-supervised video representation learning methods. Each point denotes a method, with bubble size proportional to its pretraining epochs. Our method, VideoMSN\xspacewith DINO-v3 backbone, achieves state-of-the-art performance with 160×\textbf{160}\times and 60×\textbf{60}\times fewer training epochs compared to VideoMAE [64] and SMILE [63] respectively.

Yet efforts to scale these methods to videos are hindered by the need to capture both spatial and temporal dynamics. Spatio-temporal processing often relies on heavy 3D CNNs or video vision transformers [1, 7], which are computationally intensive and memory-demanding, necessitating large-scale, sophisticated GPUs for progress. Interestingly, 2D image models on 2D super images created by rearranging video frames into multiple rows and columns [18] have proven to be a powerful machinery for video action recognition without being parameter heavy. Naturally, recasting video action recognition as 2D image classification enables a plethora of highly efficient self-supervised learning approaches designed for images to be equally applicable for scalable and efficient self-supervised video representation learning.

In this work, we propose Video Masked Siamese Network (VideoMSN\xspace) that masks super images created from unlabeled videos and leverages a masked Siamese network for invariant representation learning across masked and original views of the super images. Given a super image created by rearranging a sequence of input video frames into multiple rows and columns, VideoMSN\xspace randomly masks patches from one view while leaving the other view unchanged. To avoid temporal redundancy across frames, we make sure that if a region is masked in any frame, the same region is also masked across all the frames avoiding information leakage [64]. Masking in images helps to get a strong encoder by breaking the spatial continuity of the images. We conjecture that the complexity in videos due to the additional temporal dynamics demands breaking the temporal structure and forces the encoder to be invariant to this. Thus we propose a more aggressive temporal masking by dropping whole frames and the encoder learns to match the representations of an unmasked super image and super images with dropped frames. Unlike reconstruction-based methods, VideoMSN\xspace eliminates the need for pixel-level reconstruction making the framework decoder-free. An additional benefit of our decoder-free design is the ability to leverage off-the-shelf image encoders (e.g., ViT) pretrained solely on image data, thereby significantly reducing the amount of self-supervised pretraining with video data. In contrast, encoder-decoder based MIM approaches typically require training from scratch due to the lack of image-pretrained decoders which results in longer pretraining schedules.

By leveraging a 2D image classification pipeline for video understanding and incorporating view-invariant masked image modeling, our approach is highly efficient across multiple dimensions: it is parameter-efficient, GPU memory friendly and label-efficient. It reaches state-of-the-art performance with significantly fewer self-supervised pretraining epochs with video data. Fig. 1 shows a comparison of the number of self-supervised pretraining epochs of different approaches along with video classification accuracy. This shows that the compute-heavy pretraining with video data for VideoMSN\xspace can be upto 160×160\times fewer epochs compared to state-of-the-art approaches like VideoMAE [64], MGM [16] or MME [59], while maintaining the performance on standard video benchmarks. We perform extensive experiments on four benchmark datasets and demonstrate the superiority of VideoMSN\xspace over contemporary self-supervised video action recognition approaches.

Our key contributions are as follows:

  • •

    To the best of our knowledge, VideoMSN\xspace is the first work that successfully extends the masked Siamese network to videos leveraging an image classification pipeline for self-supervised video representation learning.

  • •

    In addition to dropping spatial patches from frames, we show that the strategy of dropping or masking frames completely in the temporal dimension to create the masked view of the videos results in better representation learning.

  • •

    Our parameter-efficient encoder-only design enables self-supervised adaptation of image-pretrained vision transformers to videos, requiring up to 160×160\times fewer additional video pretraining epochs than recent generative masked video modeling approaches while remaining competitive or superior on benchmark datasets.

2 Related Work

2D Action Recognition. Since videos contain information along both spatial and temporal dimensions, spatio-temporal or 3D processing for recognizing actions has been the mainstay for a long time. However, 2D image models have also been used for action recognition to enhance memory and compute efficiency. Early approaches [75, 54] used a single image from videos to recognize actions. Later works create a representative image from the video by informative frame synthesis [52], adaptive spatio-temporal distillation [62] and adversarial video distillation [61]. However, using a single image to represent actions is limiting and hurts the performance. Several later works [67, 76, 39, 19] used different aggregation modules on 2D image backbones. TSM [39] and its improvement TAM [19] shift channels of 2D-CNNs along the temporal dimension. Another set of approaches predicts actions by identifying key frames of an activity [70, 44, 58]. With the success of vision transformers, action recognition frameworks started exploring them [46, 17, 74]. Recently, action recognition in videos is cast as an image classification problem [18, 32] in which the frames are combined in a spatial grid to form a super image and classified using a Swin Transformer for images [40].

Self-supervised Video Representation Learning. These approaches, being supervised, are critically dependent on large datasets requiring labels. Self-supervised representation learning models address this by leveraging unlabeled data. These approaches can be categorized into three broad paradigms. The first, transformation prediction, uses pretext tasks such as solving space-time puzzles [33, 36], predicting clip order [45, 43, 23] or estimating playback speed [6, 12]. The second, contrastive learning, trains the model to align different augmentations of the same clip while separating others [50, 21, 53]. The third, masked video modeling [16, 31, 59], adopts a mask-and-predict framework, where models like VideoMAE [64] mask a large portion of video tokens and reconstruct them using an encoder-decoder architecture. However, such reconstruction-based methods try to capture unnecessary visual details and incur high computational cost. While SMILE [63] attempts to improve semantic learning by predicting the CLIP features corresponding to synthetic motion, it still relies on an encoder-decoder architecture. VideoMSN\xspace, on the other hand, eliminates the decoder entirely. Our approach replaces pixel reconstruction with feature alignment between masked and unmasked super images, using clustering losses with learnable prototypes and entropy maximization. This decoder-free design retains high-level spatio-temporal semantics and substantially reduces pretraining epochs with videos.

Masked Input Modeling for Vision Transformers. The idea of reconstructing masked inputs as a generative pretext task originated with denoising autoencoders [51] and was later scaled to vision transformers through MAE [28], which showed that masked reconstruction can produce highly transferable representations. Building on this, various image-based methods [28, 47, 69] have advanced masked image modeling, while approaches like JEPA [4] shifted the objective to predicting masked tokens in latent space, achieving strong performance. In videos, VideoMAE [64] introduced masked autoencoding, prompting follow-up works to refine reconstruction targets [69, 68] and incorporate motion-aware masking strategies [59]. These masked autoencoders reconstruct original signals from corrupted inputs through an encoder-decoder setup. Parallel to this, Siamese networks [8] have enabled robust representation learning through contrastive objectives [29, 10] and have recently been explored in conjunction with masked modeling [77, 2]. However, no prior work has explored an asymmetric, decoder-free masked Siamese design for videos. In this paper, we propose VideoMSN\xspace, the first adaptation of Masked Siamese Networks for videos. We position VideoMSN\xspaceas a highly efficient integration of super image representations, image-pretrained Vision Transformers, and Masked Siamese Learning, rather than a fundamentally new architectural paradigm.

3 Methodology

In this section, we briefly revisit the Masked Siamese Network��[2] (MSN) used in images. Then, we describe VideoMSN\xspace and its components in detail.

3.1 Preliminaries

MSN [2] is a self-supervised learning framework introduced for image representation learning using a discriminative mask-denoising process. MSN combines masked image modeling with a Siamese learning objective. The method constructs two augmented views of the same image: one is partially masked (the anchor view), and the other is unmasked (the target view). Both views are passed through a shared ViT encoder. The model is trained to align the representation of the masked view with that of the unmasked view using a soft-distribution over a set of prototypes for both the anchor and target views. This encourages the network to produce semantically meaningful features from incomplete visual inputs. MSN achieves strong performance by learning visual representations without labels.

3.2 Masking Strategy in VideoMSN\xspace

Refer to caption
Figure 2: Creation of super image and a random view. Starting from an input video, a super image SiS^{i} is formed by arranging MM sampled frames into a 2D grid in a row-major format. Each frame is then divided into non-overlapping patches. A random binary mask is applied such that patches at the same spatial location across all frames are masked or retained together enforcing temporal consistency during masking. This results in the Random View (RiR^{i}). Masked patches are shown in gray for illustration. In practice, the token from that patch is dropped and not passed through the encoder.

Let 𝒟u={Ui}i=1Nu\mathcal{D}_{u}=\{U^{i}\}_{i=1}^{N_{u}} be a collection of NuN_{u} unlabeled videos and let BB be the number of such videos in each mini-batch during pretraining. For each video UiU^{i}, we sample MM equidistant frames from non-overlapping temporal segments [67]. Following SIFAR [18], these frames are arranged in a fixed grid layout to form a 2D super image SiS^{i} (ref. Fig. 2). Each super image (SiS^{i}) gives one target view (TiT^{i}) and a set of anchor views (AiA^{i}). Note that the target view does not go through any masking; however, it is patchified into a set of non-overlapping patches and augmented. Anchor views, on the other hand, go through masking (detailed below) after similar patchification and augmentation.

Following the masking strategy proposed in MSN [2], an anchor view, in our case, comprises of random views and focal views coming from the super image. A random view RiR^{i} of a super image SiS^{i} is a result of masking spatial patches from the super image. Specifically, each frame in SiS^{i} is first divided into N×NN\times N non-overlapping patches, resulting in a set of N2N^{2} tokens for each frame. A random binary mask is applied over the patches such that all patches at the same spatial location across different frames are either simultaneously masked or retained. This design is consistent with the temporal tube masking formulation proposed in [64] and is shown to be effective in avoiding shortcuts in masked modeling for videos. An illustration of the masking process is shown in Fig. 2.

Refer to caption
Figure 3: Illustration of two-stage masking of focal views. Temporal Masking drops full frames to generate two different views: the Fast View (VfiV^{i}_{f}), retaining (approximately) half the frames of the target view of the super image and the Slow View (VsiV^{i}_{s}) containing approximately a quarter of the same. Subsequently, Focal Masking is applied by selecting a spatially contiguous region and masking all surrounding patches, resulting in the Temporally Masked Fast Focal View V^fi\hat{V}^{i}_{f} and Temporally Masked Slow Focal View V^si\hat{V}^{i}_{s}. This hierarchical masking strategy helps the model attend to the informative regions in both space and time, enabling it to better capture motion patterns and spatial details across different action speeds.

To obtain a focal view (FiF^{i}) of a super image (SiS^{i}) we propose a two-stage masking mechanism. The first stage is temporal masking which is illustrated in Fig. 3. It creates a temporally downsampled version of the super image by masking the whole of certain frames in the super image. As a result, not all MM sampled frames are used in creating this view. After this we create a fast view and a slow view from the frames that are retained.

  • •

    Fast View VfiV^{i}_{f} is obtained by dropping approximately half the number of frames present in SiS^{i}, such that the number of frames forming VfiV^{i}_{f} is approximately M2\frac{M}{2}.

  • •

    Slow View VsiV^{i}_{s} is obtained by dropping an additional half the number of frames from VfiV^{i}_{f}, making the number of frames in VsiV^{i}_{s} approximately M4\frac{M}{4}.11 1 Slow/Fast nomenclature is motivated by SlowFast Networks [20].

The reason, the number of frames in the super image after temporal masking are kept approximately M2\frac{M}{2} and M4\frac{M}{4}, is that a super image works best when the layout is square i.e., equal number of rows and columns form the super image [18]. Different from focal views in MSN applied to images, temporal masking in VideoMSN\xspace helps exploit different temporal sparsity of the same action, offering varied motion dynamics for the model to learn from.

The second stage involves focal masking on each of these temporally masked views. Specifically, a contiguous region in each frame is randomly selected within each temporally masked view and retained, while the surrounding patches are masked out. Similar to the random views, we propose to mask out the same spatial area in each frame of the super image. Following such a strategy, we get multiple focal views via multiple random selections of contiguous regions from each of the temporally masked views. Mathematically, the fast view VfiV^{i}_{f} gives rise to nn temporally masked fast focal views {V^f,ji}j=1n\{\hat{V}^{i}_{f,j}\}_{j=1}^{n}. Similarly, nn temporally masked slow focal views {V^s,ji}j=1n\{\hat{V}^{i}_{s,j}\}_{j=1}^{n} are obtained from the slow view VsiV^{i}_{s}. The set of anchor views (AiA^{i}) contains random views, temporally masked fast focal views and temporally masked slow focal views. This progressive masking strategy helps learn from both global context and localized high-resolution cues, improving spatio-temporal feature alignment across views of different temporal granularity.

Refer to caption
Figure 4: Overview of VideoMSN\xspace. Given an unlabeled video clip UiU^{i}, equidistant frames are sampled and arranged into a 2D super image. Two different views are generated: the jt​hj^{th} anchor view AjiA^{i}_{j}, which is randomly masked after patchification and the target view TiT^{i} which remains unmasked. Both views are processed using a Siamese architecture. AjiA^{i}_{j} is passed through the encoder fθ​(⋅)f_{\theta}(\cdot) to obtain embeddings 𝐎ji\mathbf{O}^{i}_{j}, while TiT^{i} is passed through the EMA updated target encoder fθ¯​(⋅)f_{\bar{\theta}}(\cdot) to produce 𝐎+i\mathbf{O}^{i}_{+}. These representations are then assigned to cluster prototypes, generating a predicted distribution 𝐩ji\mathbf{p}^{i}_{j} for the anchor and a target distribution 𝐩+i\mathbf{p}^{i}_{+} for the target. The training objective is to align 𝐩ji\mathbf{p}^{i}_{j} with 𝐩+i\mathbf{p}^{i}_{+} using cross-entropy loss H⁡(𝐩ji,𝐩+i)H(\mathbf{p}^{i}_{j},\mathbf{p}^{i}_{+}), encouraging the masked anchor representation to match that of the unmasked target.

3.3 Learning in VideoMSN\xspace

Fig. 4 provides an overview of our VideoMSN\xspace approach. Following clustering-based self-supervised learning frameworks [9, 10, 3], the super images corresponding to the target and the anchor views are passed through an encoder to produce feature representations, which are then projected onto a set of learnable prototypes. The projections are converted into distributions and the encoder is encouraged to produce similar distributions coming from the masked anchor views and the unmasked target views using cross-entropy loss.

Let fθ​(⋅)f_{\theta}(\cdot) denote the parameterized anchor encoder and let 𝐎ji=fθ​(Aji)∈ℝd\mathbf{O}^{i}_{j}=f_{\theta}(A^{i}_{j})\in\mathbb{R}^{d} represent the output embedding obtained from the jt​hj^{th} anchor view AjiA^{i}_{j}. Note that the set of anchor views contains both random and temporally masked focal views. Similarly, let fθ¯​(⋅)f_{\bar{\theta}}(\cdot) be the target encoder, parameterized by θ¯\bar{\theta}, and let 𝐎+i=fθ¯​(Ti)∈ℝd\mathbf{O}^{i}_{+}=f_{\bar{\theta}}(T^{i})\in\mathbb{R}^{d} denote the embedding computed from the target view TiT^{i}. Following MSN [2], the target encoder weights θ¯\bar{\theta} are updated using an exponential moving average (EMA) of the anchor encoder parameters θ\theta [27]. Both encoders share the same ViT architecture [14] and the final representation is taken from the [CLS] token at the output of the transformer.

Prototype-driven Learning. We adopt a prototype-based training objective to learn video representations without labels inspired by [2]. Specifically, we maintain a set of PP learnable prototypes each with dimension dd. The collection of prototype vectors are denoted as 𝐪∈ℝP×d\mathbf{q}\in\mathbb{R}^{P\times d}. We compute anchor and target predictions by measuring cosine similarity between the encoder outputs and the prototypes. For the jt​hj^{th} anchor representation 𝐎ji\mathbf{O}^{i}_{j}, the prediction is computed as:

𝐩ji=softmax⁡(𝐪⋅𝐎jiτ),\mathbf{p}^{i}_{j}=\mathrm{softmax}\left(\frac{\mathbf{q}\cdot\mathbf{O}^{i}_{j}}{\tau}\right),\vskip-5.69046pt (1)

with a temperature parameter τ∈(0,1)\tau\in(0,1), while the target prediction 𝐩+i\mathbf{p}^{i}_{+} is similarly obtained using a sharper temperature parameter τ+<τ\tau^{+}<\tau:

𝐩+i=softmax⁡(𝐪⋅𝐎+iτ+).\mathbf{p}^{i}_{+}=\mathrm{softmax}\left(\frac{\mathbf{q}\cdot\mathbf{O}^{i}_{+}}{\tau^{+}}\right)\>.\vskip-5.69046pt (2)

The training objective consists of two components. First, we minimize the cross-entropy loss HH, between anchor prediction 𝐩ji\mathbf{p}^{i}_{j} and target prediction 𝐩+i\mathbf{p}^{i}_{+}:

ℒalign=1K​B​∑i=1B∑j=1KH⁡(𝐩ji,𝐩+i),\mathcal{L}_{\text{align}}=\frac{1}{KB}\sum_{i=1}^{B}\sum_{j=1}^{K}H(\mathbf{p}^{i}_{j},\mathbf{p}^{i}_{+}),\vskip-5.69046pt (3)

where KK denotes the total number of anchor views including random and all focal views. We include a Mean Entropy Maximization (ME-MAX) regularizer [3, 34] to encourage uniform prototype usage. We first, compute the average prediction 𝐩¯\bar{\mathbf{p}} across all the anchor views in the batch as,

𝐩¯:=1K​B​∑i=1B∑j=1K𝐩ji,\bar{\mathbf{p}}:=\frac{1}{KB}\sum_{i=1}^{B}\sum_{j=1}^{K}\mathbf{p}^{i}_{j},\vskip-5.69046pt (4)

and subsequently add the ME-MAX regularizer making the overall objective,

ℒtotal=ℒalign−λ​H​(𝐩¯),\mathcal{L}_{\text{total}}=\mathcal{L}_{\text{align}}-\lambda H(\bar{\mathbf{p}}),\vskip-5.69046pt (5)

where λ\lambda is the weight of the regularizer. This joint objective has been shown to prevent representation collapse while promoting discriminative embeddings [3, 34].

4 Experiments

In this section, we present comparative evaluation of our VideoMSN\xspace framework. We also perform comprehensive ablation studies to verify the effectiveness of different components and hyperparameter sweep to choose crucial hyperparameters.

Datasets and Backbones. We evaluate VideoMSN\xspace on four widely used datasets: Kinetics-400 [35], UCF101 [57], HMDB51 [37] and Something-Something V2 (SSV2) [26]. As the backbone to our encoder, we use the small and base versions of Vision Transformer [14] (denoted as ViT-S and ViT-B respectively). We used the pretrained DeiT-v3 and DINO-v3-distilled checkpoints made available by the authors of [65] and [56] respectively to initialize our encoder. These two variations are denoted as VideoMSN\xspace-DeiT and VideoMSN\xspace-DINO respectively.

Implementation Details. Our pre-training pipeline follows [18]. We adopt uniform frame sampling and apply standard multi-scale jittering followed by Gaussian blur. The frames are randomly cropped to a resolution of 224×224224\times 224 before forming the super images. Consistent with [2], we set the number of prototypes PP to 10241024 with dimension d=256d=256. The hyperparameters τ,τ+\tau,\tau^{+} and λ\lambda are set to 0.1,0.0250.1,0.025 and 5.05.0 respectively. The number of focal views is set to 66 unless otherwise specified. Half of the focal views are fast and the rest are slow focal views. We use a masking ratio of 0.70.7 i.e., 70%70\% of tokens are dropped while creating the anchor views, except for SSV2 where the value is taken to be 0.50.5. Drop Path and weight decay values are kept both 0.010.01. All experiments were conducted on a server with 4 NVIDIA H100 GPUs. All VideoMSN\xspace-DeiT/VideoMSN\xspace-DINO models undergo the proposed self-supervised pretraining for 50/1050/10 epochs, including a 5/15/1 epoch linear warm-up. We used AdamW optimizer [41] and follow a cosine learning rate.

For full fine-tuning, we used Mixup [72] and CutMix [71] with mixing coefficients of 0.80.8 and 1.01.0, respectively and apply label smoothing with coefficient 0.10.1. All models are fine-tuned for 30 epochs unless otherwise mentioned. We used AdamW optimizer with a weight decay of 0.010.01 and follow a cosine learning rate scheduler for all datasets. For evaluation, we used 5​c​l​i​p​s×3​c​r​o​p​s5~clips\times 3~crops for Kinetics-400 and UCF101, 10​c​l​i​p​s×3​c​r​o​p​s10~clips\times 3~crops for HMDB51 and 2​c​l​i​p​s×3​c​r​o​p​s2~clips\times 3~crops setup for SSV2. Performances are shown in terms of Top-1 accuracy averaged across 2 random seeds, unless otherwise mentioned.

Choosing hyperparameters. We identified some crucial hyperparameters by conducting a sweep on a representative subset of the Kinetics-400 training set, created by randomly sub-sampling 25%25\% of the data per class. The configuration that maximized performance on the full validation set was subsequently adopted for all experimental benchmarks. More details of this analysis are provided in the Appendix.

Comparison. We compare against state-of-the-art MAE based approaches like VideoMAE [64] and its architectural variants as well as very recent approaches like SIGMA-DINO [55], SMILE [63] etc. While VideoMAEv2 [66] introduces additional improvements, it relies on significantly larger backbones and heavy distillation from large teacher models, requiring computational resources beyond our scope and thus we do not directly compare with this in our setting. Therefore, our comparisons focus on methods operating under efficient, distillation-free pretraining settings. SMILE [63] relies on additional synthetic motion signals beyond unlabeled videos. Importantly, since our formulation leverages super images rather than video data, we retain the standard 2D patch embedding without any 3D inflation. This results in a reduction of parameters, as shown in Table 1.

Method Backbone Decoder Epochs Params Top-1 (↑\uparrow)
(M ↓\downarrow)
VideoMAE [64] (NeurIPS’22) ViT-S ✓ 800 22 79.0
SIGMA-DINO [55] (ECCV’24) ViT-S ✓ 800 22 79.4
SMILE (motion) [63] (CVPR’25) ViT-S ✓ 800 22 79.5
VideoMSN\xspace-DeiT (Ours) ViT-S ✗ 50 21.5 80.0
VideoMSN\xspace-DINO (Ours) ViT-S ✗ 10 21 80.8
SVT [53] (CVPR’22) ViT-B ✗ 20 121 78.1
MGM [16] (ICCV’23) ViT-B ✓ 800 87 80.8
CMAE-V [42] (ArXiv’23) ViT-B ✓ 800 87 80.2
MGMAE [31] (CVPR’23) ViT-B ✓ 800 87 81.2
OmniMAE [25] (CVPR’23) ViT-B ✓ 800 87 80.8
ViC-MAE [30] (ECCV’24) ViT-B ✓ 800 87 80.8
ST-MAE [22] (NeurIPS’22) ViT-B ✓ 800 87 81.3
VideoMAE [64] (NeurIPS’22) ViT-B ✓ 800 87 80.0
SIGMA-DINO [55] (ECCV’24) ViT-B ✓ 800 87 81.6
VideoMAE [64] (NeurIPS’22) ViT-B ✓ 1600 87 81.5
MGM [16] (ICCV’23) ViT-B ✓ 1600 87 81.7
MME [59] (CVPR’23) ViT-B ✓ 1600 87 81.8
SMILE (motion) [63] (CVPR’25) ViT-B ✓ 600 87 83.1
VideoMSN\xspace-DeiT (Ours) ViT-B ✗ 50 86.5 82.0
VideoMSN\xspace-DINO (Ours) ViT-B ✗ 10 86 83.3
Table 1: Comparison on Kinetics-400. Our VideoMSN\xspaceis initialized from pretrained image ViT weights followed by pretraining and fine-tuning on Kinetics-400. We achieve superior top-1 accuracy in both VideoMSN\xspace-DeiT and VideoMSN\xspace-DINO after pretraining only for 50 and 10 epochs respectively. Note: The parameter counts for the competing methods are reported from the respective papers and they correspond to the learnable parameters in the encoder. However, most of these methods employ an additional decoder during pretraining, increasing actual training parameters. In contrast, our approach is decoder-free.

4.1 Experimental Results and Analysis

Results on Kinetics-400. We conduct self-supervised pretraining on Kinetics-400 for only 50 and 10 epochs in case of VideoMSN\xspace-DeiT and VideoMSN\xspace-DINO respectively. Remarkably, VideoMSN\xspace-DeiT (ViT-S) surpasses the second best model by 0.5%0.5\% while VideoMSN\xspace-DINO surpasses it by 1.3%1.3\%. For ViT-B architecture, our best model VideoMSN\xspace-DINO improves the state-of-the-art by 0.2%0.2\%. It is worth noting that the second best approach SMILE requires 600 epochs of pretraining (a 60×60\times increase) yet our performance remains superior. For some of the MAE based approaches e.g., VideoMAE, MGM and MME, the increase in pretraining epochs is 160×160\times. We attribute this substantial efficiency gain to the MSN architecture of VideoMSN\xspace, which, unlike encoder-decoder frameworks, eliminates the need for training uninitialized decoder weights from scratch and relies less on low-level detail learning for reconstruction. To compare, we initialize the ViT-B encoder of VideoMAE with ImageNet-21K pretrained weights like ours and run pretraining for 5050 epochs followed by 3030 epochs of finetuning on Kinetics-400. The model, not surprisingly, yields only 57% top-1 accuracy. One possible reason is that such a limited amount of video pretraining may be insufficient for effectively training the randomly initialized decoder. This observation suggests that our decoder-free design is better suited to efficiently adapt strong image-pretrained encoders to video representation learning under short video training schedules. The tendency of MAEs to emphasize on unnecessary low-level details may necessitate longer training.

Dataset Backbone MoCo v3 VideoMAE VideoMSN\xspace-DeiT VideoMSN\xspace-DINO
UCF101 ViT-B 81.7 91.3  (3200 ep) 92.0  (50 ep) 95.4  (10 ep)
HMDB51 ViT-B 39.2 62.6  (4800 ep) 62.6  (50 ep) 70.4  (10 ep)
SSV2 ViT-S - 66.8  (2400 ep) 64.8  (50 ep) 65.8  (10 ep)
SSV2 ViT-B 54.2 70.8  (2400 ep) 69.0  (50 ep) 69.4  (10 ep)
Table 2: Comparisons with the results of previous self-supervised pre-training methods on UCF101, HMDB51, and SSV2, using 16-frame inputs and ViT-S/ViT-B backbones. All methods utilize unlabeled training data for pre-training and consider the labels only for fine-tuning. VideoMSN\xspace  delivers significantly improved performance achieving state-of-the-art results on UCF101 and HMDB51. While slightly trailing on SSV2, our approach offers immense efficiency, very less pretraining epochs across all the datasets compared to VideoMAE. The values reported in this table use a single fixed seed due to computational constraints.

Results on UCF101 and HMDB51. As shown in Table 2, our proposed VideoMSN\xspace consistently achieves superior top-1 recognition accuracy compared to other self-supervised frameworks such as MoCo v3 [11] and VideoMAE on small-scale datasets like UCF101 and HMDB51. Remarkably, while prior methods rely on prohibitively long pretraining schedules (3200 and 4800 epochs for UCF101 and HMDB51, respectively), VideoMSN\xspace-DINO attains higher performance with only 10 epochs of pretraining, yielding a 320×320\times reduction in training epochs on UCF101 and 480×480\times on HMDB51. This result underscores the architectural efficiency and inductive strength of our framework without requiring extensive compute. Such training efficiency is particularly advantageous in regimes with limited labeled data or constrained computational budgets.

Results on SSV2. As shown in Table 2, VideoMSN\xspace-DINO lags behind VideoMAE by a small margin (1.0% for ViT-S and 1.4% for ViT-B). We contend that this is a direct and well-justified trade-off for our framework’s superior computational efficiency. This performance is achieved using 240×240\times fewer pre-training epochs for VideoMSN\xspace-DINO. Achieving competitive results on a difficult, motion-centric benchmark like SSV2 with a fraction of the computational budget highlights the superior scalability and practical utility of our approach.

4.2 Low-shot classification

In this section, we show the generalizability of the proposed self-supervised video representation learning approach. Following established protocols [55, 63], we chose the ViT-B variant of VideoMSN\xspace-DeiT and VideoMSN\xspace-DINO pretrained on 1616 frame input from Kinetics-400 and perform the Low-shot classification experiment. Here, we evaluate the learned representation on action recognition with few training samples per-category. We follow the setup in [63] and finetune with a total of 1000 training examples randomly sampled from UCF101. Table 3 shows the results, in which our method substantially outperforms the rest, demonstrating its strong few-shot action recognition capability in videos.

Dataset VideoMAE MME SIGMA SMILE VideoMSN\xspace VideoMSN\xspace
NeurIPS’22 CVPR’23 ECCV’24 CVPR’25 DeiT DINO
UCF101 74.6 79.2 84.1 86.4 88.0 89.0
Table 3: Comparison of low-shot action recognition on UCF101 using only 10001000 training videos, 16-frame inputs and ViT-B backbones for fine-tuning. Despite using substantially fewer video pretraining epochs, both VideoMSN\xspace-DeiT and VideoMSN\xspace-DINO outperform prior approaches, with VideoMSN\xspace-DINO achieving the best top-1 accuracy of 89.0%89.0\%, improving over SMILE by 2.6%2.6\%. The results highlight the strong transferability and label efficiency of the representations learned by VideoMSN\xspace.

4.3 Additional Experiments and Ablation Studies

In this subsection, we first provide additional analysis to better understand the contribution of the proposed VideoMSN\xspacepretraining on top of strong image-pretrained initialization. We then present comprehensive ablation studies to validate the effectiveness of the different design choices in VideoMSN\xspace. Unless otherwise specified, all experiments are conducted using VideoMSN\xspace-DeiT with a ViT-S backbone on Kinetics-400 with 16-frame input. The pretrained models are fine-tuned for 30 epochs and evaluated using a consistent inference protocol of 55 clips ×\times 33 crops.

Method DeiT-S DeiT-B DINO-S DINO-B
w/o VideoMSN 77.3 78.2 80.7 83.1
w/ VideoMSN 80.0 82.0 80.8 83.3
Table 4: Effect of VideoMSN\xspacepretraining on Kinetics-400. All models are initialized from the same image-pretrained DeiT-v3 or DINO-v3 checkpoints. “w/o VideoMSN\xspace” denotes direct fine-tuning of the image-pretrained backbone for action recognition, whereas “w/ VideoMSN\xspace” first performs the proposed lightweight self-supervised video pretraining stage before fine-tuning. VideoMSN\xspaceconsistently improves performance across all backbones, with particularly large gains for DeiT-S and DeiT-B (+2.7% and +3.8% top-1 accuracy, respectively), while also providing modest improvements for the already strong DINO-S and DINO-B initializations.

Effect of VideoMSN\xspace Pretraining over Image-pretrained Initialization.

We analyze the impact of the proposed VideoMSN\xspace pretraining by starting from strong image-pretrained DeiT-v3 and DINO-v3 checkpoints and evaluating two training settings on Kinetics-400. In the first setting (w/o VideoMSN\xspace), the image-pretrained backbone is directly fine-tuned for action recognition without any self-supervised video adaptation. In the second setting (w/ VideoMSN\xspace), we first perform the proposed VideoMSN\xspace self-supervised video pretraining stage before fine-tuning on Kinetics-400. As shown in Table 4, VideoMSN\xspace consistently improves performance across all image-pretrained initializations. The gains are particularly pronounced for DeiT-S and DeiT-B, yielding improvements of +2.7%+2.7\% and +3.8%+3.8\% top-1 accuracy, respectively. Even for the stronger DINO-v3 initialization, VideoMSN\xspace provides additional improvements, demonstrating that lightweight self-supervised video adaptation can effectively enhance image-pretrained representations for spatio-temporal video understanding and help in the case where adaptation in a short time is necessary.

Method Epochs (↓\downarrow) Wall-clock Time (↓\downarrow) Total FLOPs (↓\downarrow)
VideoMAE 1600 266.7 h 40.89 E
VideoMSN\xspace-DINO (ours) 10 21.7 h 4.58 E
Reduction 160×\times 12.3×\times 8.9×\times
Table 5: Total pretraining epochs, wall-clock time, and FLOPs for VideoMAE and VideoMSN\xspace-DINO on Kinetics-400, showing the substantially lower overall training cost of VideoMSN\xspace.

Training Time Compute. To better quantify efficiency of VideoMSN\xspace, we report the total video pretraining time and total FLOPs in Table 5 using the same hardware (4×\!\times\!H100 GPUs) for VideoMAE and VideoMSN\xspace. The wall-clock time is measured as # epochs (ep) ×\!\times\! time/epoch (t/ep) giving 266.7 hours for VideoMAE (1600 ep×\!\times\!10 min/ep) and 21.7 hours for VideoMSN\xspace-DINO (10 ep×\!\times\!130 min/ep); a reduction in video pretraining time by 12.3×\!\times\!. Total FLOPs is given by GFLOP/sample ×\times sample/ep ×\times ep. VideoMAE requires (×240K×1600)≈40.89(106.5\!\times\!240\mathrm{K}\!\times\!1600)\!\approx\!40.89 EFLOPs, whereas VideoMSN\xspace-DINO requires (×240K×10)≈4.58(1911.9\!\times\!240\mathrm{K}\!\times\!10)\!\approx\!4.58 EFLOPs implying a reduction of ∼⁣×\!\sim\!8.9\!\times\! in total FLOPs. Although VideoMSN\xspace has higher FLOPs per epoch, the substantially shorter video pretraining schedule lowers the total computation.

Approach Top-1
w/o Temp Aug 78.6
w/ Temp Aug 80.0
Table 6: (a) Temporal Augmentations: We evaluate the impact of temporal masking on representation learning. Incorporating temporal augmentations by using 3×3 and 2×2 focal view improves top-1 accuracy by 1.4 points over the baseline without augmentation.
Approach Top-1
w/o Sinkhorn 80.0
w/ Sinkhorn 79.1
Table 6: (b) Sinkhorn Normalization: We analyze the impact of Sinkhorn normalization and observe that excluding it leads to a 0.9% gain in top-1 accuracy, indicating that the ME-MAX regularizer alone is sufficient to preserve representation diversity in our model.
Anchor Views Top-1
Only Random View 78.5
Only Focal View 77.1
Both Views 80.0
Table 6: (c) Anchor View: ‘Both Views’ (Random and Focal) perform better, achieving the optimal 80.0% accuracy. ‘Both Views’ outperforms using ‘Only Random View’ (78.5%) or ‘Only Focal View’ (77.1%) in isolation, validating our multi-view masking design.
Focal Views Top-1
Only Fast Views 78.6
Only Slow Views 77.5
Both Views 80.0
Table 6: (d) Focal Views: Use of both fast and slow views performs the best (80.0%), which significantly outperforms only fast (78.6%) or only slow views (77.5%). It demonstrates the importance of training the model on varied temporal sampling rates.
Table 6: Ablation studies on Kinetics-400 with a 16-frame ViT-S backbone. All models are pre-trained for 50 epochs and fully fine-tuned for evaluation. We adopt a consistent inference protocol of 5 clips × 3 crops. The default configuration (highlighted) achieves optimal performance across key components: (a) Temporal augmentations, (b) Sinkhorn normalization, (c) Anchor view, and (d) Focal views. The values reported in this table use a single fixed seed due to computational constraints.

Effect of Temporal Augmentations. We examine how adding temporal augmentations affects the performance of our VideoMSN\xspace-DeiT model. Table  shows that incorporating temporal masking boosts accuracy by 1.4 percentage points compared to the variant that omits this augmentation. As depicted in Fig. 3, temporal masking employs a hierarchical scheme to create six small focal views. Starting with 16 uniformly sampled frames, we drop 7 frames to build 3×33\times 3 super images which act as the fast views. We create 3 such fast views. Similarly, we drop additional 5 frames (i.e., 12 frames are dropped altogether) to build 2×22\times 2 super images which act as the slow views. We created 3 different views for this case also. The outcome of this augmentation appears in the second row of Table . For the baseline without temporal augmentation, all 16 frames are retained for each of the six focal views. All such focal views, in this baseline, are arranged as a 4×44\times 4 super image. This configuration leads to a 1.41.4 percentage point decrease in performance, as listed in the first row of Table . We attribute the improvement to the ability of the temporal augmentation to encourage the model to learn motion-aware features. Since 4×44\times 4 focal views may dilute motion cues and action semantics by overly compressing frames, using smaller grids like 3×33\times 3 and 2×22\times 2 helps preserve and highlight the actions of the video.

Effect of Sinkhorn Normalization. For VideoMSN\xspace, we follow the default setting of Masked Siamese Network and set the ME-MAX regularization weight λ\lambda to 5.05.0. However, in this experiment, we explore the interaction between this regularization strength and Sinkhorn normalization. To investigate this, we study the impact of Sinkhorn normalization during pretraining. Table  shows that omitting Sinkhorn while keeping λ\lambda to 5.05.0 produces the best performance. This result indicates that Sinkhorn normalization may interfere with optimal feature learning when used alongside strong regularization.

Ablation on Anchor Views. Table  compares the efficacy of our masking strategies using our VideoMSN\xspace-DeiT (ViT-S). The results clearly show a synergistic effect: using only random views yields 78.5% accuracy, and using only focal views results in 77.1%. The combination of both random and focal views, as used in our full model, produces the highest Top-1 accuracy of 80.0%.

Ablation on Focal Views. Table  evaluates the impact of our temporal masking strategy. The results demonstrate that using both Fast and Slow focal views along with the Random view achieves the best top-1 accuracy of 80%. This outperforms configurations using only Fast views (78.6%) or only Slow views (77.5%) along with the Random view, highlighting the importance of learning from different temporal granularities. Additional experimental results are provided in the appendix.

5 Conclusion

In this work, we introduce VideoMSN\xspace, the first adaptation of Masked Siamese Networks to video representation learning. By representing video frames as 2D super images composed of frames sampled from videos, the method enables standard 2D Vision Transformers to learn effective spatio-temporal representations without relying on computationally expensive 3D architectures. The proposed decoder-free formulation replaces pixel-level reconstruction with feature alignment between masked and unmasked views, significantly reducing the computational cost of self-supervised pretraining. Our experiments demonstrate that VideoMSN\xspace not only outperforms competing methods but does so with very less pre-training with videos. VideoMSN\xspace demonstrates strong performance in low-shot classification confirming the quality and transferability of the learned representations. More broadly, our findings indicate that strong image foundation models can be efficiently adapted to the video domain through lightweight self-supervised learning, reducing the need for extensive video pretraining while maintaining competitive performance. We hope this work motivates further research into efficient video adaptation strategies that leverage the rapidly growing ecosystem of image-pretrained foundation models.

6 Acknowledgement

This work was partially supported by ANRF Grant CRG/2023/005010. We acknowledge the National Supercomputing Mission (NSM) for providing computational resources via the DGX GPU Cluster at IIT Kharagpur and C-DAC for providing additional computing support through the ParamRudra cluster.

References

  • [1] A. Arnab, M. Dehghani, G. Heigold, C. Sun, M. Lučić, and C. Schmid (2021) Vivit: A video vision transformer. In IEEE international conference on computer vision, pp. 6836–6846. Cited by: §1.
  • [2] M. Assran, M. Caron, I. Misra, P. Bojanowski, F. Bordes, P. Vincent, A. Joulin, M. Rabbat, and N. Ballas (2022) Masked Siamese Networks for Label-efficient Learning. In European conference on computer vision, pp. 456–473. Cited by: §C.2, §1, §2, §3.1, §3.2, §3.3, §3.3, §3, §4.
  • [3] M. Assran, M. Caron, I. Misra, P. Bojanowski, A. Joulin, N. Ballas, and M. Rabbat (2021) Semi-supervised learning of visual features by non-parametrically predicting view assignments with support samples. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 8443–8452. Cited by: §3.3, §3.3, §3.3.
  • [4] M. Assran, Q. Duval, I. Misra, P. Bojanowski, P. Vincent, M. Rabbat, Y. LeCun, and N. Ballas (2023) Self-supervised Learning from Images with a Joint Embedding Predictive Architecture. In IEEE Conference on Computer Vision and Pattern Recognition, pp. 15619–15629. Cited by: §1, §2.
  • [5] A. Bar, F. Bordes, A. Shocher, M. Assran, P. Vincent, N. Ballas, T. Darrell, A. Globerson, and Y. Lecun (2024) Stochastic Positional Embeddings Improve Masked Image Modeling. In International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 235, pp. 2944–2958. Cited by: §1.
  • [6] S. Benaim, A. Ephrat, O. Lang, I. Mosseri, W. T. Freeman, M. Rubinstein, M. Irani, and T. Dekel (2020) Speednet: learning the speediness in videos. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 9922–9931. Cited by: §2.
  • [7] G. Bertasius, H. Wang, and L. Torresani (2021) Is space-time attention all you need for video understanding?. In Proceedings of the 38th International Conference on Machine Learning, Vol. 139, pp. 813–824. Cited by: §1.
  • [8] J. Bromley, I. Guyon, Y. LeCun, E. Säckinger, and R. Shah (1993) Signature verification using a ”siamese” time delay neural network. In Proceedings of the 7th International Conference on Neural Information Processing Systems, NIPS’93, San Francisco, CA, USA, pp. 737–744. Cited by: §2.
  • [9] M. Caron, I. Misra, J. Mairal, P. Goyal, P. Bojanowski, and A. Joulin (2020) Unsupervised learning of visual features by contrasting cluster assignments. Advances in neural information processing systems 33, pp. 9912–9924. Cited by: §3.3.
  • [10] M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, and A. Joulin (2021) Emerging properties in self-supervised vision transformers. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 9650–9660. Cited by: §2, §3.3.
  • [11] X. Chen, S. Xie, and K. He (2021) An empirical study of training self-supervised vision transformers. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Vol. , pp. 9620–9629. External Links: Document Cited by: §4.1.
  • [12] H. Cho, T. Kim, H. J. Chang, and W. Hwang (2021) Self-supervised visual learning by variable playback speeds prediction of a video. IEEE Access 9, pp. 79562–79571. Cited by: §2.
  • [13] C. Doersch, A. Gupta, and A. A. Efros (2015) Unsupervised Visual Representation Learning by Context Prediction. In IEEE international conference on computer vision, pp. 1422–1430. Cited by: §1.
  • [14] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby (2021) An image is worth 16x16 words: transformers for image recognition at scale. In International Conference on Learning Representations, External Links: Link Cited by: §3.3, §4.
  • [15] A. Dosovitskiy, P. Fischer, J. T. Springenberg, M. Riedmiller, and T. Brox (2016) Discriminative Unsupervised Feature Learning with Exemplar Convolutional Neural Networks. IEEE Transactions on Pattern Analysis and Machine Intelligence 38 (9), pp. 1734–1747. Cited by: §1.
  • [16] D. Fan, J. Wang, S. Liao, Y. Zhu, V. Bhat, H. Santos-Villalobos, R. MV, and X. Li (2023) Motion-guided masking for spatiotemporal representation learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 5619–5629. Cited by: §1, §2, Table 1, Table 1.
  • [17] H. Fan, B. Xiong, K. Mangalam, Y. Li, Z. Yan, J. Malik, and C. Feichtenhofer (2021) Multiscale vision transformers. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 6824–6835. Cited by: §2.
  • [18] Q. Fan, C. Chen, and R. Panda (2022) Can an Image Classifier Suffice for Action Recognition?. In International Conference on Learning Representations, Cited by: §1, §2, §3.2, §3.2, §4.
  • [19] Q. Fan, C. R. Chen, H. Kuehne, M. Pistoia, and D. Cox (2019) More Is Less: Learning Efficient Video Representations by Big-Little Network and Depthwise Temporal Aggregation. In Neural Information Processing Systems, pp. 2261–2270. Cited by: §2.
  • [20] C. Feichtenhofer, H. Fan, J. Malik, and K. He (2019) Slowfast Networks for Video Recognition. In IEEE/CVF international conference on computer vision, pp. 6202–6211. Cited by: footnote 1.
  • [21] C. Feichtenhofer, H. Fan, B. Xiong, R. Girshick, and K. He (2021) A large-scale study on unsupervised spatiotemporal representation learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 3299–3309. Cited by: §2.
  • [22] C. Feichtenhofer, Y. Li, K. He, et al. (2022) Masked autoencoders as spatiotemporal learners. Advances in neural information processing systems 35, pp. 35946–35958. Cited by: Table 1.
  • [23] B. Fernando, H. Bilen, E. Gavves, and S. Gould (2017) Self-supervised video representation learning with odd-one-out networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 5729–5738. Cited by: §2.
  • [24] S. Gidaris, P. Singh, and N. Komodakis (2018) Unsupervised Representation Learning by Predicting Image Rotations. In International Conference on Learning Representations, External Links: Link Cited by: §1.
  • [25] R. Girdhar, A. El-Nouby, M. Singh, K. V. Alwala, A. Joulin, and I. Misra (2023) Omnimae: single model masked pretraining on images and videos. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 10406–10417. Cited by: Table 1.
  • [26] R. Goyal, S. Ebrahimi Kahou, V. Michalski, J. Materzynska, S. Westphal, H. Kim, V. Haenel, I. Fruend, P. Yianilos, M. Mueller-Freitag, et al. (2017) The “Something Something” Video Database for Learning and Evaluating Visual Common Sense. In Proceedings of the IEEE international conference on computer vision, pp. 5842–5850. Cited by: Appendix A, §4.
  • [27] J. Grill, F. Strub, F. Altché, C. Tallec, P. Richemond, E. Buchatskaya, C. Doersch, B. Avila Pires, Z. Guo, M. Gheshlaghi Azar, et al. (2020) Bootstrap your own latent-a new approach to self-supervised learning. Advances in neural information processing systems 33, pp. 21271–21284. Cited by: §3.3.
  • [28] K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick (2022) Masked autoencoders are scalable vision learners. In IEEE conference on computer vision and pattern recognition, pp. 16000–16009. Cited by: §1, §2.
  • [29] K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick (2020) Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 9729–9738. Cited by: §2.
  • [30] J. Hernandez, R. Villegas, and V. Ordonez (2024) Vic-mae: self-supervised representation learning from images and video with contrastive masked autoencoders. In European Conference on Computer Vision, pp. 444–463. Cited by: Table 1.
  • [31] B. Huang, Z. Zhao, G. Zhang, Y. Qiao, and L. Wang (2023) Mgmae: motion guided masking for video masked autoencoding. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 13493–13504. Cited by: §2, Table 1.
  • [32] O. Iqbal, O. Chakraborty, A. Hussain, R. Panda, and A. Das (2024) SITAR: Semi-supervised Image Transformer for Action Recognition. In International conference on pattern recognition, pp. 114–130. Cited by: §2.
  • [33] L. Jing, X. Yang, J. Liu, and Y. Tian (2018) Self-supervised spatiotemporal feature learning via video rotation prediction. arXiv preprint arXiv:1811.11387. Cited by: §2.
  • [34] A. Joulin and F. Bach (2012) A convex relaxation for weakly supervised classifiers. arXiv preprint arXiv:1206.6413. Cited by: §3.3, §3.3.
  • [35] W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, P. Natsev, et al. (2017) The Kinetics Human Action Video Dataset. arXiv preprint arXiv:1705.06950. Cited by: Appendix A, §4.
  • [36] D. Kim, D. Cho, and I. S. Kweon (2019) Self-supervised video representation learning with space-time cubic puzzles. In Proceedings of the AAAI conference on artificial intelligence, Vol. 33, pp. 8545–8552. Cited by: §2.
  • [37] H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre (2011) HMDB: a large video database for human motion recognition. In 2011 International Conference on Computer Vision, Vol. , pp. 2556–2563. External Links: Document Cited by: Appendix A, §4.
  • [38] G. Larsson, M. Maire, and G. Shakhnarovich (2017) Colorization as a Proxy Task for Visual Understanding. In IEEE conference on computer vision and pattern recognition, pp. 6874–6883. Cited by: §1.
  • [39] J. Lin, C. Gan, and S. Han (2019) TSM: Temporal Shift Module for Efficient Video Understanding. In IEEE International Conference on Computer Vision, pp. 7083–7093. Cited by: §2.
  • [40] Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo (2021) Swin transformer: hierarchical vision transformer using shifted windows. In 2021 IEEE/CVF international conference on computer vision (ICCV), pp. 9992–10002. Cited by: §2.
  • [41] I. Loshchilov and F. Hutter (2017) Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101. Cited by: §4.
  • [42] C. Lu, X. Jin, Z. Huang, Q. Hou, M. Cheng, and J. Feng (2023) CMAE-V: Contrastive Masked Autoencoders for Video Action Recognition. arXiv preprint arXiv:2301.06018. Cited by: Table 1.
  • [43] D. Luo, C. Liu, Y. Zhou, D. Yang, C. Ma, Q. Ye, and W. Wang (2020) Video cloze procedure for self-supervised spatio-temporal learning. In Proceedings of the AAAI conference on artificial intelligence, Vol. 34, pp. 11701–11708. Cited by: §2.
  • [44] Y. Meng, C. Lin, R. Panda, P. Sattigeri, L. Karlinsky, A. Oliva, K. Saenko, and R. Feris (2020) AR-Net: Adaptive Frame Resolution for Efficient Action Recognition. In European Conference on Computer Vision, pp. 86–104. Cited by: §2.
  • [45] I. Misra, C. L. Zitnick, and M. Hebert (2016) Shuffle and learn: unsupervised learning using temporal order verification. In European conference on computer vision, pp. 527–544. Cited by: §2.
  • [46] D. Neimark, O. Bar, M. Zohar, and D. Asselmann (2021) Video transformer network. In 2021 IEEE/CVF international conference on computer vision workshops (ICCVW), pp. 3156–3165. Cited by: §2.
  • [47] D. Nguyen, V. Aggarwal, Y. Li, M. R. Oswald, A. Kirillov, C. G. Snoek, and X. Chen (2023) R-mae: regions meet masked autoencoders. arXiv preprint arXiv:2306.05411. Cited by: §2.
  • [48] M. Noroozi and P. Favaro (2016) Unsupervised Learning of Visual Representations by Solving Jigsaw Puzzles. In European conference on computer vision, pp. 69–84. Cited by: §1.
  • [49] M. Noroozi, H. Pirsiavash, and P. Favaro (2017) Representation Learning by Learning to Count. In IEEE international conference on computer vision, pp. 5898–5906. Cited by: §1.
  • [50] T. Pan, Y. Song, T. Yang, W. Jiang, and W. Liu (2021) Videomoco: contrastive video representation learning with temporally adversarial examples. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 11205–11214. Cited by: §2.
  • [51] D. Pathak, P. Krahenbuhl, J. Donahue, T. Darrell, and A. A. Efros (2016) Context encoders: feature learning by inpainting. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2536–2544. Cited by: §2.
  • [52] Z. Qiu, T. Yao, Y. Shu, C. Ngo, and T. Mei (2021) Condensing A Sequence to One Informative Frame for Video Recognition. In IEEE International Conference on Computer Vision, pp. 16311–16320. Cited by: §2.
  • [53] K. Ranasinghe, M. Naseer, S. Khan, F. S. Khan, and M. S. Ryoo (2022) Self-supervised video transformer. In IEEE conference on computer vision and pattern recognition, pp. 2874–2884. Cited by: §2, Table 1.
  • [54] M. Safaei and H. Foroosh (2019) Still Image Action Recognition by Predicting Spatial-Temporal Pixel Evolution. In Winter Conference on Applications of Computer Vision, pp. 111–120. Cited by: §2.
  • [55] M. Salehi, M. Dorkenwald, F. M. Thoker, E. Gavves, C. G. Snoek, and Y. M. Asano (2024) Sigma: sinkhorn-guided masked video modeling. In European Conference on Computer Vision, pp. 293–312. Cited by: §4.2, Table 1, Table 1, §4.
  • [56] O. Siméoni, H. V. Vo, M. Seitzer, F. Baldassarre, M. Oquab, C. Jose, V. Khalidov, M. Szafraniec, S. Yi, M. Ramamonjisoa, et al. (2025) DINOv3. arXiv preprint arXiv:2508.10104. Cited by: §4.
  • [57] K. Soomro, A. R. Zamir, and M. Shah (2012) Ucf101: a dataset of 101 human actions classes from videos in the wild. arXiv preprint arXiv:1212.0402. Cited by: Appendix A, §4.
  • [58] X. Sun, R. Panda, C. R. Chen, A. Oliva, R. Feris, and K. Saenko (2021) Dynamic Network Quantization for Efficient Video Inference. In IEEE International Conference on Computer Vision, pp. 7375–7385. Cited by: §2.
  • [59] X. Sun, P. Chen, L. Chen, C. Li, T. H. Li, M. Tan, and C. Gan (2023) Masked Motion Encoding for Self-supervised Video Representation Learning. In IEEE conference on computer vision and pattern recognition, pp. 2235–2245. Cited by: §1, §2, §2, Table 1.
  • [60] C. Tao, X. Zhu, W. Su, G. Huang, B. Li, J. Zhou, Y. Qiao, X. Wang, and J. Dai (2023) Siamese Image Modeling for Self-Supervised Vision Representation Learning. In IEEE Conference on Computer Vision and Pattern Recognition, pp. 2132–2141. Cited by: §1.
  • [61] M. Tavakolian, M. Sabokrou, and A. Hadid (2019) AVD: Adversarial Video Distillation. arXiv preprint arXiv:1907.05640. Cited by: §2.
  • [62] M. Tavakolian, H. R. Tavakoli, and A. Hadid (2019) AWSD: Adaptive Weighted Spatiotemporal Distillation for Video Representation. In IEEE International Conference on Computer Vision, pp. 8019–8028. Cited by: §2.
  • [63] F. M. Thoker, L. Jiang, C. Zhao, and B. Ghanem (2025) SMILE: Infusing Spatial and Motion Semantics in Masked Video Learning. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 8438–8449. Cited by: Figure 1, §2, §4.2, Table 1, Table 1, §4.
  • [64] Z. Tong, Y. Song, J. Wang, and L. Wang (2022) VideoMAE: Masked Autoencoders are Data-efficient Learners for Self-supervised Video Pre-training. Advances in neural information processing systems 35, pp. 10078–10093. Cited by: Figure 1, §1, §1, §2, §2, §3.2, Table 1, Table 1, Table 1, §4.
  • [65] H. Touvron, M. Cord, and H. Jégou (2022) DeiT III: Revenge of the ViT. In European conference on computer vision, pp. 516–533. Cited by: §4.
  • [66] L. Wang, B. Huang, Z. Zhao, Z. Tong, Y. He, Y. Wang, Y. Wang, and Y. Qiao (2023) Videomae v2: scaling video masked autoencoders with dual masking. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 14549–14560. Cited by: §4.
  • [67] L. Wang, Y. Xiong, Z. Wang, Y. Qiao, D. Lin, X. Tang, and L. V. Gool (2016) Temporal Segment Networks: Towards Good Practices for Deep Action Recognition. In European conference on computer vision, pp. 20–36. Cited by: §2, §3.2.
  • [68] R. Wang, D. Chen, Z. Wu, Y. Chen, X. Dai, M. Liu, Y. Jiang, L. Zhou, and L. Yuan (2022) BEVT: BERT Pretraining of Video Transformers. In IEEE conference on computer vision and pattern recognition, pp. 14733–14743. Cited by: §2.
  • [69] R. Wang, D. Chen, Z. Wu, Y. Chen, X. Dai, M. Liu, L. Yuan, and Y. Jiang (2023) Masked Video Distillation: Rethinking Masked Feature Modeling for Self-supervised Video Representation Learning. In IEEE conference on computer vision and pattern recognition, pp. 6312–6322. Cited by: §2.
  • [70] Z. Wu, C. Xiong, C. Ma, R. Socher, and L. S. Davis (2019) Adaframe: Adaptive Frame Selection for Fast Video Recognition. In IEEE Conference on Computer Vision and Pattern Recognition, pp. 1278–1287. Cited by: §2.
  • [71] S. Yun, D. Han, S. Chun, S. J. Oh, Y. Yoo, and J. Choe (2019) CutMix: regularization strategy to train strong classifiers with localizable features. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV), Vol. , pp. 6022–6031. External Links: Document Cited by: §4.
  • [72] H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz (2018) Mixup: beyond empirical risk minimization. External Links: 1710.09412 Cited by: §4.
  • [73] R. Zhang, P. Isola, and A. A. Efros (2016) Colorful Image Colorization. In European conference on computer vision, pp. 649–666. Cited by: §1.
  • [74] Y. Zhang, X. Li, C. Liu, B. Shuai, Y. Zhu, B. Brattoli, H. Chen, I. Marsic, and J. Tighe (2021) Vidtr: Video Transformer without Convolutions. In IEEE international conference on computer vision, pp. 13577–13587. Cited by: §2.
  • [75] Z. Zhao, H. Ma, and S. You (2017) Single Image Action Recognition using Semantic Body Part Actions. In IEEE international conference on computer vision, pp. 3391–3399. Cited by: §2.
  • [76] B. Zhou, A. Andonian, A. Oliva, and A. Torralba (2018) Temporal Relational Reasoning in Videos. In European Conference on Computer Vision (ECCV), pp. 803–818. Cited by: §2.
  • [77] J. Zhou, C. Wei, H. Wang, W. Shen, C. Xie, A. Yuille, and T. Kong (2022) iBOT: Image BERT Pre-Training with Online Tokenizer. In International Conference on Learning Representations, Cited by: §2.

Appendix A Dataset Description

UCF101. The UCF101 [57] dataset is a widely-used benchmark for action recognition, comprising 13,320 unconstrained video clips collected from YouTube. It spans 101 human action categories, hierarchically grouped into five broad types: (i) human-object interaction, (ii) body-motion only, (iii) human-human interaction, (iv) playing musical instruments, and (v) sports activities. Each video is encoded at 25 frames per second (fps) with a resolution of 320×\times240 pixels, and has an average duration of approximately 7.2 seconds. The dataset is structured into 25 groups, where each group contains 4 to 7 clips per action class, sharing commonalities in background, camera viewpoint, and other visual conditions. This grouping aids in benchmarking models under intra-class variations. The dataset is publicly accessible at: https://www.crcv.ucf.edu/data/UCF101.php.

HMDB51. The HMDB51 [37] dataset (Human Motion Database) is a curated collection of 6,766 video clips, covering 51 distinct human action categories, with a minimum of 101 clips per class. Videos are sourced from a variety of real-world settings including movies, public databases, and YouTube. All clips are standardized to 30 fps, with the frame height fixed at 240 pixels, and width adjusted to preserve aspect ratio. The action categories are organized into five high-level groups: (i) General Facial Actions, (ii) Facial Actions with Object Manipulation, (iii) General Body Movements, (iv) Body Movements with Object Interaction, and (v) Body Movements for Human Interaction. This dataset poses significant challenges due to its diverse scenes, viewpoints, and motion dynamics, making it a valuable benchmark for action recognition research. It is publicly available at: https://serre-lab.clps.brown.edu/resource/hmdb-a-large-human-motion-database/.

Kinetics-400. The Kinetics-400 [35] dataset, curated by DeepMind, is one of the largest and most diverse datasets for video action recognition. It comprises 306,245 video clips uniformly distributed across 400 action categories, with each class containing at least 400 clips. Videos have an average duration of 10 seconds, and depict actions in unconstrained environments sourced from YouTube. The actions are broadly categorized into three types: (i) person actions (e.g., drawing, drinking, laughing), (ii) person-person interactions (e.g., hugging, kissing, shaking hands), and (iii) person-object interactions (e.g., mowing the lawn, opening a present, washing dishes). Kinetics-400 sets a high standard for large-scale video understanding, particularly in real-world, diverse contexts. The dataset is publicly available at: https://deepmind.com/research/open-source/kinetics.

Something-Something V2. The Something-Something V2 (SSV2) [26] dataset is a challenging, large-scale benchmark for video action recognition. With over 220K videos and 174 action classes, it focuses on fine-grained human-object interactions. Unlike other datasets, SSV2 is considered motion-heavy due to its emphasis on the nuanced gestures and directional aspects of actions, such as ”turning something upside down” or ”pushing something from left to right”. This design makes it a rigorous test for models that need to understand motion and context. The dataset is publicly available at: https://www.qualcomm.com/developer /software/something-something-v-2-dataset.

lr Top-1
1e-4 68.6
1e-5 69.1
1e-6 66.2
Table 7: (a) Learning Rate: We evaluate different learning rates and observe that 1​e−51e^{-5} yields the best performance.
Weight Decay Top-1
0.05 68.8
0.01 69.1
0.1 68.9
Table 7: (b) Weight Decay: A decay value of 0.01 provides the best trade-off.
Patch Drop BS/GPU Top-1
0.3 8 68.9
0.5 10 68.9
0.7 12 69.1
0.9 14 68.8
Table 7: (c) Patch Drop: Increasing patch drop enables larger batch sizes and improves generalization.

# Views Throughput Top-1 6 42 v/s 69.1 8 35 v/s 69.2 10 32 v/s 69.3


Table 7: (d) Focal Views: More views improve accuracy but reduce throughput. We select 6 views for the best trade-off.
Table 7: Experiments to determine optimal hyperparameter settings on a class-wise uniformly sampled 25% subset of the Kinetics-400 dataset with VideoMSN\xspace-DeiT (ViT-S) using 4×44\times 4 superimage. We analyze the effect of (a) learning rate, (b) weight decay, (c) patch drop, and (d) focal views. The default configurations (highlighted) are the ones that achieves the best overall performance.

Appendix B Impact of Hyperparameters

In this section, we analyze the impact of key model hyperparameters, including (a) learning rate, (b) weight decay, (c) patch drop ratio, and (d) the number of focal views on Kinetics-400 dataset using VideoMSN\xspace-DeiT with ViT-S backbone as the default backbone. Experiments, in the appendix are run using a single fixed seed due to computational constraints.

B.1 Effect of Learning Rate (lr)

Table  presents an ablation study evaluating the influence of different learning rates (lr) on the Top-1 accuracy (Top-1). A learning rate of 1​e−51e-5 yields the highest Top-1 accuracy of 69.1%, indicating it is the most optimal among the tested configurations. Increasing the learning rate to 1​e−41\mathrm{e}{-4} lowers accuracy to 68.6%, while decreasing it to 1​e−61\mathrm{e}{-6} further drops accuracy to 66.2%.

B.2 Effect of Weight Decay

Table  reports the impact of varying weight decay values on Top-1 accuracy. A weight decay of 0.01 achieves the best performance with 69.1% Top-1 accuracy. Using 0.05 or 0.1 slightly reduces performance to 68.8% and 68.9%, respectively. This suggests that 0.01 is an optimal choice for regularization in our setup.

B.3 Effect of Patch Drop Ratio

Table  presents an ablation study on the impact of varying patch drop ratios, where a fixed proportion of input spatio-temporal patches are masked (Tube masking) during training. Increasing the patch drop ratio reduces the number of visible tokens passed to the encoder, thereby lowering per-sample memory and compute requirements. This allows for larger batch sizes (BS) per GPU, as reflected in the table. Among the tested values, a drop ratio of 0.7 yields the best Top-1 accuracy of 69.1%, indicating an effective trade-off between information sparsity and model learning capacity (more towards increasing BS/GPU). Lower ratios—0.3 and 0.5—retain 70% and 50% of the patches, respectively, and result in slightly reduced accuracies of 68.9% for both, corresponding to performance drops of 0.2% for both. A higher ratio of 0.9 retains only 10% of the patches, leading to a performance drop of 0.2% (Top-1: 68.8%) likely due to excessive loss of informative content.

B.4 Effect of Number of Focal Views

Table  presents an ablation study evaluating the impact of varying the number of focal views on both Top-1 accuracy and inference throughput (measured in videos per second). As the number of focal views increases from 6 to 10, we observe a slight improvement in Top-1 accuracy—from 69.1% to 69.3%—suggesting that additional focal views provide marginal gains in performance. However, this comes at the cost of significantly reduced throughput: from 42 videos/s at 6 views down to 32 videos/s at 10 views. This trade-off highlights a key design consideration: while higher numbers of focal views can improve accuracy, they also impose greater computational overhead. The configuration with 6 focal views offers the best balance between accuracy and efficiency, making it preferable in scenarios where both performance and scalability are critical.

Appendix C Additional Ablations

Tables 8, 9 presents a set of ablation experiments conducted on the full Kinetics-400 dataset using a ViT-S backbone. The purpose of these studies is to validate and optimize key hyperparameters and design choices for the VideoMSN\xspace-DeiT model.

Patch Drop Dataset Top-1 Acc. (%)
70% SSv2 64.8
50% SSv2 65.8
25% SSv2 65.3
Table 8: Effect of different patch drop ratios during the pretraining of our VideoMSN\xspace-DINO model with a ViT-S backbone on the SSv2 dataset.
Focal Views Top-1 Acc. (%)
No Temporal Masking 65.8
Fast View Only 65.7
Slow View Only 65.5
Both Views 65.8
Table 9: Effect of different focal view configurations during pretraining of our VideoMSN\xspace-DINO model with a ViT-S backbone using a patch drop ratio of 0.5 on the SSv2 dataset.
Masking Strategy Top-1 Acc.(%)
Random Masking 79.0
Tube Masking 80.0

Table 10: We compare the tube masking with the conventional random masking and observe a 1.0% gain in top-1 accuracy. Tube masking enforces consistent spatiotemporal masking by applying the identical spatial masks across frames.
No. of Prototypes Top-1
512 78.6
1024 80.0
2048 80.0

Table 11: Effect of Number of Prototypes: On increasing the prototype count from 512 to 1024 provides a significant 1.4% performance boost (78.6% to 80.0%). However, a further increase to 2048 yields no additional gain, establishing 1024 as the optimal configuration.

C.1 Effect of Tube Masking.

We find that tube masking achieves better performance with an increment of 1% (ref. Table 10) than plain random masking. We attribute these interesting observations to the redundancy and temporal correlation in videos. Tube masking effectively addresses the temporal redundancy in videos by masking continuous spatio-temporal regions, rather than random patches. As seen in Fig. 2, from a super image SiS^{i} when we get a Random View RiR^{i}, we decide for a random mask for the first frame and then apply the same mask in all the frames of the formed super image. Then, all patches at the same spatial location across different temporal indices are either simultaneously masked or retained. This reduces information leakage caused by frame-to-frame similarity and prevents the model from learning shortcut features. As a result, it encourages the learning of meaningful spatio-temporal representations and enables successful training even with simple backbones like vanilla ViT on small-scale datasets. These findings are consistent with the results of prior works like VideoMAE, which also demonstrated that tube masking significantly improves performance compared to randomly removing patches from all the frames in video masked image modeling.

C.2 Effect of Number of Prototypes.

Table 11 investigates the influence of the number of learnable prototypes of our VideoMSN\xspace-DeiT (ViT-S) performance. The purpose of this study is to validate the choice of this key hyperparameter and design choice for the VideoMSN\xspace-DeiT model. The results show that while increasing the number of prototypes from 512 to 1024 significantly improves the Top-1 accuracy from 78.6% to 80.0%, a further increase to 2048 does not yield any additional performance gain. This confirms that 1024 prototypes is the optimal number for this model architecture, consistent with  [2].

C.3 Ablation on Patch Drop Ratio

Table 8 analyzes the effect of different patch drop ratios during pretraining of our VideoMSN\xspace-DINO (ViT-S) model. A higher drop ratio of 70% leads to a reduced performance of 64.8%, indicating that excessive patch removal limits the model’s ability to learn meaningful representations. Conversely, a smaller drop ratio of 25% achieves 65.3%, suggesting that insufficient patch dropping provides weaker regularization. The best performance of 65.8% is obtained with a drop ratio of 50%, which provides an effective balance between representation learning and regularization.

C.4 Ablation on Focal View Configurations

Table 9 analyzes the impact of different focal view configurations during pretraining of our VideoMSN\xspace-DINO (ViT-S). Using only Fast views (3×33\times 3 super-images) achieves a Top-1 accuracy of 65.7%, while using only Slow views (2×22\times 2 super-images) yields 65.5%. A configuration without temporal masking using six 4×44\times 4 focal views obtains 65.8% accuracy. Our proposed setup, which combines Fast and Slow views (three 3×33\times 3 and three 2×22\times 2 super-image focal views), also achieves the best performance of 65.8%, indicating that integrating multiple spatial scales improves representation learning.