Sign in to view Divay’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Sign in to view Divay’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Bengaluru, Karnataka, India
Sign in to view Divay’s full profile
Divay can introduce you to 10+ people at Meesho
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
2K followers
500+ connections
Sign in to view Divay’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
View mutual connections with Divay
Divay can introduce you to 10+ people at Meesho
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
View mutual connections with Divay
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Sign in to view Divay’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
About
Welcome back
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
New to LinkedIn? Join now
Activity
2K followers
-
Divay Jindal posted thisExcited to share that our paper “TrendPulse: A Simple yet Efficient Framework for Capturing Viral E-Commerce Spikes via LLM-Driven Contextualization” has been accepted at ACL 2026 (Industry Track) TrendPulse addresses the challenge of capturing short-lived demand spikes by detecting regional search momentum, transforming it into semantic trends using LLMs, and powering personalized recommendations via attention framework. This would not have been possible without an incredible team. Thanks to Devashish Gupta Arin Jain Bhavuk Singhal Also thankful for the support and guidance from Vinit Rongata Ravindra Yadav Debdoot Mukherjee Looking forward to presenting at ACL! #ACL2026 #MachineLearning #NLP #RecommenderSystems #LLM
-
Divay Jindal reposted thisDivay Jindal reposted thisHow do you personalize for users you know nothing about? At Meesho, we tackled the cold start problem head-on: blending demographics and early-session signals to deliver relevant recommendations from the very first scroll. Curious how we approached the cold start problem? 👉 https://lnkd.in/gfmg3b_b #Personalization #Ecommerce #MeeshoTech #MLForBharat Vinit Rongata Debdoot Mukherjee Ravindra Yadav Divay Jindal Devashish GuptaPersonalization from Day One: Solving the Cold Start Problem at MeeshoPersonalization from Day One: Solving the Cold Start Problem at Meesho
-
Divay Jindal shared thisGreat answerWhat skills do I need to be a data scientist at Google or Facebook?What skills do I need to be a data scientist at Google or Facebook?
-
Divay Jindal shared thisAwesome glossary. Have a look
-
Divay Jindal liked thisDivay Jindal liked thisTokens are all you need - even for recommendation systems. A new paper from researchers at YouTube and Google DeepMind tackles one of the biggest scaling problems in industrial recommenders: the "Memory Wall." Modern ranking and retrieval models depend on massive, dense floating-point embedding tables, and streaming those vectors during training and serving creates severe I/O and memory bandwidth bottlenecks - especially with user histories now reaching tens of thousands of items. Their answer: Dual-purpose Semantic IDs. How it works under the hood: Each item's high-dimensional content embedding (from a pre-trained multimodal model) is compressed via hierarchical residual quantization (RQ-VAE) into a short sequence of discrete tokens drawn from a shared codebook - achieving roughly 50–100x compression. These tokens then do double duty inside the model: 1. Collaborative Identity: The tokens act as categorical features mapped to learnable embedding tables, using unigram, overlapping bigram, nested n-gram, or SentencePiece representations. Since semantically similar items share token prefixes, the model generalizes naturally to cold-start and long-tail content. 2. Semantic Decoding (SiDec): Instead of logging and joining dense vectors, a static codebook lookup plus a lightweight trainable decoder (an MLP or shallow Transformer) reconstructs an approximation of the original content embedding on-the-fly, inside the model graph. This replaces the costly "join" of dense features to user history - which can hit around 200 KB per training example - with a simple integer token lookup. The results are compelling. Offline, raw dense embeddings degraded training throughput by 28%; SiDec recovered most of that speed while actually improving retrieval hit rate, and pairing it with Transformer scaling delivered the best quality at 20% faster training than dense ingestion. Online, the framework shipped in production ranking and retrieval systems with statistically significant engagement gains, disproportionately helping new users and long-tail content. The broader takeaway: almost any dense continuous feature can be tokenized, letting recommenders trade disk-bound I/O for compute-bound reconstruction - the same scaling regime LLMs enjoy. This post is in free😅 support of friends at Source Strong AI (Do check them out)
-
Divay Jindal liked thisDivay Jindal liked thisA Google Distinguished Engineer just open-sourced how Google builds production agents. 291 pages. 22 chapters. All code public. This is the playbook Google engineers use internally. "Building AI Agents: From Design Patterns to Production" Written by Antonio Gulli (Sr Director, CTO Office) and Anant Nawalgaria (Senior Staff ML Architect). The core loop: Perceive → Plan → Act → Observe. 4 patterns implemented end-to-end: → ReAct → Chain-of-Thought → Reflection → Plan-and-Execute Each chapter builds Atlas, a research assistant that evolves from a 50-line script into a self-correcting multi-agent system. 100% of author royalties go to Save the Children. Learn Agent Engineering → Support Kids → https://lnkd.in/eMQEVVEk
-
Divay Jindal liked thisDivay Jindal liked thisI’m excited to share that our work has been accepted at NeurIPS 2026! 🎉 Incredibly happy to have our research featured at one of the world’s premier AI/ML conferences! We address: How can a robot plan its next move when its sensors are noisy and it can only observe part of its state? 🤖 We introduce K-PWM (Koopman Probabilistic World Model), a framework for learning structured probabilistic world models from noisy data under partial observation and using them for uncertainty-aware predictive control. The goal: predict how a system will evolve, quantify uncertainty in those predictions, and use that uncertainty to guide control decisions. Grateful to our collaborators at Cornell University and West Virginia University. Preprint and code coming soon! Looking forward to connecting with others working on world models, robotics, and control. The Ohio State University Department of Mechanical and Aerospace Engineering #NeurIPS2026 #WorldModels #Robotics #AI #ML #PhysicalAI #MPC #MPPI
-
Divay Jindal liked thisDivay Jindal liked thisLast week at the Indian Institute of Management Bangalore campus, I had the privilege of interacting with senior transitioning officers from the Indian Army, Navy, and Air Force belonging to the IIM Bangalore Business Management Programme for Defence Officers (BMPDO) cohort. To be invited to this elite gathering, engage in strategic dialogues on corporate transition, and be presented with a special memento by officers of such distinguished caliber was a deeply humbling experience. The military exemplifies strategic discipline, sovereign leadership, and crisis management—values that resonate profoundly with the core of Organizational Development (OD) and Human Capability. My sincere gratitude to IIM Bangalore and the senior defense officials for this rare gesture of recognition. Moments like these serve as a powerful reminder of the impact we create when corporate strategy aligns with purposeful, sovereign leadership. Honored to receive the memento from Maj Gen Harish Bhutani (Retd) and Meeting some extraordinary officers: Sunil Raj Supandeep Singh Col Ravjeet S.. Vivek Murugesh Brig Vikas Sharma (Retd) Sharma Rati Bhargava Suhas Malannavar It was also a privilege to meet Rajesh Pal. Sincere gratitude to Rafikul Raheman for the invite. Thank you, Srinivas Raghavendrarao Thondamuthur for keeping me motivated! #ExecutiveLeadership #OrganizationalDevelopment #BMPDO #LeadershipDevelopment #Learning&Development IIMBangalore #DefenseTransition #CorporateVisionary
-
Divay Jindal liked thisDivay Jindal liked thisGoogle DeepMind + Massachusetts Institute of Technology + Stanford University are letting AI regenerate software from design docs. 🤯🤖 What if code stopped being the source of truth? Their new system, SMART, treats code as a disposable build artifact. 🔄 Instead, humans maintain natural-language design docs and AI agents regenerate the entire library from those docs. 🤖 The workflow: → 📝 Design docs become a dependency DAG → 🤖 AI agents implement each document → ♻️ The entire library is regenerated → 🧪 Generated code is validated against reference models → 🚀 Repeat And the main branch contains almost no code. The clever part? The docs include worked examples showing step-by-step execution, intermediate values, shapes, and exact expected cost expressions. So documentation becomes: 📚 Specification + 💡 examples + 🧪 tests. SMART already spans ~50 design docs and ~9,000 lines of specification covering TPU topology, collective communication, schedulers, numerics, and frontier models. A complete regeneration takes 1.5–3 hours and costs roughly $100. The bigger idea is fascinating: Humans maintain intent. AI maintains implementation. 🚀
-
Divay Jindal liked thisDivay Jindal liked thisWhat if a recommender could spell out the item it wants to recommend? I’ve been reading about semantic IDs. LLMs process text as sequences of tokens. Semantic IDs let us represent items as short token sequences too, with similar items sharing parts of their codes. We can then train a transformer or fine-tune an LLM to recommend items by generating those tokens, one at a time. One result caught my attention: in Spotify’s GLIDE experiments (D’Amico et al., 2026), a simpler clustering method beat RQ-VAE, the neural approach used to build these IDs in Google’s TIGER (Rajput et al. ,2023). I wonder whether good starting embeddings leave less work for the extra neural network to do. That’s a hypothesis, rather than something the experiment establishes. This week on my blog, I explain briefly how this works, and what happens when the model spells out an ID that doesn’t correspond to a real item: https://lnkd.in/eUbusmWf. Papers: 1. TIGER: https://lnkd.in/eMZ3Egcs 2. GLIDE: https://lnkd.in/e3jatsjp Source of image: Rajput et al. (2023), Figure 3. Disclaimer: personal notes on public research. Opinions are my own.
-
Divay Jindal liked thisDivay Jindal liked thisI thought I understood GPU utilization until I started working on this handbook. I kept collapsing three different things into one mental picture => the work a kernel launches, the work that is actually resident on the GPU, and the work that is ready to issue right now. They are not the same thing. Take a kernel with 240 blocks and 256 threads per block. That is 61,440 logical threads, or 1,920 NVIDIA warps. But those warps are not all physically resident at once. Blocks are admitted onto SMs as registers, shared memory, thread limits, warp limits and block limits allow. Even after a warp becomes resident, it may still be waiting on memory, an arithmetic dependency or synchronization. The scheduler can only choose from warps that are actually eligible to issue. That changed how I read performance numbers too. Occupancy is a residency number => resident warps compared with the architectural maximum. GPU utilization measures something different. In NVML it is basically the fraction of the sampling window during which at least one kernel was executing. So a GPU can show 100% utilization without Tensor Cores being saturated, without HBM bandwidth being saturated, and while one subsystem is the bottleneck and large parts of the GPU remain underused. Memory has the same kind of traps. Registers, shared memory, L1, L2 and HBM are not just increasingly slower boxes. They have different scope, capacity, management and access behavior. A model fitting in HBM tells you that it fits. It tells you almost nothing about how many bytes move, whether accesses coalesce well, how much reuse you get, or whether memory is what is holding the kernel back. FlashAttention is a nice example => the dense attention math stays the same, but the execution schedule is reorganized to reduce traffic between HBM and on-chip storage. Tensor Cores were another thing I had mentally oversimplified. They are specialized matrix multiply-accumulate hardware and they matter enormously for AI, but they do not “run the model.” Reductions, indexing, elementwise work, synchronization, memory movement, launches and plenty of other instructions still go through other parts of the GPU. After a while I stopped asking "why isn’t the GPU at 100%?" and started asking a better question => what is actually limiting useful progress right now? That ended up becoming a 42-page handbook. SMs, warps, schedulers, occupancy, latency hiding, registers, shared memory, caches, HBM, coalescing, Tensor Cores, GEMM mapping, precision, utilization and the performance numbers that are very easy to misread. Do read.
-
Divay Jindal liked thisDivay Jindal liked thisBanger paper from Google DeepMind and colleagues. (bookmark it) A model reads its entire KV cache on every generated token, even though it ends up attending to a tiny slice of it. In other words, if you ask about one detail from a 1M-token conversation the global attention layers re-read all of it, per token. The usual fix is to guess the relevant tokens first with cheap proxy scores, which still costs O(N) every step. Declarative Attention asks the model instead. The model declares where it needs to look, inside its own chain-of-thought. In this way, generation splits into three modes: global reads the full context, focus reads one specific region, and local reads only recent output. The inference engine parses those declarations the same way it parses tool calls and skips most of the cache read. On zero-shot on off-the-shelf weights across 15 long-context tasks, attended tokens during decoding drop 52.0% on Gemma-4-31B and 31.1% on Qwen-3.6-27B.
-
Divay Jindal liked thisDivay Jindal liked thisYour recommender may be deleting the right answer before it ever gets ranked. A new study from researchers at the University of Illinois Urbana-Champaign takes apart Semantic IDs (SIDs), the backbone of modern generative recommendation, and finds a structural conflict hiding in plain sight. Quick refresher on how these systems work under the hood: an encoder (Sentence-T5, BGE, Qwen) turns item content into a continuous embedding, a quantizer (RQ-VAE, RQ-KMeans, OPQ/PQ, and others) compresses it into a short sequence of discrete tokens, and an autoregressive decoder like TIGER generates that token sequence with constrained beam search. The catch: every generated token is also a filtering step. Discard a prefix, and every item under it is permanently gone, before final ranking even begins. What the paper found across three Amazon domains and eight SID constructors: - SID neighborhoods recover only 32.2% of the encoder's ten nearest neighbors on average. Coarse category structure survives; fine local structure does not. - Simply reordering an item's description fields (same words, same facts) changes 38.4% of exact SIDs on average, even though the encoder still retrieves the correct item first in 99.57% of controlled cases. Exact SIDs are not determined by meaning alone. - The consequence at decoding time: after the final semantic token, TIGER retains only 35.4% of Office targets and 24.4% of Scientific targets that an independent item scorer had ranked in its top 10. The model's own tokenization prunes items a separate model considers highly relevant. Their fix, Item-Supported Decoding (ISD), is refreshingly pragmatic. No retraining, no new parameters. Before generation, build one ranked item set from the user's history using any ranker: simple training statistics (item transitions plus co-consumption), SASRec, or UniSRec. During beam search, each candidate prefix gets two ranks: one from the decoder's sequence probability and one from the best-ranked item still living under that prefix. Reciprocal rank fusion merges them, so a prefix can survive either because the decoder likes it or because it shelters a highly relevant item. After generation, the same ranking reorders the output. Asymptotic decoding complexity stays unchanged. Result: NDCG@10 improves in every evaluated setting across TIGER and LIGER, with gains up to 31.2%. The takeaway is bigger than the method: SIDs are good at organizing items, but their fine token boundaries should not be the sole authority on which items survive generation. Identification and selection are different jobs. This post is in free😅 support of friends at Source Strong AI (Do check them out)
Experience & Education
-
Meesho
********* **** *********
-
**********
**** ********* **
-
*** ****** ******
******* ******** **********
-
******** ********* ** ********** *******
********** ****** ******** ******* undefined
-
-
*** ****** ******
****** ********* ********* undefined *****
-
View Divay’s full experience
See their title, tenure and more.
Welcome back
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
New to LinkedIn? Join now
or
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Publications
-
Basic Computer Skills
-
It is a 250 pages book covering all the basics skills needed to be taught to people such as Microsoft Office , basic computer architecture , introduction data entry tools and many more useful tools aiming at the skill development of the people in need so that could generate employment for themselves . This is approved under "Kaushal Vikas Yojna".
Other authors
Courses
-
Convolutional Neural Networks for Visual Recognition
CS231n
-
Linear Algebra
18.06 Gilbert Strang
-
Machine Learning
Abu Mostafa
-
Machine Learning
Andrew Ng
-
Probability And Statistics
STAT 110
-
RL Course by David Silver
Reinforcement learning
-
Statistical Rethinking
Bayesian Machine Learning
Organizations
-
GYANSAGAR
Co-Founder
- PresentGyansagar is a Non Profit Organisation which aims at the skill development of the needy people. We teach English,Mathematics and Science to about 200 students living in the villages near to the campus of NIT Silchar. We conduct computer classes aiming at developing basic computer skill amongst the students .
View Divay’s full profile
-
See who you know in common
-
Get introduced
-
Contact Divay directly
Other similar profiles
-
Mihir Shekhar
Mihir Shekhar
International Institute of Information Technology Hyderabad (IIITH)
2K followersNoida
Explore more posts
-
Naveen Manwani
AIMonk Labs Private Ltd • 7K followers
🚨 NeurIPS 2025 Spotlight Paper Alert 🚨 ➡️Paper Title: STITCH-OPE: Trajectory Stitching with Guided Diffusion for Off-Policy Evaluation 🌟Few pointers from the paper 🎯Off-policy evaluation (OPE) estimates the performance of a target policy using offline data collected from a behavior policy, and is crucial in domains such as robotics or healthcare where direct interaction with the environment is costly or unsafe. 🎯Existing OPE methods are ineffective for high-dimensional, long-horizon problems, due to exponential blow-ups in variance from importance weighting or compounding errors from learned dynamics models. 🎯To address these challenges, authors of this paper proposed “STITCH-OPE”, a model-based generative framework that leverages denoising diffusion for long-horizon OPE in high-dimensional state and action spaces. 🎯Starting with a diffusion model pre-trained on the behavior data, STITCH-OPE generates synthetic trajectories from the target policy by guiding the denoising process using the score function of the target policy. 🎯STITCH-OPE proposes two technical innovations that make it advantageous for OPE: (1) prevents over-regularization by subtracting the score of the behavior policy during guidance, and (2) generates long-horizon trajectories by stitching partial trajectories together end-to-end. 🎯They provide a theoretical guarantee that under mild assumptions, these modifications result in an exponential reduction in variance versus long-horizon trajectory diffusion. 🎯Experiments on the D4RL and OpenAI Gym benchmarks show substantial improvement in mean squared error, correlation, and regret metrics compared to state-of-the-art OPE methods. 🏢Organization: University of Toronto, University of Toronto Robotics Institute, Toyota Research Institute, Vector Institute 🧙Paper Authors: Hossein Goli, Michael Gimelfarb, Nathan Samuel de Lara, Haruki Nishimura, Masha Itkina, Florian Shkurti 📝 Read the Full Paper here: https://lnkd.in/gz2-DSQn 🗂️ Project Page: https://lnkd.in/gV5vdcF4 🧑💻 Code: https://lnkd.in/g9AvhmpE 🎥 Be sure to watch the attached Technical Summary Video - Sound on 🔊🔊 Find this Valuable 💎 ? ♻️REPOST and teach your network something new Follow me 👣, Naveen Manwani, for the latest updates on Tech and AI-related news, insightful research papers, and exciting announcements. #NeurIPS2025
11
-
Youssef Hosni
Solita • 118K followers
Build Self-Learning Agents — No Fine-Tuning Required Fine-tuning isn’t always the answer. In our latest newsletter, Hari Ohm Prasath Rajagopal breaks down how you can build self-improving LLM agents without touching model training — using two simple but powerful patterns: 🔹 Dynamic Tooling: Let agents create and load new tools on the fly when existing ones aren’t enough. 🔹 Dynamic System Prompts: Allow prompts to evolve automatically based on real user or reviewer feedback. The key idea: Instead of expensive fine-tuning and slow human feedback loops, let agents learn from usage by evolving their tools and instructions in real time. He also walks through how this works in practice using Strands (AWS’s agentic framework), with minimal code and tight AWS integration. If you’re a software engineer building agentic systems and want: - Better domain accuracy - Faster iteration - Less ML complexity This one’s for you. 👉 Read the full article here: https://lnkd.in/dW3jtJ2W
16
1 Comment -
Nishantha Ruwan
IWROBOTX Software Inc. • 2K followers
REFRAG: Rethinking RAG based Decoding The authors argue that in retrieval-augmented generation (RAG) applications — where a large language model (LLM) ingests many retrieved passages along with user queries — the standard decoding process wastes large amounts of computation and memory. This inefficiency arises because many retrieved passages are irrelevant to a given prompt and, due to low semantic similarity and deduplication, yield sparse cross-attention patterns. Nonetheless, existing systems allocate full token-level key/value cache and perform attention over the entire concatenated context, leading to high latency and memory load, limiting throughput and practical deployment when context is long. To address this, the paper introduces REFRAG — a decoding framework that compresses retrieved passages into compact chunk embeddings, feeds those embeddings rather than full token sequences into the decoder, and then selectively expands only those chunks deemed important via a small, lightweight policy (learned with reinforcement learning). This “compress-sense-expand” approach shortens the decoder’s input sequence roughly by the chunk size factor, drastically reducing memory and attention overhead, without modifying the underlying LLM architecture. Experiments show REFRAG achieves up to ~30.85× speed-up in time-to-first-token (TTFT) — a substantial improvement over prior methods — while maintaining comparable or better accuracy (perplexity, summarization quality) across a variety of long-context tasks. https://lnkd.in/gyaZsAft
-
Nilesh Barla
Adaline • 2K followers
Wrote a blog explaining "LoRA Fine-tuning Efficiency Under Different Loss Functions." LoRA made fine-tuning cheap, but not always efficient. This post explores how loss functions like Cross-Entropy, Label Smoothing, and Focal Loss change convergence, stability, and generalization in lightweight LLMs. Read here: https://lnkd.in/gh_vm2kG
2
-
Nagesh Singh Chauhan
PRISM • 27K followers
Quantization in Large Language Models (LLMs) As LLMs scale, intelligence is no longer the bottleneck — memory, latency, and cost are. Quantization turns out to be one of the most powerful (and misunderstood) techniques enabling real-world deployment of large models. In this blog, I break down: • Why LLMs are surprisingly quantization-friendly • INT8 vs INT4 vs NF4 — what actually works in practice • How modern architectures absorb low-bit noise • Where quantization fails and how to design around it Quantization is no longer an optimization step — it’s foundational infrastructure for production LLM systems. Link: https://lnkd.in/gqGje75G #LLMs #AIEngineering #Quantization #MachineLearning #GenAI #MLOps
9
2 Comments -
Aryan Srivastava
Digital Gateway • 553 followers
Learn in 60 sec.. Sampling Parameter - Setting that controls the randomness and creativity in LLMs output. Filter and truncate the unlikely tokens before selection Temperature is a parameter used in language models that controls the randomness of the generated text. It alteers the shape of probability across all tokens. High value like 1 means createive and low value like 0.2 means deterministic Top-K sampling - Method use in LLM to get K highest probability token, preventing the model from ever choosing obscure or out-of-context words Vector DBs Designed to store high dimensional vectors. Great at managing unstrucutured data by enabling similarity search. Vecctor dbs use inteding techniques like Appproximate NEarest Neighbour(ANN) to handle large dataset. Usacase Simarity Search Classifications recommendation Engine Anamoly detection - check if data points are simliar to existing More - https://linktr.ee/aryansri #AI #MachineLearning #LLM #VectorDatabases
5
-
Robin Reni
McDonald's Global Office in… • 5K followers
Deepseek Moment for India When the DeepSeek models were launched, the global AI community took notice. They quickly emerged as a strong contender alongside models from OpenAI and Anthropic, showing that new players could challenge the status quo in foundation models. Now, it feels like India is having a similar moment. The Sarvam team has introduced Indic open LLMs - Sarvam 30B and Sarvam 105B, marking an exciting milestone for the Indian AI ecosystem. Building powerful open models tailored for Indic languages is a huge step toward making AI more inclusive and locally relevant. Kudos to the team for pushing the boundaries and contributing to the open AI movement. They have also launched an arena platform where users can test and compare models, which is a great way for the community to explore the capabilities of these models firsthand. Really excited to experiment with these models and see how they can be applied in real-world use cases. To Try: https://lnkd.in/gbYBpsbW To Learn More about Sarvam Models: https://lnkd.in/gBHdsDWc #India #AI #LLMs #GenAI
11
-
Rajeev Singh Sisodiya
Teleperformance India • 991 followers
Projecr Titel: Strassen’s Algorithm for matrix multiplication Contributor: Rajeev Singh Sisodiya, ORCID: 0009-0003-6585-4744 Strassen’s algorithm is a divide-and-conquer algorithm for matrix multiplication that reduces the computational complexity compared to the standard (naïve) method. Instead of performing 8 multiplications when dividing matrices into blocks, Strassen reduces it to 7 multiplications by increasing the number of additions/subtractions. Strassen’s algorithm is much deeper than “just a faster multiplication trick.” It is fundamentally about tensor rank reduction, which connects it directly to fast linear algebra and even quantum algorithms. The standard algorithm corresponds to tensor rank = 8, because it uses 8 scalar multiplications. After Strassen, research focused on minimizing the tensor rank of matrix multiplication. Key milestones: 1.Strassen (1969) 2.Coppersmith–Winograd (1987) 3.Alman–Vassilevska Williams (2021) The relationship between HHL, block-encoding, and QSVT (Quantum Singular Value Transformation) is essentially the evolution of quantum linear-algebra algorithms from a specialized construction to a general operator-transformation framework. You can think of it historically and structurally: HHL (2009) → first quantum linear-system solver Block-encoding (2015–2019) → unified operator representation QSVT (2018–2019) → universal matrix-function compiler https://lnkd.in/gHqfydKT
1
Explore collaborative articles
We’re unlocking community knowledge in a new way. Experts add insights directly into each article, started with the help of AI.
Explore More