Zaloguj się, aby wyświetlić pełny profil użytkownika Souraj Mishra
lub
Jesteś nowym użytkownikiem LinkedIn? Dołącz teraz
Klikając Kontynuuj, aby dołączyć lub się zalogować, wyrażasz zgodę na warunki LinkedIn: Umowę użytkownika, Politykę ochrony prywatności i Zasady korzystania z plików cookie.
Zaloguj się, aby wyświetlić pełny profil użytkownika Souraj Mishra
lub
Jesteś nowym użytkownikiem LinkedIn? Dołącz teraz
Klikając Kontynuuj, aby dołączyć lub się zalogować, wyrażasz zgodę na warunki LinkedIn: Umowę użytkownika, Politykę ochrony prywatności i Zasady korzystania z plików cookie.
Kraków, Woj. Małopolskie, Polska
Zaloguj się, aby wyświetlić pełny profil użytkownika Souraj Mishra
Souraj może przedstawić Cię 9 osobom z Revolut
lub
Jesteś nowym użytkownikiem LinkedIn? Dołącz teraz
Klikając Kontynuuj, aby dołączyć lub się zalogować, wyrażasz zgodę na warunki LinkedIn: Umowę użytkownika, Politykę ochrony prywatności i Zasady korzystania z plików cookie.
16 tys. obserwujących
500+ kontaktów
Zaloguj się, aby wyświetlić pełny profil użytkownika Souraj Mishra
lub
Jesteś nowym użytkownikiem LinkedIn? Dołącz teraz
Klikając Kontynuuj, aby dołączyć lub się zalogować, wyrażasz zgodę na warunki LinkedIn: Umowę użytkownika, Politykę ochrony prywatności i Zasady korzystania z plików cookie.
Wyświetl wspólne kontakty z użytkownikiem Souraj Mishra
Souraj może przedstawić Cię 9 osobom z Revolut
lub
Jesteś nowym użytkownikiem LinkedIn? Dołącz teraz
Klikając Kontynuuj, aby dołączyć lub się zalogować, wyrażasz zgodę na warunki LinkedIn: Umowę użytkownika, Politykę ochrony prywatności i Zasady korzystania z plików cookie.
Wyświetl wspólne kontakty z użytkownikiem Souraj Mishra
lub
Jesteś nowym użytkownikiem LinkedIn? Dołącz teraz
Klikając Kontynuuj, aby dołączyć lub się zalogować, wyrażasz zgodę na warunki LinkedIn: Umowę użytkownika, Politykę ochrony prywatności i Zasady korzystania z plików cookie.
Zaloguj się, aby wyświetlić pełny profil użytkownika Souraj Mishra
lub
Jesteś nowym użytkownikiem LinkedIn? Dołącz teraz
Klikając Kontynuuj, aby dołączyć lub się zalogować, wyrażasz zgodę na warunki LinkedIn: Umowę użytkownika, Politykę ochrony prywatności i Zasady korzystania z plików cookie.
Informacje
Witamy ponownie
Klikając Kontynuuj, aby dołączyć lub się zalogować, wyrażasz zgodę na warunki LinkedIn: Umowę użytkownika, Politykę ochrony prywatności i Zasady korzystania z plików cookie.
Jesteś nowym użytkownikiem LinkedIn? Dołącz teraz
Aktywność
16 tys. obserwujących
-
Souraj Mishra udostępnił(a) toKaggle - Bronze Medal - Top 7% in Quick, Draw! Doodle Recognition Challenge by Google AI. Special thanks to Gábor Fodor, for the suggestions in the Kernels. #deeplearning #googleai #kaggle #computervision
-
-
Souraj Mishra zareagował(a) na toSouraj Mishra zareagował(a) na toCareer Update: I am delighted to announce that I will be joining the Indian Institute of Science (IISc), Bangalore as an Assistant Professor in the Department of Computer Science and Automation. After spending nine years in Europe and the USA, I am excited to return to India and contribute to the Indian academic and research ecosystem. It is an honor to join the company of distinguished researchers at IISc. I look forward to contributing to the Indian research ecosystem in AI and ML theory and joining the community of researchers dedicated to building a strong foundation for the future of AI in India. This marks another life-changing moment in my career. In the near future, I will be seeking PhD students and research assistants with strong backgrounds in statistics, computer science, mathematics, or physics. Please stay tuned for further updates.
-
Souraj Mishra polecił(a) toSouraj Mishra polecił(a) toExcited to share that our research on the "Hidden Biases of Policy Gradient for Continous action spaces" is accepted at #ICML2022 #icml. Joint work with Amrit Singh Bedi Alec Koppel Pratap Tokekar and team. Grateful for the excellent collaboration and guidance provided for my 1st publication as a Ph.D. student (1st year) at the University of Maryland – College of Computer, Mathematical, and Natural Sciences The research focuses on a very fundamental issue in the analysis of policy gradient with gaussian policies for continuous action space - "Bounded Score Functions" and we leverage heavy-tailed parametrization with momentum-based tracking to tackle the issue. Our algorithm shows improved performance in several continuous control benchmarks and we also derive the convergence analysis for our algorithm. Link: https://lnkd.in/g_QRTVW3 (Preprint) Code repository will be posted soon, stay tuned. #machinelearning #deeplearning #reinforcementlearning #robotics #statistics #probability #statisticaldataanalysis
-
Souraj Mishra polecił(a) toSouraj Mishra polecił(a) toIf you're trying to find the shortest path on a map, there are better alternatives to Dijkstra's algorithm (DA). One of them is A* search, a heuristic based extension to DA and I made this visualization to compare their performance. Let's say you're planning a road trip from Bangalore to Delhi and want to find the shortest route between these two cities. Intuitively, you know that you need to head North to get there. But DA is not guided by any sense of direction and nodes are greedily explored in order of their distance from the source node. As you can see from the visualization, this is like a wavefront spreading equally in all directions. The edges are colored in the order they were visited. A consequence of this approach is that even edges which take us in the wrong direction are explored - you can see that all of the roads in South India are explored. Can this redundancy be reduced or eliminated? In some cases, yes. For instance, if we knew the exact distance from each node to the destination, we could prioritize nodes which take us closer to the target. But this is as hard as solving the original problem. A* search circumvents this by using a heuristic - an approximation of the distance from any node to the target node. Through this approximation, exploration of edges is largely in the direction of the target node - the visualization shows a more focused search pattern resulting in fewer edges explored vs DA. In this case, the heuristic is the straight line distance which is easy to compute if you have the x,y coordinates of each node. Can any approximation be used for A* search? No, it needs to be consistent. One way of creating heuristics is to relax constraints - knock down walls, fly over obstacles (see comments for more information). Have you come across A* search in an algorithms class? Would love to know if you've used it for any coursework or projects! If you think someone in your network might like this visualization, please share it with them! Map of roads obtained from - https://lnkd.in/gjaVnejS #algorithms #cp #dsa #artificialntelligence
-
Souraj Mishra polecił(a) toSouraj Mishra polecił(a) toPlease Help me....... I'm in depression day by day because of I'm jobless......My family conditions is also not good....... Please don't ignore my post........ Even if you don't have a job for me, don't ignore this post...... Please like comment so that it reaches to your connections and someone may offer me a job. Please Help me anyone Contact: 01717265434(WhatsApp) I have experience in 18.3years Corporate & Channel Sales Development, Dealers Development.
Doświadczenie i wykształcenie
-
Revolut
**** ******** ********
-
******
******* ********* **
-
************ *****
**** *********
-
****** ********* ** *********** ******
****** ** ******* ***** *********** *** ********** ********* undefined
–
-
*******
********** ******* ******** **********
–
Wyświetl pełne doświadczenie użytkownika Souraj Mishra
Zobacz jego/jej stanowisko, okres zatrudnienia i więcej.
Witamy ponownie
Klikając Kontynuuj, aby dołączyć lub się zalogować, wyrażasz zgodę na warunki LinkedIn: Umowę użytkownika, Politykę ochrony prywatności i Zasady korzystania z plików cookie.
Jesteś nowym użytkownikiem LinkedIn? Dołącz teraz
lub
Klikając Kontynuuj, aby dołączyć lub się zalogować, wyrażasz zgodę na warunki LinkedIn: Umowę użytkownika, Politykę ochrony prywatności i Zasady korzystania z plików cookie.
Licencje i certyfikaty
Projekty
-
Machine Translation
–
As a part of NLP Nanodegree, Built a deep neural network that functions as part of an end-to-end machine translation pipeline. The completed pipeline accepts English text as input and returns the French translation.
-
Part of Speech Tagging
–
Build as part of NLP Nanodegree. Used several techniques, including table lookups, n-grams, and hidden Markov models, to tag parts of speech in sentences, and compare their performance.
Wyświetl pełny profil użytkownika Souraj Mishra.
-
Zobacz, jakich macie wspólnych znajomych
-
Zostań przedstawiony(-a)
-
Skontaktuj się bezpośrednio z użytkownikiem Souraj Mishra
Inne podobne profile
-
Mihir Shekhar
Mihir Shekhar
International Institute of Information Technology Hyderabad (IIITH)
2 tys. obserwującychNoida
Odkryj więcej publikacji
-
Rajan Chavada
Bree • 2 tys. obserwujących
🦞 There is no industry standard for how AI coding agents understand your project. Every IDE invented its own hidden file format. Engineering teams lose productivity duplicating conventions across GitHub Copilot (`.github/copilot-instructions.md`), Cursor (`.cursorrules`), Claude Code (`.claude/skills/`), Windsurf (`.windsurf/`), Codex CLI (`.agent/`), Kilo Code, Continuedev, OpenClaw (`.openclaw/`) etc. Context engineering, the discipline of architecting how information flows to your AI model, is the most important thing today for reliable agentic workflows. But I realized while building project's, that agent context is scattered across 8 different file formats with no way to keep them in sync. So I am building Rosetta, it tries to fix this. It scaffolds your project's Global Brain in .ai/master-skill.md and generates standalone IDE wrappers for every tool, no symlinks, maximum compatibility. npx rosettablueprint scaffold # Scaffold from repo analysis npx rosettablueprint migrate # OpenClaw to Cursor, Copilot to OpenClaw npx rosettablueprint ideate # Generate repo-specific skills via IDE agents npx rosettablueprint sync # Update all IDE wrappers at once It reads your entire repository, creates ideation templates, and generates domain-specific skills like Supabase guardrails that enforce RLS policies on every code change, or Jest triage agents that categorize failing tests and suggest fixes. 100% local, no APIs. As a university student, I have free token access scattered across Gemini, Perplexity, Copilot, OpenRouter, and more. Every IDE gives me different usage limits and different agent capabilities. Instead of being locked into one, I built a tool that treats all of them as interchangeable frontends to the same project brain. One source of truth. Every IDE stays in sync. Switching tools = zero context loss. GitHub (feel free to contribute!): https://lnkd.in/e9HH2VBb https://lnkd.in/ePihpSVQ
57
5 komentarzy -
Ariel Silahian
HFT Advisory • 29 tys. obserwujących
A $500M daily notional desk applied every textbook C++ optimization. Their p99 latency got worse by 23%. With their team, we traced it to three compounding root causes, all simultaneously present in the same codebase. No new hardware required. The first surprise was CMOV. The desk had aggressively replaced conditional branches with CMOV instructions (textbook advice). The problem: CMOV forces the processor to evaluate both execution paths before committing, serializing the entire dependency chain. On a predictable hot path like order routing (branch prediction accuracy above 95%), that is pure cost with no upside. An LLVM 2024 RFC confirmed this directly: converting selects back to branches reduced critical path from 21 cycles to 13.25 cycles, a 37% reduction. Bilokon, Lucuta, and Shermer documented in 2024 that branches outperform CMOV whenever prediction accuracy exceeds approximately 75%. The textbook assumes unpredictable branches. HFT order routing does not. The second root cause was Transparent Huge Pages. The desk had THP enabled system-wide, believing they were getting the benefits of pre-allocated large pages. They were not. The khugepaged compaction thread can stall for 6.4ms under no memory pressure at all. Under production load on a NUMA system, compaction events can trigger soft lockups exceeding 22 seconds. Hudson River Trading documented this in production. MongoDB, Redis, and TiDB all require THP disabled before deployment. The p50 looked clean because compaction events are infrequent. The p99 absorbed every one of them. The third was synchronization topology. The desk was routing market data from feed to strategy threads through an MPMC lock-free queue. That is a one-writer, many-reader pattern. For that specific topology, seqlocks eliminate cache coherence overhead because readers never write to shared state. Generic MPMC queues show latency scaling linearly with reader count, CAS contention on a shared cache line. Seqlocks in read-heavy configurations have been documented at two orders of magnitude faster than standard RwLock equivalents. The qualifier matters: specific to one-writer, many-reader dissemination, not a universal synchronization claim. The sketch in this post shows the before/after p99 distribution across all three changes. Three changes. No new hardware. p99 cut in half. The pattern across desks of this size: the optimization backlog targets p50. The desk bleeds capital at p99. The 2025 P99 Conference confirmed this as the primary performance frontier, with tail latency amplifying under load in ways that average-latency metrics never surface. If your optimization backlog is built from profiler averages and your p99 runs more than 3x your p50... you are governing the wrong percentile, and the spread is telling you so every session. #cpp #hft #lowlatency #electronictrading
66
4 komentarze -
Eliad Cohen
Priority Software • 2 tys. obserwujących
"I Gave My Team AI Agents. Shipping Got Slower." sounds familiar? (Post 1 of 4) I gave every engineer on my team an AI agent. Writing code got dramatically faster. Shipping a feature together got slower. The bottleneck didn't disappear. It moved. It used to be "how fast can I write this endpoint." Now it's "is the backend actually ready yet", "did the frontend guess the payload shape right", "why did QA write tests against an API that changed yesterday without telling anyone." Cross-team coordination is now the slowest part of the loop. And it's the one part nearly nobody (that I know) redesigned for a world where every team has an agent that can produce a full implementation in minutes - but will happily hallucinate an endpoint, a header, or a status code if nobody tells it otherwise. An agent doesn't say "I don't know." It guesses, confidently, and it can be halfway through an implementation before anyone notices the guess was wrong. So I've been building a methodology for this. I call it SDPD - Spec-Driven Parallel Development. The core move is simple to state: put a contract at the center of the board before anyone starts. (here is the repo with a full description https://lnkd.in/dgcEzWku) A commitment about what will be built, that every team and every agent is bound to before the first line of code. The part I didn't expect to matter as much as it does is the firewall around that contract. The contract, the specs, the test design - call this the sources layer - is human-authored and immutable to any LLM. It reads it, it never edits it. Everything an LLM is allowed to generate - summaries, overviews, orientation notes - lives in a separate wiki layer, and it's derived, regenerable, never authoritative. Here's the rule that makes the split do actual work instead of just organizing folders: the wiki is allowed to explain the contract. It is never allowed to restate it. That sounds pedantic until you watch it happen. A wiki page that says "the checkout endpoint accepts email and amount" is a second copy of the contract. Second copies drift. An agent that reads the friendly summary instead of the source will eventually build against a payload shape the summary got slightly wrong - which is exactly the failure I was trying to prevent, just relocated one layer up, wearing a nicer font. If a question is "what shape is this payload" - that's answered from the contract, directly, every time. If the contract doesn't cover it, the correct answer is "the contract doesn't cover this." Not a best guess. None of this needs to be heavy. Below roughly 15-20 sources, a vault is just the specs plus a one-page overview. The rule that matters on day one isn't a populated wiki - it's the boundary around the contract. Next post: the mistake most teams make when they split the work itself, and why the obvious split is usually the wrong one once agents are in the loop. #SDPD #SDD #SoftwareEngineering #AIAgents #EngineeringLeadership #AIDevelopment
3
-
Himanshu Jain
Citi • 4 tys. obserwujących
Why LLMs are unreliable arithmetic engines A recurring assumption I encounter among decision-makers is that one can "provide the data to the model and let it derive the measures." It is an understandable position — but it conflates two distinct capabilities: linguistic fluency and computational correctness. A large language model is a next-token predictor. Its training objective is to generate the most probable continuation of a sequence, not to execute a mathematical operation. Arithmetic, under this objective, is imitated from patterns observed in training data rather than computed. Two mechanisms explain the resulting fragility: 1. Tokenization. Numbers are not encoded as numbers. A value such as 47,829 is segmented into sub-word tokens, and that segmentation is frequently inconsistent across inputs. The model therefore lacks a stable, digit-aligned representation to operate on. 2. Absence of an algorithm. There is no internal carry-and-place procedure. The model matches against previously observed sums rather than computing them, and accuracy degrades systematically as operands grow and as reasoning chains lengthen — the regions where training coverage is sparsest. This compositional limitation is documented in the literature (Dziri et al., "Faith and Fate," 2023), and subsequent work on number representation and length generalization corroborates it. For many applications, this approximation is tolerable. In banking and financial services, it is not. A marginal error in an exposure, a balance, or a risk feature can propagate into reconciliation breaks, regulatory findings, and audit exceptions. The difficulty is compounded by calibration: an incorrect result is presented with the same confidence as a correct one. The appropriate response is not to abandon LLMs but to bound their role. Material quantities should be computed in deterministic code (SQL, pandas, Spark); calculation should be delegated to tool-calling or code execution; and the model should be reserved for interpreting intent, routing tasks, and explaining outputs — not for deriving metrics internally. LLMs are exceptional language models and unreliable arithmetic engines. Sound system design begins with that distinction. I would be interested to hear how others are formalizing this boundary between model-mediated reasoning and deterministic computation in their pipelines. #AIEngineering #DataScience #BankingAI #LLMs #MLOps Jagadeesh Jagarlamudi
17
1 komentarz -
Andrii Shchur
Seeking Alpha • 3 tys. obserwujących
From folding towels to reasoning robots 🧠🤖 Problem For decades, robotics has been trapped in Moravec’s Paradox: Computers handle high-level reasoning easily, but struggle with basic physical skills. A robot can solve equations — yet folding a shirt breaks it. Why? Because we tried to control reality with explicit logic. • Millions of lines of code for every joint angle • Hand-written collision rules • Hard-coded friction and geometry The result: precision without robustness. Change the lighting. Change the friction. Change reality by a millimeter — and the robot fails. That’s why industrial robots lived in cages. They weren’t intelligent agents. They were replaying scripts. ------------------------------------------------------------- 🧠 Brain — Alpamayo Alpamayo changes the paradigm. It’s a Vision-Language-Action (VLA) model where actions are tokens, just like words. Turning a wrist by 5 degrees = choosing the next word in a sentence. The real breakthrough is reasoning before action. Alpamayo brings System-2 thinking into robotics: • Perception — fast • Reasoning — deliberate • Action — fast Before moving, the robot simulates outcomes, predicts consequences, and can explain why it acted — a critical step for safety. ------------------------------------------------------------- 🧪 Simulator — AlpaSim Training robots in the real world is slow, expensive, and risky. AlpaSim moves learning into simulation. Developers connect Alpamayo to a digital robot and let it live thousands of parallel lifetimes — learning to grasp, move, and manipulate without touching real hardware. The feedback loop becomes: • fast • scalable • reproducible ------------------------------------------------------------- 🌍 World Model — Cosmos Cosmos is a world foundation model, not a physics engine. It learns physics from video — the way humans do. Trained on millions of hours of footage, it understands: • gravity • materials • fabric • fluids • lighting Not as equations — but as probabilities. By generating thousands of randomized realities (lighting, textures, friction, noise), Cosmos finally helps close the sim-to-real gap. ------------------------------------------------------------- ⚙️ Hardware — Jetson Reasoning only matters if it’s fast. In robotics, latency = safety. That’s why edge platforms like NVIDIA Jetson matter — enabling real-time reasoning under ~100 ms, right next to the sensors. ------------------------------------------------------------- Result The era of “dumb” industrial robots is ending. What comes next won’t be scripted machines — but agents that predict consequences before they move. This feels like a Stable Diffusion moment for robotics. And it’s just beginning. #Robotics #ArtificialIntelligence #DeepLearning #IndustrialRobotics #EdgeAI #NVIDIA #Jetson
1
-
Nickolaus Lachman
Valley Bank • 1 tys. obserwujących
Not every improvement in AI comes from buying more GPUs. DeepSeek just put out a paper that’s a good reminder of that. At a high level, most modern models are just deep stacks of layers passing information forward. As those stacks get deeper, information degrades. We solved part of that problem years ago with residual connections. Basically shortcuts so important signals don’t disappear as models scale. That idea is now standard across models like GPT and AlphaFold. But a single shortcut can become a bottleneck. When too much information is forced through one path, different ideas start interfering with each other. Hyperconnections tried to fix this by letting models run multiple parallel “thinking streams.” More flexibility, better reasoning... but also instability and rapidly rising training costs. DeepSeek’s approach, which they call manifold-constrained hyperconnections, keeps the parallel streams but adds guardrails so they don’t turn into chaos. The highway analogy works pretty well: Multiple lanes, but with dividers, speed limits, and rules of the road. What stood out to me is that they get consistent performance gains and stable training without dramatically increasing compute. That’s a strong signal that architecture still matters. That we’re not done finding efficiency gains just by designing models better. DeepSeek also tends to publish research like this right before shipping something real. With Chinese New Year coming up, I wouldn’t be surprised to see a new model soon. Feels like a reminder that the next step forward in AI may be quieter than the last one.. less brute force, more design. Link to the paper here: https://lnkd.in/epSk5zSH
26
2 komentarze -
Rangel Isaías Alvarado Walles
Grupo LAFISE • 7 tys. obserwujących
Push, Press, Slide: Mode-Aware Planar Contact Manipulation via Reduced-Order Models Arxiv: https://lnkd.in/eRkT-Wux Project: [Link not provided] 🔁 At a Glance 💡 Goal: Simplify non-prehensile planar manipulation by abstracting complex mechanics into intuitive modes. ⚙️ Approach: Mode-aware reduced-order models: Map contact topologies to simple non-holonomic templates Contact topology selection: Choose the model based on task feasibility and environment Algebraic force allocation: Compute contact forces instantly without optimization Kinematic feasibility: Integrate manipulator constraints for long-horizon planning 📈 Impact (Key Results) 🧪 Faster Control: Optimization-free, closed-form motion planning Successful simulation of diverse single- and bimanual tasks 🔄 Robust Manipulation: Consistent force distribution within friction constraints Stable control using body-fixed stabilization points 🤖 Enhanced Mobility: Unified models bridging car-like, unicycle, and differential drive behaviors 🔬 Experiments 🧪 Benchmarks: Drake hydroelastic contact model, simulated in diverse task settings 🎯 Tasks: pushing, press-and-slide, pivoting, object fitting 🦾 Setup: Two KUKA iiwa7 arms, Franka Panda robot, various object sizes 📐 Inputs: Object properties, contact modes → object trajectories and force commands 🛠 How to Implement 1️⃣ Identify task contact mode based on environment and goals 2️⃣ Assume corresponding reduced-order model for motion planning 3️⃣ Generate object and arm trajectories using the mode-specific model 4️⃣ Use algebraic allocator to distribute forces within friction limits 5️⃣ Regulate normal forces with PI controllers, execute via hybrid force-impedance control 📦 Deployment Benefits ✅ Fast, optimization-free planning enabling real-time control ✅ Unified approach for single- and bimanual manipulation ✅ Robustness against unmodeled friction variations ✅ Compatibility with diverse robotic platforms 📣 Takeaway This mode-aware reduced-order framework unlocks real-time, scalable non-prehensile manipulation. By abstracting complex mechanics into simple models, robots can perform agile, stable contact tasks. The approach bridges the gap between theoretical mechanics and practical control, paving the way for advanced robotic manipulation in unstructured environments. Follow me to know more about AI, ML and Robotics!
16
-
Alexander Berkovich
Multiple Startups • 9 tys. obserwujących
🚨On why 𝐋𝐋𝐌 𝐢𝐧𝐟𝐞𝐫𝐞𝐧𝐜𝐞 is unique with a few tips: ① Due to variable length input and output consider 𝐂𝐨𝐧𝐭𝐢𝐧𝐮𝐨𝐮𝐬 𝐁𝐚𝐭𝐜𝐡𝐢𝐧𝐠 🌀 ② 𝐒𝐩𝐥𝐢𝐭 the 𝐏𝐫𝐞𝐟𝐢𝐥𝐥 𝐬𝐭𝐚𝐠𝐞, which is computer heavy, from 𝐃𝐞𝐜𝐨𝐝𝐞 𝐬𝐭𝐚𝐠𝐞, which is memory heavy ⚡ ③ To avoid recalculation, and share what’s possible, apply 𝐬𝐦𝐚𝐫𝐭 𝐆𝐏𝐔 𝐦𝐞𝐦𝐨𝐫𝐲 𝐦𝐚𝐧𝐚𝐠𝐞𝐦𝐞𝐧𝐭 💾 ④ Reuse prefix data with 𝐏𝐫𝐞𝐟𝐢𝐱 𝐀𝐰𝐚𝐫𝐞 𝐑𝐨𝐮𝐭𝐢𝐧𝐠 to different replicas of the model 🔀 ⑤ Apply a 𝐦𝐢𝐱𝐭𝐮𝐫𝐞 𝐨𝐟 𝐞𝐱𝐩𝐞𝐫𝐭 𝐦𝐨𝐝𝐞𝐥𝐬 & route query to most relevant one 🎯 [Thanks to Robert Nishihara, Cofounder of Anyscale] Got great tips? Comment below 👇
1
-
Yuvraj Singh Bhadoria
Bank of America • 8 tys. obserwujących
SpecActor, Explained — Why Rollout, Not Math, Is the Real Bottleneck in LLM Post-training A paper from ByteDance Seed and SJTU that's the sharpest systems argument I've read this year: you don't lose RL post-training time to compute — you lose it to waiting. The core insight: rollout eats 75–80% of post-training time, and on real traces ~50% of that GPU time is idle — every worker stalled on the slowest one still solving its task. Generation is sequential and memory-bound, so you can't buy your way out: 2× the GPUs nets 1.2–1.3×. The fix: take speculative decoding — the inference trick you already know — and retrofit it to training. Fast draft path, parallel verification, lossless by construction. But naive speculation fails at training batch sizes: at per-worker batch 128, verification hits the compute limit and the gain goes to zero or negative. SpecActor fixes that with two moves. Move 1 — Decoupled speculation. Put drafter and verifier on separate GPUs. The drafter runs ahead without waiting for verification, so the compute-hungry verifier gets real GPU time instead of idling behind a tiny draft model. Aggression is capped by a draft window (w): at most 2w−1 tokens wasted on a rejected guess. Move 2 — Fastest-of-N speculation. Stop committing to one draft method per request. Build a draft ladder offline (which drafter wins at which acceptance rate), pick the best guess by the batch's average acceptance rate — statistically stable even when individual requests vary — then, as workers finish, launch additional drafters (0.5B, 1.5B, n-gram…) at the long-tailed stragglers on the freed GPUs: 1. Pass the request through every live drafter in parallel 2. Accept the first draft method that emits an accepted EOS 3. Remove the request everywhere — the batch finishes when the fastest finishes Benefits: - 2.0–2.4× mean rollout speedup, up to 2.7× - 1.1–2.6× faster than vanilla speculative rollout - 1.4–2.3× faster end-to-end post-training - Lossless — exact token matching, zero accuracy change - Algorithm-agnostic: works on GRPO, DAPO, and PPO; dense and MoE - Drop-in replacement of the inference component in veRL Key takeaways for anyone scaling post-training: 1. Buy hardware last — sequential waits don't yield to silicon; 2× GPUs bought 1.2× 2. Decouple the dependency you took as law: draft-then-verify coupling was starving the verifier of GPU time 3. Don't pick one drafter — when the batch shrinks, bet on the tail with parallel drafters 4. Measure batch-completion time, not token throughput — those optimize different systems Source: Fast LLM Post-training via Decoupled and Fastest-of-N Speculation, arXiv:2511.16193
17
3 komentarze -
Mayank A.
Confidential • 184 tys. obserwujących
DeepSeek discovered a new 'Scaling Knob' for LLMs. (But the implementation is the real story) Everyone knows "Depth" and "Width." DeepSeek just unlocked "Residual Width", adding multiple parallel streams of information flow between layers. Normally, this is too slow. Mixing information across 4 parallel streams requires massive memory bandwidth. So they re-engineered the GPU pipeline. The paper is a masterclass in CUDA optimization: ➡️ Kernel Fusion: They fused the Sinkhorn projection and the Matrix Multiplication into a single kernel to minimize HBM reads. ➡️ DualPipe: They overlap the communication of these new "Hyper-Connections" with the main computation stream. ➡️ Selective Recomputation: They only store specific activations, trading compute for VRAM. The Result => They achieved a 4x expansion in information bandwidth for only a 6.7% increase in training time. This is how you beat NVIDIA at their own game.
42
7 komentarzy