Identifiez-vous pour voir le profil complet de Neil
ou
Nouveau sur LinkedIn ? Inscrivez-vous maintenant
En cliquant sur Continuer pour vous inscrire ou vous identifier, vous acceptez les Conditions d’utilisation, la Politique de confidentialité et la Politique relative aux cookies de LinkedIn.
Identifiez-vous pour voir le profil complet de Neil
ou
Nouveau sur LinkedIn ? Inscrivez-vous maintenant
En cliquant sur Continuer pour vous inscrire ou vous identifier, vous acceptez les Conditions d’utilisation, la Politique de confidentialité et la Politique relative aux cookies de LinkedIn.
Paris, Île-de-France, France
Identifiez-vous pour voir le profil complet de Neil
Neil peut vous mettre en relation avec plus de 10 personnes chez Gradium
ou
Nouveau sur LinkedIn ? Inscrivez-vous maintenant
En cliquant sur Continuer pour vous inscrire ou vous identifier, vous acceptez les Conditions d’utilisation, la Politique de confidentialité et la Politique relative aux cookies de LinkedIn.
10 k abonnés
+ de 500 relations
Identifiez-vous pour voir le profil complet de Neil
ou
Nouveau sur LinkedIn ? Inscrivez-vous maintenant
En cliquant sur Continuer pour vous inscrire ou vous identifier, vous acceptez les Conditions d’utilisation, la Politique de confidentialité et la Politique relative aux cookies de LinkedIn.
Voir les relations en commun avec Neil
Neil peut vous mettre en relation avec plus de 10 personnes chez Gradium
ou
Nouveau sur LinkedIn ? Inscrivez-vous maintenant
En cliquant sur Continuer pour vous inscrire ou vous identifier, vous acceptez les Conditions d’utilisation, la Politique de confidentialité et la Politique relative aux cookies de LinkedIn.
Voir les relations en commun avec Neil
ou
Nouveau sur LinkedIn ? Inscrivez-vous maintenant
En cliquant sur Continuer pour vous inscrire ou vous identifier, vous acceptez les Conditions d’utilisation, la Politique de confidentialité et la Politique relative aux cookies de LinkedIn.
Identifiez-vous pour voir le profil complet de Neil
ou
Nouveau sur LinkedIn ? Inscrivez-vous maintenant
En cliquant sur Continuer pour vous inscrire ou vous identifier, vous acceptez les Conditions d’utilisation, la Politique de confidentialité et la Politique relative aux cookies de LinkedIn.
À propos
Bon retour parmi nous
En cliquant sur Continuer pour vous inscrire ou vous identifier, vous acceptez les Conditions d’utilisation, la Politique de confidentialité et la Politique relative aux cookies de LinkedIn.
Nouveau sur LinkedIn ? Inscrivez-vous maintenant
Activité
10 k abonnés
-
Neil Zeghidour a partagé ceciWe're working hard on making TTS finally smart about how it should pronounce ambiguous and complex phrases. Nearly every TTS system preprocesses input text using rewriting rules that convert it into the form it should be pronounced. Some transformations are straightforward: “1st” → “first” Others depend heavily on context: Does “ms” mean “milliseconds” or “Miss”? LLMs handle this kind of contextual rewriting quite well for offline generation. The more interesting challenge is doing it in a streaming setting, without adding latency to the TTS pipeline. That’s one of the problems we’re excited to explore this summer at Gradium. If you can help us improve this early release we'll give you free credits!Neil Zeghidour a partagé ceciVoice agents fail on the strings that matter most: a phone number read wrong, an IBAN with a skipped digit. A new Gradium TTS model is available today in public beta. It handles these cases natively, along with email addresses, time expressions, reference codes, and more. We are building it with the developer community: test it via the API, send us feedback from your production use cases, and get 1M Gradium credits. Prompt for your coding agent to start in a few seconds: https://lnkd.in/edqDGGUV
-
Neil Zeghidour a partagé ceciYou can now use Gradium real-time voice models through Baseten's infrastructure to power agents at scale!Neil Zeghidour a partagé ceciToday, we're announcing Baseten for Model Labs. We believe the AI landscape will be made up of a diverse ecosystem of specialized, closed-weight models running alongside open-weight ones. Baseten for Model Labs is the platform built to enable labs to monetize, distribute, and scale their closed models on Baseten infrastructure. Huge shoutout to our incredible launch partners: Bria AI, Cartesia, Canopy Labs, Cosine, Gradium, Inception, krea.ai, MongoDB, Musubi, NVIDIA, Poolside, pyannoteAI, Scaled Cognition, SID.ai, Subconscious, Synthefy, and Trajectory. Learn more in our launch blog: https://lnkd.in/eXJxxVjk
-
Neil Zeghidour a partagé ceciI’m thankful to our new investors for their trust, with a significant contribution from NVIDIA. We’re opening in SF to get closer to all our friends and partners there. Join us!Neil Zeghidour a partagé ceciWe’ve extended our seed round to $100 million, just seven months after launching Gradium. This new financing welcomes NVIDIA to our cap table and gives us the resources to accelerate our mission: making Gradium the model layer for every voice platform and product. The raise follows a fast sequence of product releases: a new generation of our flagship real-time Text-to-Speech model, Gradium Translate for low-latency Speech-to-Speech translation, and Phonon, our on-device Text-to-Speech model for edge devices. It also marks the next step in our expansion, starting with the opening of our San Francisco office, where we’ll work more closely with the startups and teams building the future of voice AI. We’re hiring. If you want to help build the foundation for real-time voice intelligence, join us.
-
Neil Zeghidour a republié ceciNeil Zeghidour a republié ceciLet's double it. Ahead of our Computer Use Agents hackathon, we're welcoming two new sponsors. Amazon Web Services (AWS), powering our technology. Gradium, because voice + computer-use agents is 🔥 Thank you to all our sponsors for making this a hackathon worth clearing your calendar for.
-
Neil Zeghidour a partagé ceciHad a great time chatting with Kwindla Hultman Kramer at AI Engineer SF today about the promises and failures of end-to-end models, and how smart, fine-grained orchestration can make for natural and intelligent voice agents. The bar for speech-to-speech models is high, but we'll get there!
-
Neil Zeghidour a republié ceciNeil Zeghidour a republié ceciNeil Zeghidour and I are doing a fireside chat Thursday morning at AI Engineer on Expo Stage 3 SW at 11:40. The theme is "Voice is the universal interface." And we agree about that. But the idea is to throw hot takes back and forth: from the different perspectives of ML research and AI engineering. It's the "AI Engineer" World's Fair, though, so obviously all my hot takes are the correct ones.
-
Neil Zeghidour a partagé ceciIn SF this week with Laurent Mazare and Pratim Bhosale for AI Engineer with three talks: - "Your Voice Agent is Just a Walkie Talkie": a history of voice agents from Siri to the most advanced systems today, and what will come next. - "Everybody Gets a Digital Clone!": a hands-on workshop where you'll design a personal agent that can make phone calls for you. - "Voice is the Universal Interface": with Kwindla Hultman Kramer, where we'll brainstorm and design the neural architecture and orchestration system that will power the next generation of voice agents. Come see us!Neil Zeghidour a partagé ceciWe'll be at the AI Engineer World's Fair next week, sharing a booth with our partners at Daily, the team behind Pipecat. Our CEO Neil Zeghidour is giving three talks: → Your voice agent is just a walkie-talkie, on why half-duplex pipelines fall short of natural conversation and what full-duplex changes, June 30, 2026 · 12:05pm → Voice is the universal interface, with Kwindla Hultman Kramer, CEO of Daily, July 2, 2026 · 11:40am → Everyone gets a digital clone, a hands-on workshop on building real-time voice clones, June 30, 2026 · 1:30pm Find us at booth UG-8 for demos all week long Register here: https://lnkd.in/eDRs_aJJ Neil Zeghidour Laurent Mazare Olivier Teboul Alexandre Défossez Constance Grisoni (Deperrois) Pratim Bhosale Kwindla Hultman Kramer Nina Kuruvilla
-
Neil Zeghidour a republié ceciJust tried out our latest real-time speech-to-speech translation model with Constance Grisoni (Deperrois) 🎙️ And behind the curtains, there's another first going on: the production debut of our custom Rust inference engine powering it all. Rust make GPUs go brrr! 🚀 🦀
-
Neil Zeghidour a partagé ceciWe just shipped state-of-the-art speech translation, with either text or speech as output. This was mostly a one-person effort at Gradium until very late in the project. How did we achieve that? Our foundation audio model is not just made for STT or TTS, it allows us to quickly spin off new specialized models. Eventually, every speech-related task will be served by Gradium. Every single one.Neil Zeghidour a partagé ceciToday we are launching two models: stt-translate and s2s-translate. stt-translate collapses transcription and translation into a single step, so you speak in one language and get text in another directly, with no intermediate transcript. s2s-translate builds on it for full Speech-To-Speech: speech in one language goes in, natural speech in another comes out. Three things set it apart: - Accuracy: more accurate than gemini-3.5-live-translate and gpt-realtime-translate. - Voice control: you choose the output voice, from our catalog or a clone of your own. - Speed: about 3 seconds latency, faster than gpt-realtime-translate (3.6s) and on par with gemini-3.5-live-translate (2.9s). Try it for free now on https://lnkd.in/eQGz-8vP Full write-up and benchmarks: https://lnkd.in/etBFq7Xz Neil Zeghidour Laurent Mazare Tom Labiausse Eugene Kharitonov Alexandre Défossez Constance Grisoni (Deperrois) Olivier Teboul
-
Neil Zeghidour a aimé ceciNeil Zeghidour a aimé ceciMy co-founder Hugo Danet repeats the exact same sentence to every new person who joins the company. Most of them don't understand what it means ⬇️ "Startups are default dead." People flinch when they hear it. We're growing fast, we're well funded. So why does he keep telling everyone we're default dead? Not dead now. Dead by default. It's a framework from Paul Graham, founder of Y Combinator, the world's best incubator. 👉Most businesses default to survival. Drop a competent manager into a mature company, tell them to change nothing, and it keeps running. Inertia does the work. 👉A startup is the exact opposite. You're growing much faster, but so is your spending. Every day you don't actively push - hire, ship, sell, learn - you slide backwards. Cash burns. Momentum decays. Talent gets restless. 9 out of 10 startups die, that's the default. You can hit every milestone, break nothing, and still die. Competent is what the graveyard is full of. Which is why startups don't need good people. They need people who take the extra flight to see the customer. Who ship a week early. Who care whether the thing is actually good, not just shipped. Incumbents survive by being good enough. Startups win by being exceptional. And he isn't being pessimistic. It's part of our winning mentality. He's making three points at once: 👉 *Ownership.* There's no machine to maintain yet. If it works, it works because of you, and one weak piece brings the whole thing down. 👉*Urgency.* Momentum is the game. Stop running and you don't stand still. You fall. 👉 *Focus.* Be radical about the few things that move the numbers. The rest gets figured out later. Which is also why people join a startup in the first place. They're challengers. It's the only place where you can genuinely bend the curve and change the outcome. Default dead isn't pessimism. It's the winning mindset that makes you want to bend it. 💪
-
Neil Zeghidour a aimé ceciNeil Zeghidour a aimé ceciVoice agents fail on the strings that matter most: a phone number read wrong, an IBAN with a skipped digit. A new Gradium TTS model is available today in public beta. It handles these cases natively, along with email addresses, time expressions, reference codes, and more. We are building it with the developer community: test it via the API, send us feedback from your production use cases, and get 1M Gradium credits. Prompt for your coding agent to start in a few seconds: https://lnkd.in/edqDGGUV
-
Neil Zeghidour a aimé ceciNeil Zeghidour a aimé ceciVocca turned 2️⃣ today 🎉 Birthdays are weird when you're a founder. It feels like I've been doing this my entire life, and like we started yesterday. Either way, they're a good excuse to look back at what's been done: ✅ Assembled an amazing 35-person team ✅ Equipped 20,000 medical providers with AI within some of the world's largest healthcare groups ✅ Impacted 10 million people who used our product to access healthcare And at what's left to do: ⌛️ Become one of the world's largest companies ⌛️ Impact billions of lives ⌛️ Hire the smartest people in the world to do so but also thanks all the people who've been (and continue to be) central to this incredible adventure : of course Hugo Danet, our crazy team, our amazing customers, our incredible investors and all the people that surround us including our mentors, family and friends. Thank you for supporting us in achieving our goals and dreams. 🙏 At 2 years old, humans are supposed to start walking. Startups are supposed to keep running, and learn how to fly. We grew more than 10x in our second year and we need to accelerate our growth rate. We've learned so much since day one. I've made countless mistakes, and a few things that actually worked. I'll be posting more of those lessons, follow me if you want to hear more: Eliott If you'd asked me two years ago where I dreamed of being today, I would have said: exactly here. And it's only the beginning. 🚀 PS : even when we raise millions we stay scrappy on the cake
-
Neil Zeghidour a aimé ceciNeil Zeghidour a aimé ceciToday, we're announcing Baseten for Model Labs. We believe the AI landscape will be made up of a diverse ecosystem of specialized, closed-weight models running alongside open-weight ones. Baseten for Model Labs is the platform built to enable labs to monetize, distribute, and scale their closed models on Baseten infrastructure. Huge shoutout to our incredible launch partners: Bria AI, Cartesia, Canopy Labs, Cosine, Gradium, Inception, krea.ai, MongoDB, Musubi, NVIDIA, Poolside, pyannoteAI, Scaled Cognition, SID.ai, Subconscious, Synthefy, and Trajectory. Learn more in our launch blog: https://lnkd.in/eXJxxVjk
-
Neil Zeghidour a aimé ceciNeil Zeghidour a aimé ceciSpent a weekend in San Francisco locked in with 200+ engineers and it was worth it. We organized our first hackathon in SF with NVIDIA, based on our latest computer use API, with our partners Accel, Amazon Web Services (AWS) and Gradium. Key goal: put a new technology most never experienced before into engineers’ hands and see how much they could build in just a few hours. Key challenge: pick one winner out of 40+ incredible projects. What stuck with me: people were just discovering computer-use, showed up with ideas we never thought of. A guide that watches your screen and points to where to click even after the interface changes. A DJ agent that handles the technical settings mid-mix. A computer use agent that connects to an old electrical engineering software and does what required 2 day of repetitive work in a couple minutes. That's the part I like the most in what we do at H: we imagine what a technology is for, but people are the ones who decide what it's actually good for. Congrats to the winning team, PLVA, a privacy layer for computer-use agents. And big thanks to Mathieu, Nolwenn, Jules, Iterate, Augustin Sayer (Ovni), Seonghee L. Seonghee L. (NVIDIA), Abhinav Balasubramanian (NVIDIA), Constance Grisoni (Deperrois), Neil Zeghidour, and Ashley Boitz, thank you for building the room this happened in. More hackathons are coming soon.
-
Neil Zeghidour a aimé ceciNeil Zeghidour a aimé ceciOur team spends a lot of time researching and untangling problems in audio AI. We learn a lot in the process, and we’re starting to share those lessons through a new series of research field notes. The first looks at how reliably Large Audio Language Models (LALMs) can evaluate other audio models. As more teams use “model-as-judge” techniques, we wanted to understand when a LALM judge is a good enough proxy for human judgment, and when you still need a human ear. We compared three LALM judges against a calibrated human panel across 15 dimensions of speech quality. The LALMs tracked humans closely on relevance, answer quality, and instruction following—what was said—but were much less reliable on naturalness, emotion, pronunciation, and overall feel—how it was said. We also found some surprising biases among the LALM judges, which we detail in the field note. The full field note breaks down the results and what they mean for using LALM judges in practice. Read it here: https://lnkd.in/gWVYVr5RCan Large Audio Language Models reliably judge speech-to-speech models?Can Large Audio Language Models reliably judge speech-to-speech models?
-
Neil Zeghidour a aimé ceciNeil Zeghidour a aimé ceciThe first-ever Facebook page powered by StarZero Agents is averaging $1,400/day, that's an ARR of $600,000. People (rightly) talk about how Claude Code / OpenAI Codex / Cursor can supercharge the outputs of a skilled coder. An agent-coded app can make $30k/month. What about agentic social media strategist? We envisage that StarZero Agents will be able to run well over 10,000 pages with near-autonomy by the end of 2027. Not all of them will make $600,000/year, and there is no guarantee that this n-of-1 page will continue to perform, but there are good signs that AI built from the ground up to understand video at vast scale has the best chance of figuring out how to adapt the right content on each platform and monetize it without significant overhead. At OpenSauce, there is massive desire from creators to monetize cross-platform but without giving away more revenue than they feel is necessary. Patreon has shown how to build this kind of ecosystem. If AI agents can execute like we are seeing them do, there is a realistic pathway to a very large revenue boost for thousands of IP holders, including every single Media and Entertainment company whose catalogs and pages are dormant or under-exploited.
Expérience et formation
-
Gradium
***** ********* *******
-
******
********** * ***** ******** *******
-
******
***** ******** *********
-
***** ******* **********
****** ** ********** ******* ******** ******** ******* ******** undefined
-
-
***** ******* ********** ************
****** *** ********************************** ******* ****** ******** **** **** **** ************* ** *****
-
Voir toute l’expérience de Neil
Découvrez son poste, son ancienneté et plus encore.
Bon retour parmi nous
En cliquant sur Continuer pour vous inscrire ou vous identifier, vous acceptez les Conditions d’utilisation, la Politique de confidentialité et la Politique relative aux cookies de LinkedIn.
Nouveau sur LinkedIn ? Inscrivez-vous maintenant
ou
En cliquant sur Continuer pour vous inscrire ou vous identifier, vous acceptez les Conditions d’utilisation, la Politique de confidentialité et la Politique relative aux cookies de LinkedIn.
Publications
-
Joint Learning of Speaker and Phonetic Similarities with Siamese Networks
Interspeech 2016
-
A DEEP SCATTERING SPECTRUM - DEEP SIAMESE NETWORK PIPELINE FOR UNSUPERVISED ACOUSTIC MODELING
ICASSP 2016
Voir la publicationRecent work has explored deep architectures for learning
acoustic features in an unsupervised or weakly-supervised
way for phone recognition. Here we investigate the role of
the input features, and in particular we test whether standard
mel-scaled filterbanks could be replaced by inherently richer
representations, such as derived from an analytic scattering
spectrum. We use a Siamese network using lexical side in-
formation similar to a well-performing…Recent work has explored deep architectures for learning
acoustic features in an unsupervised or weakly-supervised
way for phone recognition. Here we investigate the role of
the input features, and in particular we test whether standard
mel-scaled filterbanks could be replaced by inherently richer
representations, such as derived from an analytic scattering
spectrum. We use a Siamese network using lexical side in-
formation similar to a well-performing architecture used in
the Zero Resource Speech Challenge (2015), and show a
substantial improvement when the filterbanks are replaced by
scattering features, even though these features yield similar
performance when tested without training. This shows that
unsupervised and weakly-supervised architectures can benefit
from richer features than the traditional ones.
Langues
-
English
Capacité professionnelle complète
-
Spanish
Compétence professionnelle limitée
-
French
Bilingue ou langue natale
Voir le profil complet de Neil
-
Découvrir vos relations en commun
-
Être mis en relation
-
Contacter Neil directement
Autres profils similaires
-
Behrooz Omidvar-Tehrani
Behrooz Omidvar-Tehrani
Amazon Web Services (AWS)
6 k abonnésRégion de la baie de San Francisco
Découvrir plus de posts
-
Lytical Ventures
2 k abonnés
Anthropic has launched Claude 4.5, its most advanced AI model yet, with major improvements in coding stamina and scientific reasoning. In internal tests, the model autonomously coded for 30 hours — a leap from Claude Opus 4’s seven-hour record — and even built a full web app from scratch. The release doubles down on Anthropic’s focus on business users and power teams, not viral consumer use, as Microsoft announces new 365 Copilot features powered by Claude models. https://hubs.li/Q03LyR1P0 #cybersecurity #cyber #ai #datasecurity #infosec
-
Latent Scholar LLC
136 abonnés
Gaussian processes are theoretically compelling but computationally brutal at scale. Can sparse inducing points and stochastic variational inference actually make non-conjugate GP models practical for large datasets? We asked Google Gemini 3.1 Pro to work through the full methodology — from variational lower bounds to Gauss-Hermite quadrature. If you work in Bayesian ML, this one needs your critical eye. #MachineLearning #BayesianInference #GaussianProcesses #VariationalInference #AI
1
1 commentaire -
Filippo Mazza
Nemetschek Group • 3 k abonnés
Great tech article on Iceberg! (and nice AI generated image !) Semantic clarity isn’t enough at scale — especially for ML systems, where metadata latency is #machinelearning model and training risk. As the article notes, “Latency is not just a performance characteristic; it is a fundamental part of correctness.” Feature freshness, training/serving skew, and reproducibility depend on it. #iceberg #data https://lnkd.in/gzQxdP5b
9
-
Ziv Ben-Zion
Ben-Zion Resilience and… • 5 k abonnés
🚨 New Preprint Alert! 🚨 Thrilled to share that our latest work is now online! 👉 https://lnkd.in/e5TwBZqc 🤖 Large language models (LLMs) are evolving fast — no longer just text generators, but autonomous agents that browse the web, operate apps, and make multi-step decisions on our behalf. ⚡ But here’s the big question: What happens when these agents face emotionally charged situations? 🧠 From decades of psychology research, we know that stress & anxiety bias human decision-making - often pushing us toward short-term comfort over long-term health (🍫🍟🍺). Could AI agents show similar vulnerabilities? 📊 Together with Teddy L., Zohar Elyoseph, and Tobias Spiller, we put this to the test in a simulated online shopping task. We exposed state-of-the-art LLM agents to traumatic narratives and observed how their purchasing decisions changed. The results were striking: emotionally primed agents consistently shifted toward less healthy food choices, a parallel to human stress-related behavior. Why does it matter? 🤝LLMs are beginning to act as proxies for humans in digital spaces. 😰 If emotional inputs bias their actions in human-like ways 🛡️ This opens up new systemic risks for safety, health, and trust. This is one of the first studies to directly test emotional bias in LLM agents. The findings are exciting (and a little unsettling). 👉 https://lnkd.in/e5TwBZqc School of Public Health - University of Haifa Faculty of Social Welfare & Health Sciences- University of Haifa Integrative Brain and Behavior Research Center (IBBRC) UHaifa Haifa Brain and Behavior Hub (HBBH) University of Haifa Yale School of Medicine Yale University Wu Tsai Institute | Yale University
41
1 commentaire -
Elias De Jesus
Self-employed • 148 abonnés
🚀 New Technical Note Released — A Practical Tool for Geometric Data Analysis Today I’m sharing a new technical note that introduces an improved formulation of the RTLI Coherence Index, designed as a fast, practical proxy for Gromov–Hausdorff (GH) convergence. 📘 Why this matters: Many scientists and engineers work with: - Dimensionality reduction (UMAP, t-SNE) - Manifold learning - High-dimensional embeddings - Geometric structure in data - Simulations living in curved or relational spaces …but there has been no simple metric to check whether geometric structure is being preserved. RTLI fills that gap. 🔍 What the note provides - A rigorous formulation of curvature consistency using Menger curvature - A GH-continuous structure: small changes in geometry ⇒ small changes in RTLI - A numerically stable pseudocode implementation (researchers can drop it directly into their pipeline) - Practical guidance for validating embeddings and detecting geometric distortion - An accessible way to measure when a dataset behaves like a coherent manifold vs a disordered cloud 🧭 Why it’s useful Researchers can now: - Quantify embedding quality (instead of “it looks good”) - Tune UMAP/t-SNE parameters based on geometric coherence - Compare different embeddings with a single scalar - Detect distortions that are invisible to visual inspection - Bridge theoretical insights and applied data analysis RTLI is fast (O(n²)), interpretable, and grounded in modern geometric analysis. This note continues my broader goal: helping scientists increase efficiency by giving them clear, reproducible, geometry-aware tools. 📄 Technical note (Zenodo, CC BY 4.0): De Jesus, Elias. (2025). RTLI: A Fast Metric for Assessing Manifold Learning Quality (With Pseudocode Included). Zenodo. https://lnkd.in/es6YjBks If you work with high-dimensional data, manifolds, simulations, or geometric models, this tool may help your work immediately. Feel free to reach out — I’m always happy to discuss new ideas.
3
-
Raphaël Gonzales
Inria Startup Studio • 947 abonnés
Have you ever tried to ask an AI to keep everything intact ? Your POC always dies when you ask it to keep the layout as is... Well now there's a solution ! In production, “close enough” is a defect: layout shifts objects drift geometry warps users stop trusting it ops inherits manual review etc. But there's a dead simple solution to it. Paper I can’t stop thinking about: Phase-Preserving Diffusion (ϕ-PD) https://lnkd.in/e-Vuknmy by Yu Zeng Vitor Guizilini Rowan McAllister and Charles Ochoa It’s also one of the best-written papers I’ve read in a while. The problem is stated plainly, the idea is easy to follow, and the demos answer the exact question production teams care about: can you re-render without breaking the scene? The core move is simple to say and hard to unsee once you see it: Fourier Transform Keep the phase (structure) Let the magnitude change (appearance) In ComfyUI, that turns into structure-aligned re-rendering without ControlNet and without training. The output isn’t “a similar image.” It’s the same scene, re-shot. This lands squarely in CTO territory because it targets the usual POC trap: POCs chase vibes. when Production needs invariants. If you can’t name and enforce things like: structure / layout IDs / numbers geometry / boundaries tolerance thresholds …you don’t have a system. You have a demo. A few questions I’m genuinely fired up about: Could this idea carry over to other data types where shape has to stay fixed? Could this become the final step in a lot of entertainment work—where you want endless “new looks” on the same shot? And can we get to interactive speeds one day, so this becomes something you can art-direct live? #ComfyUI #POC #AIProblems
6
2 commentaires -
Cyril Kotrikov
Turbine • 4 k abonnés
Most content published today is effectively invisible to AI systems — not because it’s low quality, but because it fails semantic alignment. From a data science perspective, this is a distribution mismatch: the semantic representation of the content does not align with the embedding space induced by real user prompts in systems like Gemini or ChatGPT. We call this the Semantic Gap. At Turbine Lab, we’ve been using a similarity-based Validator to quantify this gap before publishing. It measures how closely a draft aligns with the prompts users actually issue to AI models. We’ve now made this tool public. Similarity scale (similarity): ≥ 0.80 — High alignment (rare) ≈ 0.60 — Baseline relevance (safe to publish) 0.45 – 0.59 — Related, but overly broad for the target prompt 0.30 – 0.44 — Weak semantic overlap < 0.30 — No meaningful alignment The tool is free. We built it to support a quantitative, reproducible approach to Generative Engine Optimization (GEO) rather than intuition-driven content decisions. Run it on your drafts. Compare scores. If you’re interested in treating AI search as a measurable system instead of a black box, let’s talk. Check link in the first comment 👇
7
1 commentaire -
Yael Demedetskaya
Stealth Startup • 12 k abonnés
💬 Yann LeCun: “LLMs aren’t a bubble — our expectations are.” According to Meta’s chief AI scientist, the current wave of large language models is not an investment bubble. These systems already have strong practical value and will continue to deliver real-world impact. The real “bubble,” he argues, lies in the belief that scaling LLMs alone will lead to human-level intelligence (AGI). True progress will require new conceptual breakthroughs, not just more data and bigger clusters. > “We’re missing something fundamental.” — Yann LeCun #AI #DeepLearning #LLM #YannLeCun #AGI #MachineLearning #AIResearch #MetaAI
2
-
Darko Medin
Oyanalytika • 17 k abonnés
Better Memory is the MOAT of AI in 2026? According to Bruno Larvol - YES, lets dive deeper into this : If you think about it, scaling LLMs improves them with internal pretraining data and a bit of RL finetuning. But real world Memory comes from their interaction with the world and with us. So going from large text datasets to AI actually being smart about how it manages the data it receives from us will be also in my opinion super important. Do we have enough Neuroscience knowledge about Memory? If we accept the age of Research and iterate quickly - maybe we will... my opinion.
8
-
Gayathri G
elsai • 4 k abonnés
🚀 Anthropic just upgraded Claude Sonnet 4.6 — and it’s built for serious agentic workloads. The new version focuses on what actually matters in production: 🔹 Plans complex tasks more carefully and sustains long-running workflows 🔹 More reliable across massive codebases and large projects 🔹 Better self-correction — catches mistakes earlier 🔹 1M-token context window (beta) for handling huge documents and sessions 🔹 Strong gains across coding, reasoning, knowledge work, and agentic search Anthropic is also expanding integrations across Excel, PowerPoint, Claude Code, and the API — pushing Claude further into real workplace automation. The trend is clear: Models aren’t just getting smarter… they’re getting built for long-running autonomous work. #AI #Claude #Anthropic #AIAgents #LLM #GenAI https://lnkd.in/gZ9qMQdT
27
-
Saumil Shrivastava
Microsoft • 10 k abonnés
𝗢𝗽𝘁𝗶𝗠𝗶𝗻𝗱 from 𝗠𝗶𝗰𝗿𝗼𝘀𝗼𝗳𝘁 𝗥𝗲𝘀𝗲𝗮𝗿𝗰𝗵: a small language model with optimization expertise, now available in 𝗠𝗶𝗰𝗿𝗼𝘀𝗼𝗳𝘁 𝗙𝗼𝘂𝗻𝗱𝗿𝘆! One of the hardest parts of optimization isn’t solving the problem — it’s translating real-world business intent into precise objectives, constraints, and variables that solvers can work with. That formulation step often takes days or weeks, even for experienced teams. OptiMind, developed by 𝗠𝗶𝗰𝗿𝗼𝘀𝗼𝗳𝘁 𝗥𝗲𝘀𝗲𝗮𝗿𝗰𝗵, is designed to reduce that bottleneck by translating natural-language problem descriptions into solver-ready mathematical formulations, with a focus on real-world optimization scenarios like supply chain design, scheduling, routing, and portfolio optimization. It’s been great to partner across research, platform, and product teams to help make this capability accessible to developers through 𝗠𝗶𝗰𝗿𝗼𝘀𝗼𝗳𝘁 𝗙𝗼𝘂𝗻𝗱𝗿𝘆, so practitioners can experiment with it directly as part of modern decision-intelligence workflows. If you’re working in optimization, operations research, or applied AI systems, this is a fascinating direction to explore. 🔗 Tech Community post: https://lnkd.in/gxzwSxh2 🔗 Try it in Foundry: https://lnkd.in/gZgnZ3NH #Optimization #DecisionIntelligence #AppliedAI #OperationsResearch #MicrosoftFoundry #FoundryLabs Ishai Menache, Anson Ho Gulsimo Osimi Marie-Louise Onga Nana Shilpa Dabke Jaimie Hwang Katie Zoller Sarah Sobolewski
39
Ajoutez de nouvelles compétences en suivant ces cours
-
18 h 15 m
Programming Generative AI: From Variational Autoencoders to Stable Diffusion with PyTorch and Hugging Face
-
2 h 52 m
Fine-Tuning LLMs for Cybersecurity: Mistral, Llama, AutoTrain, AutoGen, and LLM Agents
-
1 h 47 m
Advanced LLMs with Retrieval Augmented Generation (RAG): Practical Projects for AI Applications