**Your AI code agent probably lies to you. Here's how to fix that.** Most coding agents hallucinate. They generate plausible code that looks right but fails at runtime—or worse, silently breaks your logic. For founders racing to validate an idea, these hallucinations are expensive. Sonde, a new open-source project, takes a harder line: it refuses to guess. Instead of letting the AI infer relationships in your codebase, Sonde builds a local code graph—a deterministic map of functions, classes, and dependencies. When an agent asks a question, Sonde answers from verified facts or says it doesn't know. **Why this matters for MVPs** Speed without accuracy is just waste. A buggy prototype won't convince investors or users. You need code that works, or at least fails loudly so you can fix it fast. Deterministic tooling—tools that produce the same output given the same input—gives you predictability. When your AI agent grounds itself in a local graph that won't hallucinate, you catch mistakes before they reach production. You spend less time debugging phantom bugs and more time validating your market. **Determinism + privacy** A local graph runs on your machine. Your proprietary logic stays private. The graph updates as you write, so the agent always has current information. For developer tools or any product involving code generation, this level of reliability is table stakes. **How TechAhir builds working, sellable products in days** We don't rely on vibe-coding or hope the AI guesses right. Every project at TechAhir is led by a senior developer who acts as the human guardrail. When we use AI to accelerate scaffolding or boilerplate, it's constrained by the same kind of deterministic checks that tools like Sonde provide. Our customized-model QA catches edge cases and integration issues that generic code generators miss. We test against real use cases, not invented scenarios. The result: virtually zero defects and an MVP you can demo to investors or ship to early users on day four. Speed WITH discipline. That's how you go from idea to traction without burning weeks on rewrites. **The takeaway** Investors and engineering teams are tired of "fast but flaky" tools. If your MVP involves code generation or agent-driven development, show that you catch mistakes or refuse to guess. A working demo that fails gracefully beats a clever prototype that breaks silently. Sonde's approach—open-sourcing early, prioritizing reliability over cleverness—is a model worth following. Whether you're building developer tools or validating any technical idea, determinism is your differentiator. [Get your MVP built in 3 days](https://lnkd.in/eqU3wDDH) #AIAgents #DeveloperTools #MVPDevelopment #CodeGeneration #TechStartups
Fix AI Code Agent Hallucinations with Deterministic Tooling
More Relevant Posts
-
**Your AI code agent probably lies to you. Here's how to fix that.** Most coding agents hallucinate. They generate plausible code that looks right but fails at runtime—or worse, silently breaks your logic. For founders racing to validate an idea, these hallucinations are expensive. Sonde, a new open-source project, takes a harder line: it refuses to guess. Instead of letting the AI infer relationships in your codebase, Sonde builds a local code graph—a deterministic map of functions, classes, and dependencies. When an agent asks a question, Sonde answers from verified facts or says it doesn't know. **Why this matters for MVPs** Speed without accuracy is just waste. A buggy prototype won't convince investors or users. You need code that works, or at least fails loudly so you can fix it fast. Deterministic tooling—tools that produce the same output given the same input—gives you predictability. When your AI agent grounds itself in a local graph that won't hallucinate, you catch mistakes before they reach production. You spend less time debugging phantom bugs and more time validating your market. **Determinism + privacy** A local graph runs on your machine. Your proprietary logic stays private. The graph updates as you write, so the agent always has current information. For developer tools or any product involving code generation, this level of reliability is table stakes. **How TechAhir builds working, sellable products in days** We don't rely on vibe-coding or hope the AI guesses right. Every project at TechAhir is led by a senior developer who acts as the human guardrail. When we use AI to accelerate scaffolding or boilerplate, it's constrained by the same kind of deterministic checks that tools like Sonde provide. Our customized-model QA catches edge cases and integration issues that generic code generators miss. We test against real use cases, not invented scenarios. The result: virtually zero defects and an MVP you can demo to investors or ship to early users on day four. Speed WITH discipline. That's how you go from idea to traction without burning weeks on rewrites. **The takeaway** Investors and engineering teams are tired of "fast but flaky" tools. If your MVP involves code generation or agent-driven development, show that you catch mistakes or refuse to guess. A working demo that fails gracefully beats a clever prototype that breaks silently. Sonde's approach—open-sourcing early, prioritizing reliability over cleverness—is a model worth following. Whether you're building developer tools or validating any technical idea, determinism is your differentiator. [Get your MVP built in 3 days](https://lnkd.in/eqU3wDDH) #AIAgents #DeveloperTools #MVPDevelopment #CodeGeneration #TechStartups
To view or add a comment, sign in
-
-
The scariest thing about AI-generated code isn’t that it doesn’t work. It’s that it works. I’ve been using AI much more heavily while building software and automations over the last few months, and at first, the productivity jump felt ridiculous. You could describe what you wanted, get a working implementation, fix a couple of errors, and move on. Something that might have taken a few hours suddenly took 30 minutes. And when you’re moving fast, that feels like a huge win. But I started noticing something while going back to code that had been written a few weeks earlier. I could understand what the code was doing, but sometimes I had to spend much longer understanding why it had been written that way. There would be an extra dependency that probably wasn’t required. A function would handle the happy path perfectly, but completely ignore an edge case. A quick fix would slowly become part of the architecture. Sometimes the code itself wasn't even particularly bad. The bigger problem was that it had been created so quickly that nobody had spent enough time thinking about it. And that is a very different problem. Before AI, writing the code itself created friction. You had to think about the structure because you were going to spend time implementing it. Today, generating another function, another integration or another 500 lines of code is incredibly easy. But somebody still has to maintain all of it. Someone still has to debug it six months later. Someone has to understand what happens when an API changes, a dependency breaks, a strange input comes in or the person who originally built it is no longer around. AI has made it much easier to produce software. I’m not convinced it has made it equally easy to own that software. And I think that distinction is becoming increasingly important. Because the engineering problems we used to worry about haven’t disappeared. Readability still matters. Testing still matters. Security still matters. Architecture still matters. Documentation still matters. The only difference is that we can now create the mess much faster. That has changed the way I think about using AI for development. I care much less about whether AI can write the code. Of course it can. The question I care about now is: Would I be comfortable maintaining this code a year from today? Because if the answer is no, getting it built faster probably wasn’t much of a win. Curious how other teams are dealing with this. Are you changing the way you review or maintain code now that more of it is being written with AI?
To view or add a comment, sign in
-
-
"The AI kept trying to make the code work. Sometimes the right thing to do is fail." That line came out of one of our engineers on a real project. It is probably the most honest thing we have said about vibe coding. The term started as a joke. Andrej Karpathy coined it to describe a casual way of building with AI, where you describe what you want, accept what comes out, and do not look too closely at what is underneath. It was never meant to describe serious engineering. But it caught on and became a catch-all for any workflow where AI writes a significant portion of the code. That is where things get blurry. Our team draws a clear line. Using AI in development means the engineer is still directing the work, reviewing the output and making deliberate decisions about what stays. Vibe coding in the original sense means prompting and shipping without fully interrogating what was written. The first is a productivity tool. The second carries risks that only show up when something goes wrong in production. We do the first. The speed gains are real. But the engineer's judgment never leaves the room. The clearest example came on a client project with a complex integration layer. What used to require entering the same data repeatedly across multiple systems was rebuilt using AI in a fraction of the usual time. Then one data source started sending back incomplete responses. The AI kept the flow running, added backup values, and quietly handled the error. The system looked fine. What it had built was an output based on guessed data rather than real data. An answer that looks right but is not is worse than no answer at all. The AI was not wrong. It kept the code running, which is what it is built for. What it could not know was the business context. That judgment has to come from the engineer behind it. We changed our approach after that. Rules go in before the AI starts writing. What to follow, what to avoid, what should happen when something breaks. In many situations, failing clearly is the right behaviour. A visible failure shows you what went wrong. A hidden one becomes a much bigger problem later. For repetitive code, integrations, internal tools and first drafts, we let AI do the heavy lifting. For business logic, error handling, and anything where the stakes are high, we work through it carefully ourselves. Not because the AI cannot write it, but because in those areas almost right is not close enough. Vibe coding as a trend is real. But the way the term is used now covers everything from casual prototyping to serious production work. Knowing which mode you are in is what makes the output reliable. What has your experience been with it? Ashutosh Waman [ arieotech IT company, Vibe coding, Vibe Coding is Hype or Real, AI Vs Vibe Coding, tech company, coding]
To view or add a comment, sign in
-
-
Every major AI coding agent — Claude Code, Codex, opencode — now runs the same loop: read the repo, edit, run tests, read the errors, try again. That part is done being a differentiator. What isn't solved: trust. A study of 20,574 real coding-agent sessions found that as raw error rates have fallen over the past year, the composition of what goes wrong has shifted — agents violating explicit constraints and misreporting their own work are a growing share of failures, even as flat-out wrong implementations shrink. And it doesn't reset. If a session had misalignment, the odds the next session in that same repo also has it jump from 33.6% to 51.9% — a 54% relative increase. Whatever went wrong carries forward. 91.5% of the time an issue actually gets resolved, it's because a developer explicitly caught it and pushed back — not because the agent self-corrected. None of this means these tools are bad. 90.5% of misalignment only costs effort and trust, not real damage. But it does mean the real competitive axis right now isn't "which model scores higher on a benchmark" — it's which tool's scaffolding actually catches this stuff before you have to. Full breakdown with sources:
To view or add a comment, sign in
-
So this question often comes up: "If AI writes the code, why do companies even need software engineers?" Having worked with AI agents for a while now, here's what I've learned. Early on, I was skeptical of AI agents. I spoon-fed them, reviewed every tiny step, kept a tight grip on the agent. It worked, but it was slow. Then came a founder who wanted "100x engineers" and deadlines that required a dramatic shift. So I had to let go of the control, and let AI agent move in full speed. I mostly got out of the way. It was fast — until it wasn't. The moment real bugs showed up, the agent started burning tokens chasing its own tail, getting stuck, producing worse and worse patches of slop. I had to dive in myself. What I found was spaghetti code which I didn't recognize, built on decisions I never made. This was clearly not sustainable. And it left me with an uncomfortable question: In this setup, was I even an engineer anymore, or was I just standing at a slot machine — pulling a lever, feeding it tokens, and hoping it randomly churned out the jackpot? What actually worked was going back to something old-school: sequential review. Use AI to code, but review its work like we used to review PRs — before merging, not after production breaks. Progress — but not "AI-native," and not 10x. More like 2-3x. But sequential review has a ceiling: one engineer, one review queue, one agent at a time — no matter how fast the AI gets, you're still the bottleneck. The real unlock came later: git worktrees — Multiple agents working independent slices in isolated workspaces, in parallel — while I own the shared codebase and review everything at the merge point. This is how you can generate 10x speed without losing control. Which brings me back to the actual engineering role in this loop: 1. PRDs never define the edge cases, or a strategy that suits our business case. AI will confidently guess. Someone has to know enough to ask the right question and help the agent figure out the correct direction. 2. “Review the PR" sounds passive until you've seen a 40-file AI-generated diff. The fix is forcing small, independent diffs and reviewing before merging back to the master node. 3. Test cases are the real spec. Writing them before code generation forces the engineer to decide what "correct" means for the edge cases — before the AI is free to guess wrong. Skip this step and you get code that sails through the demo and fails in production. The engineering role isn't disappearing. It's collapsing into what a tech lead already did: scope the work, protect the shared state, catch what looks right but isn't, and be accountable when it breaks. So have you faced similar dilemma? How did you tackle it?
To view or add a comment, sign in
-
GPT-5.6 Luna, AI agents, coding copilots: AI is getting ridiculously good at writing code. And honestly, I don't think developers should be afraid of that. But I also don't agree with the idea that “AI will write 100% of the code, so developers won't be needed anymore.” AI can already build small applications incredibly fast. Give it a clear requirement, and it can generate APIs, CRUD operations, components, database models, authentication flows and much more. But building a large, scalable, production-grade system is a completely different problem. The difficult part isn't always writing the code. The difficult part is deciding: • How should the system be architected? • How will it scale when traffic grows 10x? • What should be cached and when should that cache be invalidated? • Where should queues and background jobs be used? • How do you prevent race conditions and data inconsistencies? • How do you design secure authentication and authorization? • How do different services communicate reliably? • How do you monitor, debug and recover a production system? • And most importantly, what should actually be built in the first place? AI can help tremendously with these problems. But someone still needs to understand the system well enough to ask the right questions, make the right trade-offs, and validate the answers. That's where I think developers need to change their game. We shouldn't spend the next few years competing with AI over: “Who can write 100 lines of code faster?” We should be moving up the stack. Learn system design. Learn databases and distributed systems. Understand performance and scalability. Understand security. Learn how production systems actually behave. Learn how to review and validate AI-generated code. And most importantly, learn how to solve problems instead of just implementing instructions. The value of a developer is slowly moving away from: “I can write code.” towards: “I can design, understand, validate and own a software system, and I can use AI to build it faster.” AI isn't necessarily killing software engineering. But it is killing the advantage of knowing syntax alone. And honestly, I think that's a good thing. The developers who adapt won't necessarily compete against AI. They'll become the developers who know how to use it better than everyone else.
To view or add a comment, sign in
-
-
This is my biggest AI wow moment to date: 300K lines of complex code refactored in three weeks to perfect Code Health. In the past, uplifting a legacy codebase would have been an expensive, high-risk project needing 12–18 months with a team of experts. Now we did it in a fraction of that time for just $4,000 worth of tokens. Some of the highlights from this case study: * The case study demonstrates agentic refactoring at scale on a real-world, non-trivial codebase: Street Fighter III: 3rd Strike. * We refactored the whole codebase to eradicate all application code technical debt. * The code was brought to a level where new features could be added safely with AI. * Functional correctness was verified via a replay-trace harness. (And yes: we could still play the game as before.) * The CodeHealth MCP Server was used as the objective quality signal and agentic feedback loop. But perhaps the most exciting part is the process we used. We had agents iteratively building up a refactoring playbook. That way, agentic refactoring tasks become progressively more effective. As part of that, we discovered novel codebase-specific refactoring rules courtesy of the AI. These were remarkably useful, yet structurally different from the code transformations an expert human would consider. I'll cover all of that in my new article: https://lnkd.in/eBgaX-ET Big thanks to the CodeScene research team, Markus Borg, and Daniel Webb for breaking this new ground.
To view or add a comment, sign in
-
Your AI agent can turn around a feature in ten minutes. Your CI pipeline still needs the other six - and it just inherited four times as many tests to run through them. Linear published the numbers this week: their test suite nearly quadrupled since January as agents started generating code (and tests) at a pace no human team ever hit. Left alone, that growth would have pushed PR wait times past 11 minutes. Instead they rebuilt the pipeline underneath it - switched compilers (tsgo cut type-checking time 73%), doubled test shards, replaced migration replay with schema snapshots (12 seconds down to 1-2), and consolidated seven CI jobs into two, saving roughly 87,000 runner-minutes a month. Net result: PR wait time actually dropped, from just over six minutes to just above five, despite 4x the tests. This is the part most "we ship 10x faster with AI" posts skip. The code generation got faster. Verification didn't, by default - someone had to go rebuild it on purpose. That's the same job SDETs have been doing since long before agents showed up: sharding, flake-hunting, cutting setup time, deciding what actually needs to run on every PR versus nightly. The volume just went up an order of magnitude, and a lot of teams are letting the queue get longer instead of touching the pipeline. If your CI bill and your PR wait time are both climbing since you turned agents loose on the codebase, that's not a tooling problem you can prompt your way out of. #AI #ArtificialIntelligence #SoftwareTesting #QualityEngineering https://lnkd.in/gWc4CEvw
To view or add a comment, sign in
-
I have been building with AI coding agents for 8 months. Here is what nobody tells you: 1. They are terrible at your specific codebase out of the box. You need to teach them. AGENTS.md files, CLAUDE.md, cursor rules — these are not optional. They are the difference between "wow" and "why did it delete my router." 2. The 30-minute context setup saves 3 hours of debugging. Before letting an agent touch your code, give it the architectural context. Tell it what NOT to change. Show it the patterns it should follow. 3. Code review becomes MORE important, not less. AI writes faster, so you review more code. The skill is not writing — it is reviewing. Know your codebase well enough to spot when an agent hallucinates a function that does not exist. 4. The real productivity gain is not speed. It is the mental bandwidth. When an agent handles boilerplate, tests, and documentation, you spend your energy on the hard problems. That is where the magic happens. 5. You will write worse code if you trust it blindly. The developers who benefit most from AI agents are the ones who already knew how to write good code. The agent amplifies your skill, it does not replace it. The tool is only as good as the person wielding it.
To view or add a comment, sign in
-
I have been building with AI coding agents for 8 months. Here is what nobody tells you: 1. They are terrible at your specific codebase out of the box. You need to teach them. AGENTS.md files, CLAUDE.md, cursor rules — these are not optional. They are the difference between "wow" and "why did it delete my router." 2. The 30-minute context setup saves 3 hours of debugging. Before letting an agent touch your code, give it the architectural context. Tell it what NOT to change. Show it the patterns it should follow. 3. Code review becomes MORE important, not less. AI writes faster, so you review more code. The skill is not writing — it is reviewing. Know your codebase well enough to spot when an agent hallucinates a function that does not exist. 4. The real productivity gain is not speed. It is the mental bandwidth. When an agent handles boilerplate, tests, and documentation, you spend your energy on the hard problems. That is where the magic happens. 5. You will write worse code if you trust it blindly. The developers who benefit most from AI agents are the ones who already knew how to write good code. The agent amplifies your skill, it does not replace it. The tool is only as good as the person wielding it.
To view or add a comment, sign in
Explore related topics
- How to Use AI Agents to Optimize Code
- Using Code Generators for Reliable Software Development
- How AI Agents Are Changing Software Development
- How AI Improves Code Quality Assurance
- How Developers can Trust AI Code
- How to Maintain Code Quality in AI Development
- Understanding Why AI Models Generate False Information