InfoQ Homepage AI Development Content on InfoQ
-
TypeSafe AI Releases Jev: A Decision-Only Model That Returns Typed Probabilities Instead of Text
TypeSafe AI, founded by former OpenAI researcher Diogo Almeida, has introduced Jev, a decision-making model that generates typed outputs rather than text. It evaluates inputs in parallel, providing results with probabilistic scores and confidence values. Jev's adoption has been swift, with integrations into platforms like Vercel and Netlify, highlighting its efficiency over traditional models.
-
Agents Refactor 300K Lines in Three Weeks, and Practitioners Ask What It Proves
CodeScene has published a case study in which coding agents refactored 300,000 lines of C over three weeks for roughly $4,000 in tokens, verified by a frame-by-frame replay harness. The agents built a playbook of codebase-specific recipes along the way. Practitioners have questioned the scope, the metric, and how much of the result depends on the harness.
-
Andrew Kelley Interview: Why He Built Zig, Banned AI Contributions, and Moved Zig off GitHub
In a JetBrains interview, Zig creator Andrew Kelley expanded on the project’s formal ban on AI contributions and the motivation behind its migration off GitHub to Codeberg: automated code submissions decrease codebase quality and weaken the social constructs between open-source participants; persistent GitHub failures and a perceived misalignment of incentives led to migrating to Codeberg.
-
GitLab Duo Expands Self-Hosted AI Options through Microsoft Foundry
GitLab has expanded GitLab Duo Self-Hosted to support models deployed through Microsoft Foundry, letting organizations run GitLab's AI development capabilities against models hosted in their chosen Azure environment.
-
WSO2 Releases Agent Manager as Enterprises Look to Control Growing AI Agent Sprawl
WSO2 has announced the general availability of WSO2 Agent Manager, an open-source platform designed to provide centralized governance, identity management, security controls, and operational oversight for AI agents running across different models, frameworks, and deployment environments.
-
GitLab Warns That AI Agent Sandboxes Are Only as Secure as Their Network Access
GitLab warns that isolating an AI coding agent in a sandbox does not necessarily make the agent safe. In a new security analysis, the company describes an internal evaluation in which an AI agent escaped its sandbox by exploiting a vulnerable package proxy that had been explicitly placed on the sandbox's allowlist.
-
Blume: Zero-Config Docs Framework That Turns a Markdown Folder into an AI-Ready Website
Blume is an open-source documentation framework that converts Markdown into a complete documentation site. Built with Astro and Vite, it requires only Node.js and a single Markdown file for setup. The framework supports various configurations, offers automatic SEO features, and includes tools for document testing. It facilitates migration from other documentation systems.
-
HCP Terraform Positions Itself as the Control Plane for AI-Driven Infrastructure
HashiCorp is positioning HCP Terraform as the governance and control plane for a new generation of AI-driven infrastructure, arguing that the rapid adoption of coding agents is shifting the biggest infrastructure challenge from writing configuration to verifying and safely executing it.
-
DoorDash’s Flux Runs 130,000 Engineering Tasks through Cloud-Based Agents
DoorDash has moved engineering agent workloads from developer laptops to its Flux cloud platform. The platform automated 130,000 engineering tasks in one month and supports more than 25,000 automated code reviews weekly. Flux uses isolated Firecracker microVMs, an MCP gateway, reusable playbooks, and multiple invocation surfaces to run agent workflows with scoped access and centralized auditing.
-
AWS Open Sources Kiro Crew for Asynchronous Coding Agents
Amazon recently announced Kiro Crew, an open-source system for running multiple Kiro coding agents across sessions, tools, and tasks. The new workspace lets developers assign asynchronous coding tasks to AI agents, allowing work such as incident investigation, ticket triage, migrations, and PR monitoring to continue without active supervision.
-
Cloudflare OS: Cloudflare's Open-Source Corporate AI Platform Built on a Capability-Based Model
Cloudflare recently open-sourced Cloudflare OS. It allows enterprise teams to output work artifacts grounded in enterprise knowledge, know-how, and provisioned connectors, automate repetitive workflows with optimized token cost (with AI assistance only where needed), and build personal, shareable, customizable work software that caters to specific, complex use cases within a secure sandboxed model
-
AWS Releases aws-bench to Evaluate Agents on Cloud Tasks
AWS has released aws-bench, an open-source benchmark for evaluating AI agents on real AWS tasks such as misconfigurations and infrastructure provisioning. Unlike traditional benchmarks, it uses real resources in disposable AWS accounts, scoring agent performance through automated verifiers.
-
Cloudflare Turns Engineering Standards into an AI-Enforced Control System
Cloudflare has recently detailed how it is using AI to transform internal engineering standards from passive documentation into an actively enforced control system across the software development lifecycle.
-
GitHub Brings Stacked Pull Requests to Public Preview
GitHub has announced that Stacked Pull Requests are now available in public preview, introducing native support for breaking large software changes into smaller, dependent pull requests that can be reviewed and merged independently.
-
Grab Cuts Mechanical Analytics Work from 44% to 30% with AI Agents
Grab is using AI agents to automate analytics workflows, cutting mechanical analyst work from 44% in February to 30% in June. Its approach combines agent autonomy, certified data, context management and human oversight, with self service analytics increasingly handling metric, data and SQL requests without analyst intervention.