Tech firms, corporations, and utilities are on track to spend more than $1 trillion on AI infrastructure in the coming years, according to a recent report by Goldman Sachs. Despite this massive investment in artificial intelligence growth, significant profitability breakthroughs in AI remain elusive. To date, the hardware stack — chipmakers, hardware makers, and foundries for board and rack makers — have been the primary beneficiaries of capital expenditure for AI growth. Yet, there are still questions about whether the current output justifies the ROI.
The root challenge in artificial intelligence profitability stems from the interconnect bottleneck at the chip level, which impacts the unit economics at the top of the stack. In-package optical I/O represents a solution and turning point, with the ability to boost AI growth and next-generation GenAI profitability by up to 20x and improve interactivity by 3-4x.
In this blog post, we address three key questions that factor into future artificial intelligence profitability:
- Where are the biggest bottlenecks in large-scale AI systems impacting AI growth and AI market size today?
- What technologies and products can solve them?
- How much better can we do to accelerate artificial intelligence profitability and domains of performance for new applications to increase AI market size?
The Biggest Bottlenecks in Today’s AI Systems
It’s been nearly two years since the launch of ChatGPT, and it appears that the gold rush for SaaS companies and the AI industry growth has stalled. As AI models continue to grow in size and complexity, the capabilities and economics of the underlying infrastructure are failing to scale as needed with artificial intelligence market size. In fact, the unit economics of AI applications — such as dollars per token — remain far too high because of the high total cost of ownership (TCO) and limited performance of hardware architectures constrained by bandwidth distance bottlenecks arising from electrical I/O, limiting the potential of artificial intelligence and economic growth.
Our latest research outlines the challenges we face and highlights the need for a new approach to infrastructure for AI growth — one that harnesses light. This new approach is the key to unlocking artificial intelligence profitability, addressing the bandwidth distance bottleneck, and providing a roadmap for scale-up fabrics to meet the demands of growing AI market size.
The increasing demand and complexity of the artificial intelligence market size is growing faster than memory bandwidth and capacity, which is driving computing costs higher and higher. Traditional copper and pluggable optics fail to effectively scale compute advancements from the package to the cluster in today’s AI infrastructure. This results in inefficiencies, higher power consumption, and increasing costs. At the same time, copper interconnects limit the scale-up domain size to a single rack. This limits the cluster performance by constraining the effective memory bandwidth and capacity. Ultimately, this limits both artificial intelligence profitability and the interactivity of the inference model running on such a cluster. (Figure 1)
“Copper interconnects have already failed to support the AI workload in an economical manner. The industry is now in a situation where hardware builders need to dramatically increase the cost-effective throughput of these systems. Otherwise, we’re all heading towards a Dot-Com style crunch.”
Mark Wade, CEO, Ayar Labs
Figure 1: Comparison of AI throughput versus artificial intelligence profitability for one of today’s leading AI scale-up solutions and a potential future solution.
Scaling GenAI Inference Performance
Let’s start by considering the options available today: Enhancing performance in the scale-up fabric with electrical I/O requires adding more accelerators to the rack, which increases power density in the rack. Alternatively, pluggable optics can connect accelerators beyond a rack, improving the interactivity, but its high power and cost hurt artificial intelligence profitability.
This causes GenAI inference to be limited by two factors:
- Memory Bandwidth: Sets the time required to load attention cache and model weights from memory to GPU
- Scale-Up Fabric: Determines the time needed to communicate activations between expert/tensor parallel GPUs.
Scaling overall GenAI inference performance requires increasing the number of GPUs or accelerators working in parallel within the scale-up domain. Addressing these issues is critical and will require innovative changes in AI infrastructure — a shift toward light-based solutions — for the industry to achieve profitability.
Accelerating the Path to Artificial Intelligence Profitability for New Applications
Profitability is both a primary metric in AI applications and the end goal that every company (and investor) is vying for — with varying interactivity levels affecting the achievement of this goal. To better understand the path to artificial intelligence profitability, we have introduced three figures of merit for large-scale AI (Figure 2):
- Throughput: This metric is a key component of AI application profitability. It is defined as the number of users divided by time per output token.
- Interactivity: This is an indication of how responsive — and how interactive — an AI model can be when looking at different inference workloads that are running on top of it. It is defined as one divided by time per output token.
- Profitability: This metric is designed to provide an idea of the cost structure and the ability for applications to achieve headroom for profitability. It is defined as throughput divided by cost and power.
Figure 2: Figures of merit for large-scale AI.
By focusing on these metrics, we can better evaluate and enhance the profitability of AI applications in order to support artificial intelligence growth and innovation.
To accurately predict AI inference application-level performance, we have created a system architecture simulator (Figure 3) to arrive at these application-level figures-of-merit. This workload simulator takes into account a variety of inputs, including model specifications, technology components, model optimization, algorithm details, network fabric and cost. We’ve even taken the time and care to tune this simulator to a variety of models and GPUs on the market today and in the future.
Figure 3: Ayar Labs’ system architecture simulator.
Optical I/O: The Key to GenAI Profitability
A new approach to scale-up performance is needed to improve the unit economics of AI and the “tokenomics” of dollar per token per second. Optical I/O provides a clear solution and path to artificial intelligence profitability (Figures 4 and 5). How? By breaking the bandwidth distance bottleneck, optical I/O-based scale-up fabrics enhance application-level performance for both inference and training, optimizing TCO metrics, including throughput, interactivity, and profitability, leading to greater artificial intelligence growth.
Figure 4: Scale-up network with Ayar Labs’ optical I/O.
Figure 5: Theoretical roadmap of next-generation AI platforms.
Traditional electrical I/O requires cramming everything close together to drive scale-up performance, which creates two problems from the data center infrastructure point of view:
- The challenge of power delivery to the rack is that the data center infrastructure is not designed to deliver 120 kW per rack. (Typical CPU racks are 20 kW and accelerator racks are up to 40 kW)
- The challenge of managing power density inside the rack demands significant advancements in cooling solutions.
Optical I/O alleviates these pressures. By enabling the scale-up fabric to connect more GPUs or accelerators per cluster while spanning greater distances, optical I/O increases cluster performance without increasing the power per rack. In practical terms, you can deploy, for example, 16 or 32 GPUs or accelerators per rack and build more racks better tuned to the data center infrastructure, all connected through optical I/O. (Figure 6).
Figure 6: Ayar Labs’ optical I/O is ideally suited for AI scale-up fabrics.
Data center operators are evaluating these solutions to understand the value proposition better. The ultimate goal is to drive dramatic improvements in the “tokenomics” all the way up to the AI applications, ensuring that current investments in AI industry growth lead to significant economic output and artificial intelligence profitability. This includes profitability for data centers and users while also delivering clear economic value to customers by enhancing interactivity — the rate at which users and machines interact with AI models.
With optical I/O arriving as a key technology for AI profitability, we can expand the profitable domains of interactivity and connectivity by addressing the power and thermal density issues and breaking the bandwidth distance bottleneck.
Optical I/O’s Impact on GPT-4 Inference Workloads
Figure 7: Optical I/O can meaningfully improve artificial intelligence profitability by 6x and interactivity by 4x for batch and human-to-AI inference workloads of today’s GPT-4 models.
Our system architecture simulator has found that optical I/O can meaningfully improve profitability by 6x and interactivity by 4x for batch and human-to-AI inference workloads of today’s GPT-4 models (those with, say, 1.8 trillion parameters) (Figure 7). Here is a brief explanation of batch processing and human-to-AI interaction:
- Batch processing: When it comes to batch processing, or offline inferencing, machine learning models are run on large datasets that generate outputs on a batch of information. These batch runs are typically generated during some recurring schedule (e.g., hourly or daily). These inferences are then stored in a database or a file and can be made available to developers or end users. Examples of batch processing would include product review summaries on e-commerce websites or streaming movie recommendations based on viewing habits.
- Human-to-AI interaction: With human engagement, or real-time/online serving, instead of waiting hours or even days for predictions to be generated in batch, we can generate predictions as soon as they are needed — instantly serving those results to users. The goal of human-to-AI interactions is usually to optimize the number of transactions per second that the model can process. As you might imagine, examples include ChatGPT and any of the copilot product offerings we see from companies today.
Impact on Future GPT-X Models and Beyond
The benefits of optical I/O become even clearer when looking two or three years down the road, as GPT-X models with approximately 14 trillion parameters (an 8x model size increase, which is less than the projected trend of 16-25x every two years) become commonplace. (Figure 8)
Figure 8: GPT-4 model of today and estimate for GPT-X.
According to our system architecture simulator, optical I/O has the potential to increase artificial intelligence profitability by 20x while improving interactivity by 3-4x. In fact, optical I/O becomes the only viable technology for all use cases of GPT-X future models — and is the key to opening up the possibilities for future agentic inference models beyond today’s human-to-AI interactions. (Figure 9)
Figure 9: Optical I/O has the potential to increase profitability by 20x while improving interactivity by 3-4x for future GPT-X models.
Only optical I/O makes it possible for multiple AI copilots or agents to “talk” with one another in what is known as machine-to-machine multi-agentic interactions. Imagine, if you will, an engineer is working on a patent and leverages a co-pilot built to leverage the expertise from two agents. The first is focused on engineering and the second on law. The two agents work together to provide fast, relevant insights increasing the copilot’s value to the engineer on patent development. It’s hard to imagine based on today’s capabilities, but closer than one might think — with the help of optical I/O.
Although GenAI may be failing to meet both consumer and investor expectations, the path to artificial intelligence profitability is still possible. Optical I/O opens the floodgates for AI to improve the unit economics of AI applications, breaking the bandwidth-distance bottleneck arising from electrical I/O. This will allow AI systems and humans alike to reap the benefits of better profitability and interactivity while also enabling future multi-agentic inference models beyond today’s human-to-machine engagement.
Continue to learn more about optical I/O, our AI research and upcoming analysis by signing up for our monthly newsletter.
Related Resources:
Solution Overview | Ayar Labs’ Optical I/O Solution Improves AI Profitability by up to 20x
Solution Overview | AI Scale-Up Architecture with Optical I/O
Podcast | Ayar Labs’ Mark Wade on Optical I/O Boosting AI Performance
Video | CEO Mark Wade’s Presentation at the 2024 AI Hardware & Edge AI Summit
The Next Platform | Copper Wires Have Already Failed Clustered AI Systems
Tech Tech Potato | Solving AI’s Lumpy Problem











