The AI industry has a straightforward response to growing demand: buy more compute. Before another dollar goes to hardware, it is worth asking how much performance is still sitting inside the machines the industry already owns. The companies with the largest fleets have measured it, and the answer is: a lot.
The spending
For 2026, Amazon now expects roughly $220B of capital expenditures, Alphabet expects $195–205B, Microsoft expects roughly $175B, and Meta expects $130–145B, each per its own guidance from July 2026. That is approximately $720–745B of 2026 capital spending across four companies alone. Not all of it is AI hardware, but all four companies explicitly tie the surge to AI and technical infrastructure.
Sources: Amazon · Alphabet · Microsoft · Meta
The obvious next question is: before spending another dollar on hardware, how effectively are we using the hardware we already have?
The answer is: often, not very well. Here is what that looks like in one pull request.
What this looks like in practice
esp-dl is Espressif’s deep-learning library for ESP32 microcontrollers. Every operator in it holds its selected kernel in one library-wide handle, and that handle was a std::function.
A std::function earns its cost when a handle might hold a lambda with captured state. Nothing in esp-dl ever stored anything in it but a plain function pointer: 139 declaration sites, 933 assignments, not one exception.
The handle is called inside the innermost loop of convolution, so every output position paid for an indirect call through the std::function invoker, plus a 32-byte copy, to reach a function whose address never changed. Our change replaces it with a thin wrapper over the function pointer. No call site changes. One file, 27 lines added, one removed.
esp-dl pull request #331 · A/B on the device
The setup is one line: an ESP32-S3 at 240 MHz, ESP-IDF v5.5, the library’s human-face-detect example, performance optimization build. Every module that goes through the handle got faster. Concat, which does not, did not move, which is the control we wanted. The image is 17 KB smaller and the output was identical on every run of every build.
The limits are the ones the pull request states: one example, one board, one IDF version. The upstream pull request was merged on September 9, 2026.
It is a small change. That is the point. Nobody wrote a faster kernel and nothing about the chip changed. Every operator that goes through the handle got faster on every target, with a smaller binary and the same output. The performance was already in the machine.
The performance is already in the machine
That pattern is not specific to microcontrollers. A 2024 Microsoft Research study examined 400 real deep-learning jobs on Microsoft’s internal platform, all already identified as averaging 50% GPU utilization or less. Researchers found 706 distinct low-utilization problems, and 84.99% of them could be addressed with a small number of code or script changes.
Source: Microsoft Research, An Empirical Study on Low GPU Utilization of Deep Learning Jobs, ICSE 2024
The scope matters. Microsoft deliberately studied jobs already known to be underutilizing their GPUs, so this is not 85% of the fleet. It is something more useful: where expensive hardware was being used badly, most of the diagnosed problems were fixable in software.
Alibaba found the same problem at fleet scale. Its OSDI 2026 study analyzes a six-month production trace covering 155,410 GPUs across 37,707 servers, and its central finding is unusually direct: high GPU demand does not yield high effective utilization. Scheduling changes raised Alibaba’s GPU allocation ratio from 68% to 93%.
Source: Alibaba, Heterogeneity at Hyperscale, USENIX OSDI 2026
Underutilization takes different forms. Sometimes a GPU is waiting for work; sometimes the workload does not exercise the chip efficiently, which is what the esp-dl handle was doing at a much smaller scale. The economic consequence is the same: companies buy expensive compute that software does not fully convert into useful work.
A small software improvement is worth a lot of hardware
Suppose a company can make a fixed workload run 10% faster without changing the hardware. The same machines now do roughly 10% more work per unit time. That is economically similar to adding 10% more compute capacity without fabricating another chip or building another rack.
The comparison is only illustrative. Hyperscaler capital expenditure includes buildings, networking, storage, and power, so 10% better software does not literally save 10% of CapEx. But it puts the scale in perspective: against $720–745B of 2026 spending, even a single-digit improvement in the effective output of existing compute is worth an enormous amount of capacity. And measured software gains can be much larger than single digits.
A peer-reviewed HPDC 2026 paper called CARBS automatically tuned compiler decisions across 133 benchmarks totaling nearly 5 million lines of C/C++-heavy code, and reported an average speedup of 18.7% over -O3 on its traditional benchmarks and 10% on SPECspeed CPU 2017. These are benchmark results, not a promise that every application has 18.7% waiting. What matters is that the programs were already at the compiler’s aggressive standard setting, and there was still meaningful performance left.
Source: CARBS, ACM HPDC 2026
At the algorithmic level the gains can be larger still. Google DeepMind’s AlphaTensor used hardware performance as its feedback signal and found matrix-multiplication algorithms DeepMind reported were 10–20% faster than commonly used ones on the same NVIDIA V100 and Google TPU v2 hardware.
Source: Google DeepMind, AlphaTensor
FlashAttention is the clearest modern example of the upside. A 2024 analysis cites roughly six times the performance of native PyTorch for FlashAttention, achieved by reorganizing how data moved through the GPU. The hardware had not become six times faster. Engineers found a dramatically better way to use it.
Source: FlashAttention on a Napkin, Abbott and Zardini
Headroom the industry has measured
So why isn’t everyone doing this already?
There is already a solution: performance engineers
The fastest software in the world is not produced by pressing a better compiler button. It is produced by experts who understand both the program and the hardware.
They profile the workload. They restructure computations, fuse operations, change memory layouts, and choose specialized kernels and libraries. They write custom kernels when nothing fits. They benchmark alternatives on the real machine and repeat until the hardware is being used well.
The industry has spent decades building extraordinary tools for these engineers: NVIDIA’s cuBLAS, cuDNN, and CUTLASS; OpenAI’s Triton, now a key backend for PyTorch compilation; TorchInductor, which turns model graphs into optimized Triton and C++; Google’s XLA and OpenXLA; AWS Neuron for Trainium and Inferentia; Arm’s CMSIS-NN for microcontrollers. Nearly every specialized hardware company, from AMD and Qualcomm to Cerebras and Groq, maintains its own compiler, runtime, or kernel stack.
These libraries and the experts behind them are not what Velobyte replaces. They are what makes Velobyte possible.
The problem is that connecting a program to the best combination of those tools is still largely human work. In esp-dl the kernels were already excellent; what was missing was someone with the time to look at how the program reached them.
Why an agent cannot just “make the code faster”
Agents are increasingly capable of writing kernels and suggesting optimizations. That is necessary, but not sufficient. An unconstrained coding agent has five fundamental problems.
It does not know which implementation is actually fastest. A model can generate plausible code, but performance is often counterintuitive. The only reliable judge is the real target hardware.
We learned this the direct way. An earlier, agent-assisted rewrite of the arithmetic decoder in Google’s liblc3 codec replaced a chain of dependent multiplies on the critical path with one division. On an Apple M2 Pro it cut median decoder time by 3–5.6% with byte-identical output, and the integer identity behind it was machine-checked. The maintainers pointed out that it would regress on microcontrollers without a hardware divider, where liblc3 is most heavily deployed. They are right: a win on one machine is not a win on every machine, which is why the loop has to measure on the real target. That pull request remains open.
It does not naturally know the entire backend ecosystem. The fastest answer may already exist inside CUTLASS, cuBLAS, Triton, CMSIS-NN, or a vendor SDK. The goal is not to regenerate decades of expert work but to expose it to the agent as part of the search.
The space of possible changes is too large. There can be thousands of ways to restructure a program, combine operations, select libraries, change layouts, or specialize code for hardware. Without structure, an agent is guessing in an enormous search space.
Faster code can silently be wrong. Performance transformations can change floating-point behavior, overflow, memory semantics, or edge cases that ordinary tests miss. An autonomous optimizer needs a correctness gate, not “the agent says it looks equivalent.”
The agent does not learn a compiler. Even if an agent finds a great optimization once, a normal coding session does not turn that discovery into reusable compiler knowledge. Without a framework, the next program starts from nearly zero.
Five problems
The loop
What the infrastructure supplies- hardware decides
- libraries exposed
- search structured
- changes checked
- results kept
What it takes
This is the missing infrastructure.
What Velobyte builds
Velobyte is not building a new cuBLAS. It is not trying to write a better CUTLASS than NVIDIA, or to replace Triton, XLA, PyTorch, GCC, LLVM, or vendor SDKs. It is building the automation layer above them. Our product, EvoCompiler, gives coding agents that layer. At a high level:
Understand the program. Identify a region of computation and what it is supposed to do.
Expose the ways to execute it. Compiler paths, expert libraries, target-specific primitives, and backend implementations become options.
Let agents expand the possibilities. New mappings, a reorganization that makes an existing library usable, a composition of primitives, or a new implementation when nothing fits.
Check the transformation. Before an optimization is kept, establish that it preserves the intended behavior. The agent’s confidence is not evidence.
Measure on the real target. Compile it, run it, benchmark it, and let the hardware, not the agent, decide.
Keep what works. Successful transformations become reusable knowledge rather than one-off patches.
01 program
UnderstandA region and what it is supposed to do.
02 options
ExposeCompilers, libraries, primitives, kernels.
03 agents
ExpandNew ways to reach a library, or new code.
04 behavior
CheckIntended behavior preserved, or the candidate is out.
05 hardware
MeasureOn the real target. The chip decides.
06 knowledge
KeepWhat works becomes reusable.
kept results seed the next search
The loop is: understand → expose → expand → check → measure → keep. That is what turns a coding agent into a performance engineer.
Where this is going
This is not an empty market waiting for someone to discover the problem. A large ecosystem is converging on pieces of the solution: autotuning like Apache TVM, Ansor, and Halide; production compilers like OpenXLA, TorchInductor, TensorRT, IREE, and MLIR; equivalence-checking research like Alive2, CompCert, and Souper; agent-driven kernel work like KernelBench and AWS Neuron Agentic Development. Luminal is the closest architectural cousin: it reduces networks to a small set of primitives and searches across implementation decisions instead of relying on fixed heuristics. Velobyte’s difference is that the search space is not meant to stay fixed. The existing libraries are the starting point; agents keep finding new ways to connect programs to them.
Source: Luminal on GitHub
The market is assigning real strategic value to this layer. Qualcomm completed its acquisition of the compiler company Modular in July 2026 at approximately $3.1B, per Qualcomm’s SEC filing. NVIDIA acquired OctoAI, built by the TVM compiler’s core contributors, in 2024, and agreed to acquire the AI optimization company Deci the same year. Red Hat completed its acquisition of Neural Magic in January 2025.
Sources: Qualcomm · Qualcomm SEC filing · NVIDIA and OctoAI · NVIDIA and Deci · Red Hat and Neural Magic
Anthropic shows what the solution looks like when money is no object. It says its engineers have worked directly with AWS’s Annapurna Labs to write low-level kernels against Trainium silicon and contribute to the AWS Neuron stack, and Anthropic’s head of compute, quoted on AWS’s Trainium page, says almost a million Trainium2 chips are training and serving Claude through Project Rainier.
Sources: Anthropic on the AWS and Trainium partnership · AWS on Trainium customers
That is the solution today: build specialized hardware, assemble elite compiler and kernel engineers, hand-connect models to the best hardware primitives, profile, tune, repeat. It works. It just does not scale to every company, every workload, every chip, and every new generation of hardware.
AWS is now making that point itself. In June 2026 it published a post titled “Stop hand-tuning kernels,” announcing its Neuron Agentic Development capabilities and describing custom kernel development as the historical path to closing the gap between theoretical and achieved performance, and one that requires architectural knowledge, manual profiling, and iteration cycles few teams can afford. Its agents write NKI kernels, debug them on Trainium, and analyze hardware profiles; AWS has since added agentic model porting and numerical-equivalence validation.
Sources: AWS, “Stop hand-tuning kernels,” June 2026 · AWS Neuron Agentic Development · Neuron 2.30 equivalence validation
AWS is beginning to automate expert performance engineering inside the Trainium ecosystem. The natural next question is: why can’t an agent do this everywhere?
The thesis
The industry is racing to manufacture more compute. At the same time, Microsoft, Alibaba, academic compiler research, and the most sophisticated AI companies in the world all show the same thing: there is substantial performance available in software.
We already know how to recover it. We employ expert engineers, build specialized libraries, profile the hardware, search through implementations, hand-write kernels, and continuously tune the mapping between programs and machines. That approach is powerful, and it does not scale.
Agents can now participate in the work. An agent alone is missing the infrastructure required to search intelligently, use the best existing libraries, check that aggressive changes are safe, benchmark them on the real target, and retain what it learns. Velobyte is building that infrastructure. We are not replacing the world’s performance engineers or the libraries they built. We are turning their work into a search space that agents can use, extend, check, and optimize for every program and every machine.
The performance is already in the machine. We make your code faster.