The AI industry has a straightforward response to growing demand: buy more compute. Before another dollar goes to hardware, it is worth asking how much performance is still sitting inside the machines the industry already owns. The companies with the largest fleets have measured it, and the answer is: a lot.

The spending

For 2026, Amazon now expects roughly $220B of capital expenditures, Alphabet expects $195–205B, Microsoft expects roughly $175B, and Meta expects $130–145B, each per its own guidance from July 2026. That is approximately $720–745B of 2026 capital spending across four companies alone. Not all of it is AI hardware, but all four companies explicitly tie the surge to AI and technical infrastructure.

Sources: Amazon · Alphabet · Microsoft · Meta

The obvious next question is: before spending another dollar on hardware, how effectively are we using the hardware we already have?

The answer is: often, not very well. Here is what that looks like in one pull request.

What this looks like in practice

esp-dl is Espressif’s deep-learning library for ESP32 microcontrollers. Every operator in it holds its selected kernel in one library-wide handle, and that handle was a std::function.

A std::function earns its cost when a handle might hold a lambda with captured state. Nothing in esp-dl ever stored anything in it but a plain function pointer: 139 declaration sites, 933 assignments, not one exception.

The handle is called inside the innermost loop of convolution, so every output position paid for an indirect call through the std::function invoker, plus a 32-byte copy, to reach a function whose address never changed. Our change replaces it with a thin wrapper over the function pointer. No call site changes. One file, 27 lines added, one removed.

esp-dl pull request #331 · A/B on the device

Baseline: std::function handlePatched: function-pointer wrapper
End-to-end detect41,684 → 39,546 µs
−5.1%
Conv module37,029 → 35,001 µs
−5.5%
PRelu module907 → 800 µs
−11.8%
Concat moduledoes not use the handle
±0 · control
App image2,331,552 → 2,314,400 B
−17 KB
Measured on the deviceESP32-S3 at 240 MHz, ESP-IDF v5.5, human-face-detect example, performance optimization build. Three runs per arm, A-B-A-B, run-to-run spread about 1 µs. Bars are indexed to each row’s baseline. Detection output identical on every run of every build.

The setup is one line: an ESP32-S3 at 240 MHz, ESP-IDF v5.5, the library’s human-face-detect example, performance optimization build. Every module that goes through the handle got faster. Concat, which does not, did not move, which is the control we wanted. The image is 17 KB smaller and the output was identical on every run of every build.

The limits are the ones the pull request states: one example, one board, one IDF version. The upstream pull request was merged on September 9, 2026.

It is a small change. That is the point. Nobody wrote a faster kernel and nothing about the chip changed. Every operator that goes through the handle got faster on every target, with a smaller binary and the same output. The performance was already in the machine.

The performance is already in the machine

That pattern is not specific to microcontrollers. A 2024 Microsoft Research study examined 400 real deep-learning jobs on Microsoft’s internal platform, all already identified as averaging 50% GPU utilization or less. Researchers found 706 distinct low-utilization problems, and 84.99% of them could be addressed with a small number of code or script changes.

Source: Microsoft Research, An Empirical Study on Low GPU Utilization of Deep Learning Jobs, ICSE 2024

The scope matters. Microsoft deliberately studied jobs already known to be underutilizing their GPUs, so this is not 85% of the fleet. It is something more useful: where expensive hardware was being used badly, most of the diagnosed problems were fixable in software.

Alibaba found the same problem at fleet scale. Its OSDI 2026 study analyzes a six-month production trace covering 155,410 GPUs across 37,707 servers, and its central finding is unusually direct: high GPU demand does not yield high effective utilization. Scheduling changes raised Alibaba’s GPU allocation ratio from 68% to 93%.

Source: Alibaba, Heterogeneity at Hyperscale, USENIX OSDI 2026

Underutilization takes different forms. Sometimes a GPU is waiting for work; sometimes the workload does not exercise the chip efficiently, which is what the esp-dl handle was doing at a much smaller scale. The economic consequence is the same: companies buy expensive compute that software does not fully convert into useful work.

A small software improvement is worth a lot of hardware

Suppose a company can make a fixed workload run 10% faster without changing the hardware. The same machines now do roughly 10% more work per unit time. That is economically similar to adding 10% more compute capacity without fabricating another chip or building another rack.

The comparison is only illustrative. Hyperscaler capital expenditure includes buildings, networking, storage, and power, so 10% better software does not literally save 10% of CapEx. But it puts the scale in perspective: against $720–745B of 2026 spending, even a single-digit improvement in the effective output of existing compute is worth an enormous amount of capacity. And measured software gains can be much larger than single digits.

A peer-reviewed HPDC 2026 paper called CARBS automatically tuned compiler decisions across 133 benchmarks totaling nearly 5 million lines of C/C++-heavy code, and reported an average speedup of 18.7% over -O3 on its traditional benchmarks and 10% on SPECspeed CPU 2017. These are benchmark results, not a promise that every application has 18.7% waiting. What matters is that the programs were already at the compiler’s aggressive standard setting, and there was still meaningful performance left.

Source: CARBS, ACM HPDC 2026

At the algorithmic level the gains can be larger still. Google DeepMind’s AlphaTensor used hardware performance as its feedback signal and found matrix-multiplication algorithms DeepMind reported were 10–20% faster than commonly used ones on the same NVIDIA V100 and Google TPU v2 hardware.

Source: Google DeepMind, AlphaTensor

FlashAttention is the clearest modern example of the upside. A 2024 analysis cites roughly six times the performance of native PyTorch for FlashAttention, achieved by reorganizing how data moved through the GPU. The hardware had not become six times faster. Engineers found a dramatically better way to use it.

Source: FlashAttention on a Napkin, Abbott and Zardini

Headroom the industry has measured

84.99%of diagnosed low-utilization problems fixable with a few code or script changesMicrosoft Research, ICSE 2024
68% → 93%GPU allocation ratio after scheduling changes across 155,410 GPUsAlibaba, USENIX OSDI 2026
18.7%average speedup over -O3 on traditional benchmarks, 133 programs tuned in allCARBS, ACM HPDC 2026
10–20%quicker matrix multiplication on the same V100 and TPU v2 hardwareGoogle DeepMind, AlphaTensor
≈6×performance versus native PyTorch, as cited by the paperFlashAttention on a Napkin, Abbott and Zardini
Five measured results, five sourcesEach figure is the number its source reported, in the scope its source studied. None of them is a promise about any other workload.

So why isn’t everyone doing this already?

There is already a solution: performance engineers

The fastest software in the world is not produced by pressing a better compiler button. It is produced by experts who understand both the program and the hardware.

They profile the workload. They restructure computations, fuse operations, change memory layouts, and choose specialized kernels and libraries. They write custom kernels when nothing fits. They benchmark alternatives on the real machine and repeat until the hardware is being used well.

The industry has spent decades building extraordinary tools for these engineers: NVIDIA’s cuBLAS, cuDNN, and CUTLASS; OpenAI’s Triton, now a key backend for PyTorch compilation; TorchInductor, which turns model graphs into optimized Triton and C++; Google’s XLA and OpenXLA; AWS Neuron for Trainium and Inferentia; Arm’s CMSIS-NN for microcontrollers. Nearly every specialized hardware company, from AMD and Qualcomm to Cerebras and Groq, maintains its own compiler, runtime, or kernel stack.

These libraries and the experts behind them are not what Velobyte replaces. They are what makes Velobyte possible.

The problem is that connecting a program to the best combination of those tools is still largely human work. In esp-dl the kernels were already excellent; what was missing was someone with the time to look at how the program reached them.

Why an agent cannot just “make the code faster”

Agents are increasingly capable of writing kernels and suggesting optimizations. That is necessary, but not sufficient. An unconstrained coding agent has five fundamental problems.

It does not know which implementation is actually fastest. A model can generate plausible code, but performance is often counterintuitive. The only reliable judge is the real target hardware.

We learned this the direct way. An earlier, agent-assisted rewrite of the arithmetic decoder in Google’s liblc3 codec replaced a chain of dependent multiplies on the critical path with one division. On an Apple M2 Pro it cut median decoder time by 3–5.6% with byte-identical output, and the integer identity behind it was machine-checked. The maintainers pointed out that it would regress on microcontrollers without a hardware divider, where liblc3 is most heavily deployed. They are right: a win on one machine is not a win on every machine, which is why the loop has to measure on the real target. That pull request remains open.

It does not naturally know the entire backend ecosystem. The fastest answer may already exist inside CUTLASS, cuBLAS, Triton, CMSIS-NN, or a vendor SDK. The goal is not to regenerate decades of expert work but to expose it to the agent as part of the search.

The space of possible changes is too large. There can be thousands of ways to restructure a program, combine operations, select libraries, change layouts, or specialize code for hardware. Without structure, an agent is guessing in an enormous search space.

Faster code can silently be wrong. Performance transformations can change floating-point behavior, overflow, memory semantics, or edge cases that ordinary tests miss. An autonomous optimizer needs a correctness gate, not “the agent says it looks equivalent.”

The agent does not learn a compiler. Even if an agent finds a great optimization once, a normal coding session does not turn that discovery into reusable compiler knowledge. Without a framework, the next program starts from nearly zero.

Five problems

Does not know what is fastestPlausible code, counterintuitive performance.
Does not know the backend ecosystemThe best answer may already be in a library.
Search space too largeThousands of restructurings, no structure.
Faster code can be wrongOverflow, rounding, aliasing, edge cases.
Nothing is retainedThe next program starts from zero.

The loop

What the infrastructure supplies
  • hardware decides
  • libraries exposed
  • search structured
  • changes checked
  • results kept

What it takes

Measure on the real targetCompile, run, benchmark, compare, feed back.
Expose the expert librariesDecades of vendor work become options.
Structure the searchLegal choices derived from the program and target.
A correctness gateBehavior checked, not asserted.
Keep what worksA win becomes reusable knowledge.
Five problems, five requirementsEach thing an unconstrained agent lacks maps to one piece of infrastructure the loop has to supply.

This is the missing infrastructure.

What Velobyte builds

Velobyte is not building a new cuBLAS. It is not trying to write a better CUTLASS than NVIDIA, or to replace Triton, XLA, PyTorch, GCC, LLVM, or vendor SDKs. It is building the automation layer above them. Our product, EvoCompiler, gives coding agents that layer. At a high level:

Understand the program. Identify a region of computation and what it is supposed to do.

Expose the ways to execute it. Compiler paths, expert libraries, target-specific primitives, and backend implementations become options.

Let agents expand the possibilities. New mappings, a reorganization that makes an existing library usable, a composition of primitives, or a new implementation when nothing fits.

Check the transformation. Before an optimization is kept, establish that it preserves the intended behavior. The agent’s confidence is not evidence.

Measure on the real target. Compile it, run it, benchmark it, and let the hardware, not the agent, decide.

Keep what works. Successful transformations become reusable knowledge rather than one-off patches.

01 program

Understand

A region and what it is supposed to do.

02 options

Expose

Compilers, libraries, primitives, kernels.

03 agents

Expand

New ways to reach a library, or new code.

04 behavior

Check

Intended behavior preserved, or the candidate is out.

05 hardware

Measure

On the real target. The chip decides.

06 knowledge

Keep

What works becomes reusable.

kept results seed the next search

The loopunderstand → expose → expand → check → measure → keep. The agent proposes; the hardware and the checks decide.

The loop is: understand → expose → expand → check → measure → keep. That is what turns a coding agent into a performance engineer.

Where this is going

This is not an empty market waiting for someone to discover the problem. A large ecosystem is converging on pieces of the solution: autotuning like Apache TVM, Ansor, and Halide; production compilers like OpenXLA, TorchInductor, TensorRT, IREE, and MLIR; equivalence-checking research like Alive2, CompCert, and Souper; agent-driven kernel work like KernelBench and AWS Neuron Agentic Development. Luminal is the closest architectural cousin: it reduces networks to a small set of primitives and searches across implementation decisions instead of relying on fixed heuristics. Velobyte’s difference is that the search space is not meant to stay fixed. The existing libraries are the starting point; agents keep finding new ways to connect programs to them.

Source: Luminal on GitHub

The market is assigning real strategic value to this layer. Qualcomm completed its acquisition of the compiler company Modular in July 2026 at approximately $3.1B, per Qualcomm’s SEC filing. NVIDIA acquired OctoAI, built by the TVM compiler’s core contributors, in 2024, and agreed to acquire the AI optimization company Deci the same year. Red Hat completed its acquisition of Neural Magic in January 2025.

Sources: Qualcomm · Qualcomm SEC filing · NVIDIA and OctoAI · NVIDIA and Deci · Red Hat and Neural Magic

Anthropic shows what the solution looks like when money is no object. It says its engineers have worked directly with AWS’s Annapurna Labs to write low-level kernels against Trainium silicon and contribute to the AWS Neuron stack, and Anthropic’s head of compute, quoted on AWS’s Trainium page, says almost a million Trainium2 chips are training and serving Claude through Project Rainier.

Sources: Anthropic on the AWS and Trainium partnership · AWS on Trainium customers

That is the solution today: build specialized hardware, assemble elite compiler and kernel engineers, hand-connect models to the best hardware primitives, profile, tune, repeat. It works. It just does not scale to every company, every workload, every chip, and every new generation of hardware.

AWS is now making that point itself. In June 2026 it published a post titled “Stop hand-tuning kernels,” announcing its Neuron Agentic Development capabilities and describing custom kernel development as the historical path to closing the gap between theoretical and achieved performance, and one that requires architectural knowledge, manual profiling, and iteration cycles few teams can afford. Its agents write NKI kernels, debug them on Trainium, and analyze hardware profiles; AWS has since added agentic model porting and numerical-equivalence validation.

Sources: AWS, “Stop hand-tuning kernels,” June 2026 · AWS Neuron Agentic Development · Neuron 2.30 equivalence validation

AWS is beginning to automate expert performance engineering inside the Trainium ecosystem. The natural next question is: why can’t an agent do this everywhere?

The thesis

The industry is racing to manufacture more compute. At the same time, Microsoft, Alibaba, academic compiler research, and the most sophisticated AI companies in the world all show the same thing: there is substantial performance available in software.

We already know how to recover it. We employ expert engineers, build specialized libraries, profile the hardware, search through implementations, hand-write kernels, and continuously tune the mapping between programs and machines. That approach is powerful, and it does not scale.

Agents can now participate in the work. An agent alone is missing the infrastructure required to search intelligently, use the best existing libraries, check that aggressive changes are safe, benchmark them on the real target, and retain what it learns. Velobyte is building that infrastructure. We are not replacing the world’s performance engineers or the libraries they built. We are turning their work into a search space that agents can use, extend, check, and optimize for every program and every machine.

The performance is already in the machine. We make your code faster.