Custom inference for every hardware stack.

Evo is your AI performance engineer, continuously evolving and specializing new models and architectures to your stack. The more custom the stack, the bigger the gains.

  1. Modelthe architecture and weights you serve
  2. Workloadyour requests, batch sizes and latency target
  3. Hardwareyour machine, in the loop for every measurement
  4. Implementationdiscovered, measured and proved equivalent

The AI performance engineer

Every new model and architecture, specialized to your stack.

New models and architectures arrive before a team has tuned the last one by hand. Evo automatically discovers a faster implementation of each one and mathematically proves it equivalent to the original, getting you more performance from the hardware you already have.

  • Continuous, model after model.

    When a new model, architecture or workload lands, Evo goes again, starting from what it already learned about your stack. Performance keeps up with what you ship.

  • Built for your machine. The more custom, the better.

    Generic runtimes and kernels are written to work everywhere. Evo, an agent-driven compiler, builds new implementations where they fall short for your stack and mathematically proves each one equivalent.

  • Expert-level performance, automatically.

    Profiling, proposing, checking and measuring is the work of an expert performance engineer. Evo does that work for you and returns an implementation you can inspect, proved equivalent to the original under the preconditions you specify.

Where Evo works

Performance lives across the stack.

  1. 01

    Application / model

    Workload semantics

    • model architecture
    • attention / MoE structure
    • quantization
    • sparsity
  2. 02

    Algorithms / compiler

    Transformation and scheduling core

    • operator decomposition
    • fusion
    • tiling
    • layouts
    • precision
    • instruction selection
  3. 03

    Runtime

    Execution and data movement

    • batching
    • KV cache
    • expert routing
    • parallelism
    • communication
  4. 04

    ISA / dataflow

    The contract with the hardware core

    • tensor / matrix units
    • SIMD
    • custom instructions
    • supported datatypes
    • dataflow primitives
  5. 05

    Microarchitecture

    Compute and memory organization

    • memory hierarchy
    • bandwidth
    • SRAM / cache behavior
    • interconnect
    • scheduling constraints
  6. 06

    RTL

    Implementation

    • fixed-function blocks
    • compute organization
  7. 07

    Silicon

    Power, performance, area

    • memory capacity
    • bandwidth, area and power constraints

Core marks the compiler and the hardware contract it targets. The three layers beneath are the chip itself, and Evo does not need their internals. It measures with the hardware in the loop, and optimizes from that data and everything known about the chip.

Performance comes from how these layers meet. Evo optimizes across them with the hardware in the loop, instead of tuning one in isolation.

How Evo works

The fastest implementation is discovered, not guessed.

  • One model, many equivalent implementations.

    A model fixes what has to be computed, not how. The algorithm, kernels, layout, precision and schedule underneath are free to change.

  • Most are slower. Some are wrong.

    Every candidate is measured with your hardware in the loop, and the ones Evo keeps are mathematically proved equivalent under the preconditions you specify.

  • Evo keeps what survives.

    Search can be probabilistic. What gets deployed is deterministic and checkable: an ordinary implementation you can inspect, build and run with no agent in the loop.

Measured so far, on the target, against stock and tuned baselines

  • 1.48×

    Qwen3 on one H100: 48% ahead of stock SGLang and 12.9% ahead of vLLM + DFlash.

  • 3×

    A custom-silicon audio compression workload, measured on the hardware it ships on.

  • 1.3×

    1-line change to an embedded deep-learning library that experts had already optimized. Merged upstream.

Formal validation

Faster, and mathematically proved equivalent.

A model's outputs are the contract; the algorithm, kernels and schedule that produce them are free to change. For supported transformations, a solver proves that each change Evo keeps is equivalent to the original, for every input that meets the preconditions you specify.

  • The agent decomposes the change.

    For supported transformations, the original and the candidate are split into matching pieces.

  • It proposes why each piece is equal.

    Each claim is stated in full: what is equal, and under which of the preconditions you specify.

  • A solver checks they compose.

    The proof holds under the preconditions you specify and assumes nothing else. Unsupported or inconclusive cases are reported as such.

Example optimization record

change
conv2d dispatch · candidate 2 of 3
preconditions
n_ch ≤ 512 · stride ∈ {1, 2} · buffers do not alias
proof
mathematically proved equivalent under the preconditions you specify
measured
2.1× end to end on the target · 33.3 → 15.9 ms per frame · identical output on 200 frames
kept
frozen as a reusable pass; the build replays it

Inference is where we are applying this first. Read: the performance is already in the machine

A decade at the hardware‑software boundary.

Two founders, friends since fourth grade. Both dropped out to do this full-time. Backed by Y Combinator.

  • Kaden Cassidy

    Kaden Cassidy

    Top 20 in the Anthropic kernel optimization challenge. 600+ certification tests shipping in RISC-V International's ACT4 suite.

  • Parth Kocheta

    Parth Kocheta

    Sped up petabyte-scale robotics infrastructure at Amazon. Visual-inertial odometry for autonomous drones at CMU's AirLab.

More about the team

What could you ship with more performance?

A new model tuned for the stack you already run. A larger model on the same accelerator. More tokens per second at the latency you promise. Bring us the model and the workload. Evo gets more performance from the hardware you already have, proves each change equivalent under the preconditions you specify, and goes again when the next model lands.