Use cases
One loop. Pointed at whatever costs you the most.
Evo profiles your workload on your hardware, proposes one change at a time, keeps only what clears the noise floor and the correctness gate, and returns a pull request with the measurement attached. The request changes by team. The loop does not. Each card below says what you ask for, what the result looks like, and how far along that lane is today.
All use cases
-
“Make our CI faster.”
CMake and Bazel, Cargo, Go, Python and TypeScript monorepos
Shorter clean and incremental builds and test runs, with every required check preserved. Runner-minutes are measured alongside wall time, so the saving shows up on the bill.
What counts build and test wall time on your runners, paired against the current configuration, output digests identical
In pilot -
“Handle more API traffic on these machines.”
Go, Rust, Java, C#, Python and Node services; SaaS and transaction systems
More requests per machine at the tail latency and error rate you already commit to. Fewer machines for the same traffic, or the same machines for the next doubling.
What counts requests per second at fixed p99 and error budget on a replayed traffic sample
In pilot -
“Continuously improve this model endpoint.”
PyTorch, CUDA, Triton and C++; inference providers and AI applications on vLLM or SGLang
- Triton
- vLLM
- SGLang
More quality-qualified requests per GPU-hour. The serving recipe (parallel layout, draft depth, KV precision, batch limits, attention backend) is searched against your SLO and your traffic mix, and the loop re-runs when the model, engine or traffic changes.
What counts dollars per million output tokens that meet the SLO, with a quality gate against a reference
Measured -
“Generate images, video and speech more cheaply.”
Diffusion, video and speech pipelines
- Diffusers
More accepted outputs per GPU-hour at fixed quality and latency. Quality is a gate the search cannot move, not a knob it trades away.
What counts outputs per GPU-hour above a fixed perceptual or task-quality threshold
Roadmap -
“Finish this data job before the deadline.”
Python, SQL, C++ and Rust; analytics, ingestion and document processing
A faster complete pipeline with matching outputs and lower memory and compute, measured end to end rather than stage by stage.
What counts pipeline wall time and peak memory on a frozen input sample, outputs byte-identical
In pilot -
“Build and search our index faster.”
Embedding, vector-search and recommendation pipelines
- FAISS
- HNSW
Faster indexing or faster queries at fixed recall and freshness. Recall is part of the correctness contract.
What counts queries per second at fixed recall@k, index build time at fixed recall
Roadmap -
“Encode our media backlog faster.”
C, C++ and SIMD codecs; media pipelines
- SIMD
Faster encoding at matched quality and bitrate. The codec lane is where the equivalence prover is strongest: a candidate is accepted only when the compiled code provably produces identical bytes.
What counts throughput on a frozen corpus, output bytes identical, proof attached where the transformation allows
Measured -
“Run more scientific experiments.”
C++, Fortran, CUDA and MPI; materials, physics and life sciences
- MPI
- OpenMP
Less time to the required numerical or physical result. Tolerances are stated by the scientist and enforced by the gate, so a faster run never means a different answer.
What counts time to a stated numerical tolerance on a reference case, across the cluster's real node
Roadmap -
“Simulate this design faster.”
Verilator, C++ simulation and EDA workflows
- Verilator
- SystemC
More completed simulation work per day while observable behaviour is preserved cycle for cycle.
What counts simulated cycles per second on the design's own regression, traces identical
Roadmap -
“Meet the robot or device deadline.”
C++ and Rust; embedded inference, control and planning on the actual target
Lower application latency, memory or energy, measured on the device rather than on a developer laptop. A runner on the target is the whole setup.
What counts loop latency, memory high-water mark and energy per cycle on the shipped hardware
In pilot -
“Keep these gains through upgrades.”
Any project above, once qualified
- Your CI
When the compiler, the framework, the model or the hardware changes, the optimized implementation is revalidated, repaired and released again. The ledger of every judged change is what makes that cheap.
What counts every past gain re-measured on the new baseline; regressions caught before release
Roadmap
Nothing in this area yet. Tell us what you would point it at.
What could you ship with more performance?
Bring the repository, the workload and the machine it has to run on. We bring the loop and the proof.