We built this because inference routing was broken.
Most teams handle inference routing with a hardcoded endpoint, a retry on 503, and a prayer that the traffic stays flat. It works until it doesn't. We spent two years on the wrong side of that equation, then built the layer we needed.
From GPU dashboards at 3am to a product engineers actually use.
Erik Johansson spent 2020 to 2022 as a platform engineer at a computer vision startup, responsible for keeping their inference stack available and within budget. The team was running vLLM on a small cluster of A10 nodes. Routing was a 200-line Python script that pointed traffic at node-01 by default and fell over to node-02 when node-01 returned a 503. That was it.
The problem was not that the script was buggy. It ran fine at median traffic. The problem was that it had no concept of load, latency, or cost. Off-peak, 60% of the cluster sat idle. During batch jobs that ran overnight, the latency-sensitive user-facing workloads were competing for the same nodes with no awareness of each other. The GPU bill came in every month and nobody could explain it from first principles.
In late 2022, Erik and his co-founder Lena Mwangi started sketching a routing layer that treated latency and cost as first-class inputs rather than afterthoughts. Proximarun, Inc. was incorporated in Salt Lake City in early 2023. We have been bootstrapped from day one and plan to stay that way.
We do not sell GPU capacity. We are not a cloud provider. Proximarun is a routing layer that sits in front of the inference backends you already have and makes sure each request reaches the node that can actually serve it at your declared target. That is a narrow problem. We think narrow is good.
Erik Johansson
CEO & Co-Founder
Three principles that shape every product decision.
Specificity over abstraction
Routing decisions need concrete targets, not vibes. "Route to the best node" is not a routing policy. "Route to the node whose rolling p95 is under 400ms and whose cost per token is under $0.002" is. Proximarun forces you to declare your SLA targets explicitly, because explicit constraints produce predictable behavior and debuggable failures. We do not let you skip this step.
Latency is a feature
We treat latency as a first-class product requirement, not a footnote. The routing engine adds overhead. That overhead competes with your latency budget. So we measure it on every release and we hold the p99 of the routing decision itself to under 12ms. If we ship a change that moves that number, we roll it back. Latency is not a tradeoff we make to ship faster.
Honest defaults
Proximarun's default config should be safe and cost-aware, not just fast. An out-of-the-box configuration that maximizes throughput at the expense of unpredictable billing is not a good default. We'd rather ship sensible cost ceilings by default and let you loosen them than watch teams get surprised by their first invoice.
Questions, feedback, or just want to talk inference routing?
Proximarun, Inc.
- 222 South Main Street, Suite 1800, Salt Lake City, UT 84101
- [email protected]
- +1 (801) 463-7218
- Founded 2023 · Bootstrapped