Pricing

Routing that pays for itself.

Per-call pricing, the way an API gateway should work. No seat licenses, no per-GPU fees, no minimum commitment. At 1M calls per month on Growth, you are paying $0.000149 per routing decision. The GPU time you stop wasting costs considerably more than that.

Starter
Free
Up to 100K inference calls/month
  • Up to 3 node endpoints
  • Latency-budget routing
  • Basic cost-ceiling enforcement
  • Community support
  • 7-day telemetry history
Custom
Talk to us
Volume pricing and custom SLAs
  • Everything in Growth
  • Dedicated routing cluster option
  • Custom routing policy scripting
  • Priority support with SLA guarantee
  • On-prem deployment support
  • Quarterly architecture review

Before you ask your finance team.

One call to client.infer() counts as one inference call, regardless of the response length or the number of tokens generated. Streaming and non-streaming calls are counted the same way. Health check pings to your nodes do not count.

Yes. Proximarun is backend-agnostic. You can mix cloud-hosted nodes, on-prem edge nodes, and any OpenAI-compatible endpoint in the same node pool. On Starter you can define up to 3 nodes. Growth and Custom have no limit on the number of nodes in your pool.

The Starter tier is permanently free up to 100K calls per month. No trial period, no credit card required. At 100K calls, routing stops for the rest of the calendar month and resumes at the start of the next month. You can upgrade to Growth at any time to remove the cap.

Growth includes 5M calls per month. Above that, usage is billed at $0.30 per additional 100K calls. You will receive an email notification when you reach 80% of the included volume. There is no hard cutoff; routing continues without interruption.

No contract for Starter or Growth. Both are month-to-month. Cancel any time; your access continues through the paid period. Custom tier agreements typically include a 12-month commitment in exchange for volume discounts and dedicated infrastructure. Contact us to discuss options.

Start free. See the GPU savings first.

100K calls on Starter at no cost, no card required. Connect your node pool, run real traffic, and check the telemetry before you pay anything.