HeadlinesBriefing favicon HeadlinesBriefing.com

Router by Ramp: Cut AI Inference Costs 40%

Hacker News •
×

Router by Ramp is a unified AI routing platform that cuts inference costs by 40% on average by intelligently matching each request to the lowest-cost model that meets performance needs. Built from years of internal use at Ramp, the service routes traffic across models from OpenAI, Anthropic, and open-source providers like Kimi through a single endpoint.

Ramp’s engineering teams used Router to reduce their own AI spend by 30% without sacrificing latency or solve rates. The system dynamically responds to live latency and failure rates, automatically shifting eligible requests to more cost-efficient tiers when quality isn’t impacted.

Router is free through 2026, with the first $26 in model credits on us. Developers get one endpoint, one bill, and zero lock-in. Setup takes two lines of code and works with existing OpenAI and Anthropic SDKs. Router Strategies let teams define custom cost and performance priorities or start with Ramp’s benchmarked defaults.

The service supports over 25 models, handles trillions of tokens monthly, and includes fallback routing when providers are unavailable. Enterprise features are coming soon, but individual developers and teams in the U.S. can start today—no Ramp card or company account required.