FlexInference
An inference routing API that finds cheaper model inference within a user-defined time window to reduce cost
The product is an inference routing API that reduces model inference cost by finding cheaper providers within a user-defined time window. It is for developers and data teams in business environments building applications on OpenAI, Anthropic, and Gemini models. It is delivered as a hosted API and infrastructure service where clients point their existing SDK base URL to the router and add a start_within parameter, with bring-your-own-key or managed-key options.
Key features
- Inference routing to cheaper providers
- Time-window based flex race
- Edge compute via 300+ cities
- 1-5ms routing overhead
- OpenAI/Anthropic/Gemini client compatibility
- Route via Bedrock and Vertex AI
- Bring-your-own-key support
- Managed keys with pass-through pricing
- AES-256-GCM key encryption
- Logs analytics and cost tracking
- Spend and balance alarms
- Optional PII masking
- Optional trace storage
- Machine-readable error codes with fixes
- MCP server for agent self-fix
Social posts
- No social media activity within the last 30 days
GTM channels
- Blog
- API
- Docs
ICP
- Software developers
- Engineering teams
- Data analytics teams