InferCrane
Operates production inference behind one stable OpenAI-compatible endpoint, handling deployment, autoscaling, routing, release guarding, and monitoring
The product is an open-source inference operations platform that provides a stable OpenAI-compatible endpoint for model inference, handling deployment, autoscaling, routing, release management, and monitoring so applications do not need to change code when serving plans change. It is for developers, data teams, and operations teams in b2b companies building AI applications, agents, and retrieval workloads. It runs in the customer's infrastructure under Apache-2.0 with a cloud option in private preview and operates owned and external model APIs behind one endpoint.
Key features
- Inference Gateway with one stable endpoint
- Deployments and autoscaling for open-weight inference
- Release Guard for candidate revision comparison
- Monitoring and diagnostics for requests and latency
- Model Catalog for popular families and custom artifacts
- Optimize and Qualify for serving plan measurement
- OpenAI-compatible endpoint routing
- Connect existing vLLM, SGLang, LiteLLM endpoints
- Govern external model APIs with spend limits
- Benchmarking for TTFT, TPOT, and throughput
- No social media activity within the last 30 days
GTM channels
- API
- Docs
ICP
- Software developers
- DevOps sre teams
- Data analytics teams