Genosis
Cache optimization layer that captures provider discounts on cached LLM input tokens across Anthropic, OpenAI, and Google
The product is a cache optimization layer for LLMs run at production scale that captures 50-90% provider discounts on cached input tokens such as system prompts, tool definitions, and RAG chunks across differing provider mechanics. It is for engineering, DevOps, and IT teams at B2B companies building AI applications with significant inference volume. It is delivered as a single-tenant service deployed in the customer's VPC or as a dedicated instance, with a content-blind architecture and open-source SDK, and is paid as a percentage of measured savings reported by the provider API.
Key features
- Single-tenant deployment in customer VPC
- Content-blind architecture with anonymized fingerprints
- Open-source SDK
- Per-provider cache optimization
- Prefix-based and implicit caching support
- Automatic block ordering optimization
- Measured savings via provider API
- Performance-based pricing on savings
- No social media activity within the last 30 days
GTM channels
- Blog
- API
- Docs
ICP
- Engineering teams
- DevOps sre teams
- Software developers