Base Compute
Builds runtimes and infrastructure for on-device AI inference, optimizing performance for specific hardware and models.
An AI inference lab building runtimes and infrastructure that enable powerful AI models to run on-device, solving latency, privacy, and cost issues associated with cloud-based inference. It sells to developers, IT teams, and enterprises seeking on-premise, hybrid, or air-gapped AI deployment. The approach is differentiated by hardware-specific optimization, achieving significant speedups over alternatives like llama.cpp and MLX, and by offering a zero marginal cost model once deployed on existing hardware.
Key features
- Prefill up to 6.4x faster than llama.cpp
- Decode up to 1.33x faster than MLX
- On-device inference for privacy
- Zero marginal cost per token
- On-premise, hybrid, or air-gapped deployment
- Automated research pipelines
- GPU kernels tuned per hardware and model
- X0.2/day
GTM channels
- Blog
- Community
ICP
- Software developers
- Engineering teams
- Enterprises