Inferencer
Runs AI models locally with full control and privacy, including token inspection, model serving, and memory offloading.
Inferencer is a desktop application that runs AI models locally on the user's device, ensuring privacy by processing all data offline. It targets developers and engineering teams who need to control and inspect AI model outputs, offering features like token probability inspection, model serving over networks, and memory offloading for large models. The product is positioned as a transparent alternative to cloud-based AI services, with a focus on deep control and privacy.
Key features
- Token probability inspection
- Token entropy and selection
- Token exclusion
- Mixture of Experts control
- Prompt prefilling
- Custom tool calls
- Private server with SSL
- Ollama and OpenAI compatible APIs
- Mobile support
- Agent support
- Persistent prompt caching
- Distributed inference
- Sandboxed execution
- Markdown and LaTeX rendering
- Batching
- Model streaming
- Model compression
- Context compression
- No social media activity within the last 30 days
GTM channels
- Newsletter
- API
- Changelog
ICP
- Software developers
- Engineering teams
- DevOps sre teams