
Clusterflock
Unifies GPUs into one backend to load models, run inference, and launch autonomous agent missions
The product is a distributed AI orchestration platform that unifies GPUs across machines into a single backend for loading models, running inference, and executing autonomous agent missions. It is for developers, data teams, and operations/IT teams in businesses that operate their own GPU fleets. It is delivered as open-source, self-hosted infrastructure with an OpenAI-compatible API that routes requests across the cluster.
Key features
- Unified GPU cluster backend
- Automatic VRAM detection and bin-packing
- Self-adapting and self-healing missions
- Mixture of agents orchestration
- Autonomous sandboxed container missions
- Real-time GPU and VRAM telemetry
- Tokens per second monitoring
- Multi-backend support llama.cpp LM Studio
- Metal and CUDA support
- Mix DGX Spark consumer GPUs and Mac
- OpenAI-compatible API on port 1919
- Fanout speed and manual routing modes
- No social media activity within the last 30 days
GTM channels
- API
ICP
- Software developers
- DevOps sre teams
- Data analytics teams