Skip to content
Home

Clusterflock

Unifies GPUs into one backend to load models, run inference, and launch autonomous agent missions

www.clusterflock.netMLOpsApr 2026Notum Robotics

The product is a distributed AI orchestration platform that unifies GPUs across machines into a single backend for loading models, running inference, and executing autonomous agent missions. It is for developers, data teams, and operations/IT teams in businesses that operate their own GPU fleets. It is delivered as open-source, self-hosted infrastructure with an OpenAI-compatible API that routes requests across the cluster.

Key features

  • Unified GPU cluster backend
  • Automatic VRAM detection and bin-packing
  • Self-adapting and self-healing missions
  • Mixture of agents orchestration
  • Autonomous sandboxed container missions
  • Real-time GPU and VRAM telemetry
  • Tokens per second monitoring
  • Multi-backend support llama.cpp LM Studio
  • Metal and CUDA support
  • Mix DGX Spark consumer GPUs and Mac
  • OpenAI-compatible API on port 1919
  • Fanout speed and manual routing modes
  • No social media activity within the last 30 days
GTM channels
  • API
ICP
  • Software developers
  • DevOps sre teams
  • Data analytics teams