Modal
Cloud infrastructure platform for running AI inference, training, batch processing, and sandboxes with container startup, GPU scheduling, autoscaling and observability
Launch screenshots from Product Hunt, Nov 2022
The product is a cloud infrastructure platform for running AI inference, training, batch processing, and sandboxes that handles container startup, GPU scheduling, autoscaling, and observability without manual capacity planning. It is for software developers, data teams, and engineering teams in startups to enterprises building and operating AI applications. It is delivered as a cloud service with a Python SDK and composable primitives, offering sub-second cold starts, autoscaling from zero to thousands of GPUs across clouds, and usage-based pricing.
Key features
- Run LLM inference on H100s/A100s/A10Gs
- Multi-modal inference for image/video/audio
- Batch and async inference at scale
- Online inference with token streaming
- Fine-tuning with SFT and LoRA
- Reinforcement learning with parallel trajectories
- Multi-node training up to 128 GPUs
- Parallel hyperparameter sweeps
- Secure ephemeral sandboxes for untrusted code
- Coding and background agents execution
- GPU-accelerated research sandboxes
- Memory snapshotting for fast cold starts
- Smart filesystem with on-demand loading
- Autoscale from 0 to 1000+ GPUs
- Multi-cloud GPU capacity pooling
- Near-max GPU utilization via batching
- Volumes and Buckets storage
- Queues and Dicts data structures
- Tunnels and Proxies networking
- Real-time observability dashboard
- Granular metrics, logs and traces
- First-party telemetry integrations
- LinkedIn0.8/day
- X0.8/day
- Blog
- Partner program
- Referral program
- Marketplace
- Community
- API
- Docs
- Creative ads
- Software developers
- Engineering teams
- Data analytics teams