Vlm Run
Provides a unified API for visual intelligence, enabling interaction with images, videos, and documents through a single chat-completions interface, with capabilities like detection, segmentation, OCR, and video analysis.
A unified API gateway that enables developers to build applications that understand, reason over, and act on images, videos, and documents. It solves the problem of integrating multiple vision models and tools by providing a single interface. The platform targets developers and engineering teams at B2B companies, from startups to enterprises, who need to extract structured data from visual content. It is positioned as a drop-in replacement for the OpenAI SDK, offering an OpenAI-compatible API that combines vision-language models with specialized computer-vision tools.
Key features
- OpenAI-compatible API
- Unified interface for images, documents, videos
- Compose visual operations via conversation
- Drop-in replacement for OpenAI SDK
- Auditable outputs with visual proof
- Image captioning and tagging
- Object detection with bounding boxes
- Pixel-perfect segmentation
- Pointing with pixel-level accuracy
- Image generation and editing
- UI parsing from screenshots
- AI-powered image transformations
- Document parsing and summarization
- Structured OCR for messy forms
- Automatic PHI/PII redaction
- Video captioning and tagging
- Video generation and editing
- Video trimming and frame sampling
- Observability and debugging
- Async jobs and retries
- Fine-tuning support
- Auto-evals
- Visual observability for agentic workflows
- Trace explorer
- Batch processing
- Real-time inference
- Agentic workflows
- Visual ETL
- Swap models without rewriting pipelines
- Combine multiple models in one workflow
- X0.2/day
- LinkedIn0.1/day
- Blog
- Partner program
- Marketplace
- Community
- API
- Docs
- Software developers
- Engineering teams
- Enterprises