AI ToolAI Infrastructure

VLM Run

VLM Run Gateway consolidates access to 21+ visual models (OCR, detection, segmentation, video analysis, chat) behind a single OpenAI-compatible API, eliminating the need to manage separate integrations with OpenAI, Claude, Gemini, Qwen, and others.

VLM Run is an AI infrastructure platform, priced at $799/month on the Pro plan, integrating with OpenAI SDK, Zapier, MongoDB, and Claude Code. InnovaAI scores it 5.3/10 for agency resale.

Consider5.3/10

Agency Audit

VLM Run Gateway routes requests to 21+ visual models (OCR, detection, segmentation, video analysis) through a single OpenAI-compatible API, handling document chunking and video reassembly automatically. It outputs structured JSON schemas and integrates with OpenAI SDK, Zapier, Claude Code, and Pydantic AI. Best suited for agencies building document processing, video analysis, or physical AI workflows for clients. Resale potential exists for agencies serving construction, healthcare, or robotics verticals, but pricing is usage-based, making predictable client retainers harder to structure than fixed-tier SaaS.

ConsiderNo WLUsage Hybrid
Fit

5.3/10

Typical Margin

Depends on volume

Time-to-Value

3d about 3 days

Complexity
Low
Consider
Fit53
Visit VLM Run
Best For
  • You serve construction or healthcare clients who need document AI for blueprints, faxes, or clinical paperwork and can pass through token costs as usage charges.
  • Your agency builds internal AI tools and needs a unified API to avoid managing separate integrations with OpenAI, Claude, and Gemini for vision tasks.
  • You have robotics or physical AI clients requiring agentic data-labeling and can leverage MCP server deployment for agent-native vision workflows.
Not For
  • You want to resell a fixed-price retainer without tracking client token usage; VLM Run's per-token and per-operation pricing requires billing infrastructure or usage guardrails.
  • Your clients expect a white-labeled portal or branded interface; VLM Run does not offer a verified white-label program.
  • You need sub-second latency for real-time video processing; the platform is optimized for batch and asynchronous workflows, not live streaming.

Profit Path

Your Cost (USD)

$799/mo

Market Range

$3K–$8K/project

Revenue Model

Monthly Recurring

Planning benchmark at United States price levels. Not a measured market survey.

Platform Features

Core capabilities of VLM Run

21+ model routing via single endpoint

Route OCR, detection, segmentation, chat, and video analysis requests to 21+ visual models (OpenAI, Claude, Gemini, Qwen, Kimi, MiniMax, and others) through one OpenAI-compatible API. Agencies avoid maintaining separate integrations for each model provider.

Automatic document chunking and reassembly

Submit multi-page documents or long-form video in a single call; VLM Run chunks, processes, and reassembles outputs automatically. Reduces client-side orchestration work for document processing agencies.

Structured JSON schema enforcement

Enforce Pydantic schemas on every vision call, guaranteeing structured outputs for downstream workflows. Eliminates post-processing parsing logic and ensures consistent data formats across client projects.

MCP server for agent-native vision

Deploy VLM Run as an MCP server for Claude Code, Codex, and Pydantic AI agents. Agencies can build autonomous workflows that see, reason, and act on images and documents without custom API wrappers.

Video summarization and transcription

Summarize, transcribe, and search long-form video without manual review. Supports per-second billing for video generation and editing, enabling agencies to offer video analysis retainers to media and content clients.

Private VPC and on-premises deployment

Enterprise tier supports in-VPC deployments and on-premises installation. Agencies serving regulated industries (healthcare, finance) can offer compliant visual AI without data leaving client infrastructure.

What Makes VLM Run Different

Unique advantages vs similar tools in this niche

Cost-efficient document OCR at sub-cent per-page

vs Closed OCR APIs like Textract and Azure Doc AI

Pricing calculator shows gateway OCR costs $23-$90 per 100K pages vs $150-$1K for OCR APIs.

Handles 500-page PDFs or 2-hour videos in a single call

vs Building custom chunking and reassembly pipelines

Orchestration is built-in, eliminating pipeline development.

Open-weight models with no lock-in

vs Proprietary VLM APIs

Every model is open-weight and served through OpenAI-compatible API, allowing free model swapping.

Investment ROI Calculator

Value equation analysis for VLM Run, based on the Hormozi framework

What is the Hormozi framework? A four-factor score: (what the service delivers × how reliably it delivers) divided by (how long it takes × how much effort it requires). A higher Value Multiplier means a better return on the time and money invested: faster, easier, and more proven results.

Value MultiplierExcellent

2.8× value multiple: invest $799/mo and agencies typically charge $3K–$8K/project for the work it powers.

Outcome42
÷
Friction15

Why This Succeeds

Higher is better

Implementation Challenges

Lower is better

Strong ROI. VLM Run at $799/mo supports market rates of $3K–$8K. Its 2.8× value-equation score weighs client outcome and likelihood against the time and effort to deliver, not cost.

Best if:You serve construction or healthcare clients who need document AI for blueprints, faxes, or clinical paperwork and can pass through token costs as usage charges.Your agency builds internal AI tools and needs a unified API to avoid managing separate integrations with OpenAI, Claude, and Gemini for vision tasks.You have robotics or physical AI clients requiring agentic data-labeling and can leverage MCP server deployment for agent-native vision workflows.You need SOC 2, HIPAA, and BAA compliance for enterprise clients; the Pro plan ($799/mo) includes BAA and Zero-Data Retention, and Enterprise tier offers custom SLAs.

Pricing

VLM Run platform cost to your agency

Pro: $799/mo

Starter

Custom
  • Pay-as-you-go
  • Up to 10 requests/min
  • Community Discord
  • Basic Usage Logs

Pro

$799/mo
  • Up to 100 requests/min
  • Dedicated Slack support
  • Zero-Data Retention (ZDR)
  • Business Associate Agreement (BAA)
Enterprise

Enterprise

Custom
  • Invoiced billing
  • Tier-based pricing with volume discounts
  • Custom rate-limits
  • In-VPC deployments

How usage-based pricing works

VLM Run charges per consumption unit (per page grounding/confidence surcharge). Below are the component rates the vendor publishes. Each row is a separate charge: your total cost combines them based on your configuration and volume. Component rates range from $0.001 per page grounding/confidence surcharge.

Final agency cost = (sum of selected component rates) × client usage volume. Confirm a usage estimate with each client before quoting.

Component Rates

Cost per unit: total depends on your configuration and volume

Per page grounding/confidence surcharge
$0.001/ page grounding/confidence surcharge
Per code execution call
$0.001/ code execution call
Per image segment (fast/auto)
$0.01/ image segment (fast/auto)
Per page OCR/layout (fast/auto)
$0.01/ page OCR/layout (fast/auto)
Per image segment (pro)
$0.02/ image segment (pro)
Per image generate/edit (fast/auto)
$0.04/ image generate/edit (fast/auto)
Per page OCR/layout (pro)
$0.04/ page OCR/layout (pro)
Per second video generate/edit (fast/auto)
$0.15/ second video generate/edit (fast/auto)
Per image generate/edit (pro)
$0.24/ image generate/edit (pro)
Per 1M input tokens (vlmrun-orion-2:fast / auto)
$0.30/ 1M input tokens (vlmrun-orion-2:fast / auto)
Per second video generate/edit (pro)
$0.40/ second video generate/edit (pro)
Per 1M input tokens (kimi-2.6)
$0.66/ 1M input tokens (kimi-2.6)
Per 1M input tokens (gemini-flash-3.7)
$0.75/ 1M input tokens (gemini-flash-3.7)

Add-ons

Optional extras priced on top of any main plan

Add-on: 1M output tokens (vlmrun-orion-2:fast / auto)
$2.50
Add-on: 1M input tokens (vlmrun-orion-2:pro)
$1
Add-on: 1M output tokens (vlmrun-orion-2:pro)
$10
Add-on: 1M output tokens (kimi-2.6)
$3.41
Add-on: 1M input tokens (muse-spark-1.1)
$1.25
Add-on: 1M output tokens (muse-spark-1.1)
$4.25
Add-on: 1M output tokens (gemini-flash-3.7)
$3.75
Add-on: 1M input tokens (sonnet-5)
$2
Add-on: 1M output tokens (sonnet-5)
$10
Add-on: 1M input tokens (grok-4.5)
$2
Add-on: 1M output tokens (grok-4.5)
$6
Add-on: 1M input tokens (opus-4.8)
$5
Add-on: 1M output tokens (opus-4.8)
$25
Add-on: 1M input tokens (gpt-5.6-sol)
$5
Add-on: 1M output tokens (gpt-5.6-sol)
$30

No verified white-label program for VLM Run: client-facing delivery runs under the platform's native branding.

Reality Check

Trade-offs & Gotchas

Usage-based pricing (input/output tokens, per-page OCR, per-second video) makes client retainer pricing unpredictable. Agencies must either absorb cost variance or implement strict usage caps per client, adding operational complexity. No verified white-label program means client-facing surfaces display VLM Run branding.

Implementation Reality

Moderate effort: standard configuration with some customization needed

Effort: 3/10Time: 5/10

How This Accelerates White-Label Services

Who It's For

  • ai-infrastructure-providers
  • document-processing-agencies
  • video-analysis-agencies
  • enterprise-ai-teams

Acceleration Steps

  1. 1Create your account and complete setup wizard
  2. 2Configure route requests to 21+ visual models through one openai-compatible api
  3. 3Connect OpenAI SDK
  4. 4Launch your first client project

Academy for VLM Run

Work through it in order: the course for this service first, then the modules behind it.

Course for this service

VLM Run Agency Implementation, Building Vision AI Retainers

Learn to architect client workflows using VLM Run's multi-model routing, automatic document chunking, and structured JSON enforcement. This course teaches agencies how to design retainer-based vision AI services for construction, healthcare, and document-heavy verticals while managing usage-based costs through predictable pricing models.

Open the course

Core concepts

The mental model you need to price and scope the work.

  1. Inference Cost Pass-Through CeilingConcept

    Inference Cost Pass-Through Ceiling is the point at which an agency can no longer absorb a model provider's price or latency change inside a fixed retainer, so the cost has to move to the client or the work has to shrink. The framework asks three questions per client engagement: what share of delivery cost is metered inference, how fast can that share be re-routed to a cheaper model, and what contract language lets you reprice. Forrester's 2027 predictions flag AI growth colliding with energy and infrastructure limits, which converts compute scarcity into API price movement on agency tools. A concrete case: an agency running document analysis on a frontier API can shift bulk classification to a smaller open-weight model served through Ollama or a gateway like Helicone, keeping the frontier model only for reasoning steps. That split is the ceiling defense.

  2. Provider Substitution WindowConcept

    Provider Substitution Window is the interval during which an agency can move a client workload from one model provider to another without rewriting prompts, evals, or integration code. The window is widest at the orchestration layer and narrowest at the fine-tuned weights layer: a gateway swap takes hours, a retrained model takes a quarter. Agencies that measure this window per client account know exactly when they hold pricing leverage and when a vendor holds it. Forrester's 2027 predictions flag compute and energy constraints pushing API pricing upward, which turns a wide substitution window into a margin defense rather than an engineering nicety. A concrete case: an agency routing Claude and GPT traffic through a gateway such as Helicone or Portkey can shift a client's summarization workload in an afternoon when one provider raises rates, while a competitor with hardcoded SDK calls absorbs the increase on a fixed retainer.

  3. Margin Defense StackConcept

    Margin Defense Stack treats AI infrastructure as a layered cost structure rather than a single line item. The bottom layer is raw compute and API tokens, the middle layer is routing and caching, and the top layer is the client-facing retainer price. Agencies that only negotiate the top layer absorb every shock from the layers beneath. Forrester's 2027 predictions flag that AI expansion is colliding with energy and infrastructure limits, which translates into API price increases for agency tools and compresses margins on AI-inclusive retainers. A concrete defense: route repeat prompts through a gateway such as Helicone or Portkey so cached responses cut token spend before it reaches the client invoice, and keep a local fallback like Ollama for privacy-sensitive work. When a client asks why the AI retainer costs what it does, the stack shows exactly which layer each dollar covers.

13 modules selected for VLM Run

Frequently Asked Questions

Answers about pricing, setup, implementation

VLM Run offers 3 pricing tiers, at $799/mo (Pro). Agencies typically achieve 55% profit margins when reselling to clients.

Starter is free with pay-as-you-go pricing and up to 10 requests/min. Pro is $799/month with up to 100 requests/min, dedicated Slack support, Zero-Data Retention, and BAA. Enterprise pricing is custom and includes invoiced billing, volume discounts, custom rate-limits, in-VPC deployments, SOC 2, HIPAA, BAA, and custom SLAs. Usage add-ons include per-token charges (input/output) ranging from $0.3 to $30 per 1M tokens depending on model, plus per-operation fees: $0.01 per page OCR (fast/auto), $0.04 per page OCR (pro), $0.15 per second video generation (fast/auto), and $0.4 per second video generation (pro).

No verified white-label program: client-facing surfaces display the VLM Run brand. Agencies can integrate VLM Run's API into their own applications or workflows, but cannot present a fully branded portal or interface to end clients.

Yes. VLM Run provides OpenAI-compatible endpoints, so the OpenAI SDK works natively. Zapier integration is also supported. Additional integrations include MongoDB, Claude Code, Codex, Pydantic AI, and Mastra, enabling agencies to embed VLM Run into multi-tool workflows.

Setup time depends on deployment model. API key provisioning is immediate (minutes). For agencies building client-facing applications, integration time ranges from hours (simple document OCR) to days (multi-model orchestration with custom schemas). Enterprise deployments with in-VPC or on-premises options require coordination with the VLM Run team.

Construction (blueprint and spec analysis), healthcare (fax and clinical paperwork processing), robotics and physical AI (agentic data-labeling), and enterprise AI teams building internal vision workflows. Document processing agencies and video analysis agencies are also strong fits.