AI ToolAI Infrastructure

IQ Routing

IQ Routing is an LLM gateway that intercepts API calls and routes each request to the cheapest model capable of meeting quality requirements for that specific task.

IQ Routing is an LLM gateway, priced at $70/month on the Team plan, integrating with OpenAI, Anthropic, Google, and Claude Code. InnovaAI scores it 5.8/10 for agency resale.

Consider5.8/10

Agency Audit

IQ Routing sits between your LLM calls and OpenAI, Anthropic, or Google endpoints, automatically routing each request to the cheapest model capable of handling it while maintaining quality. Agencies building chatbots, RAG systems, or agent workflows can reduce LLM costs by 40-80% without changing application code. The tool works inside locked environments like Claude Code and Cursor, making it useful for agencies that need model consistency but want cost control. Best fit: AI product agencies and those running multi-step agent loops where per-step routing can compound savings.

ConsiderNo WLFreemium
Fit

5.8/10

Typical Margin

63%

Time-to-Value

3d about 3 days

Complexity
Low
Consider
Fit58
Visit IQ Routing
Best For
  • You have clients running LangChain agent loops or multi-step RAG pipelines where different steps have different complexity requirements, since IQ Routing's per-step resolver can route planning to a reasoning model and verification to a cheaper variant.
  • Your clients use Claude Code or Cursor and need cost control without switching model families, because IQ Routing maintains model family constraints while routing within tiers.
  • You manage 5+ client accounts and need per-team budgets and audit logs to track spend by client, since the Team plan ($70/mo) includes per-team controls and cost reports by model and team.
Not For
  • Your clients require HIPAA or FedRAMP compliance, since IQ Routing only offers SOC2 evidence on request and does not publish healthcare or government compliance certifications.
  • You need white-label client portals or branded cost dashboards, because no verified white-label program exists in the provided content.
  • Your clients run fewer than 240 requests per minute on the Free plan and cannot justify the $70/mo Team plan, since the Free tier caps at 240 requests/min and lacks per-team budgets.

Profit Path

Your Cost (USD)

$70/mo

Market Range

$1K–$3K/project

Revenue Model

Hybrid

Planning benchmark at United States price levels. Not a measured market survey.

Platform Features

Core capabilities of IQ Routing

Per-step routing in agent loops

Routes each step of a multi-step agent workflow to the cheapest model that meets quality requirements for that step. Agencies can reduce total loop cost by 58% or more by using reasoning models only where needed and cheaper variants for retrieval, tool calls, and verification.

Unified endpoint for three providers

Accepts OpenAI or Anthropic SDK calls at a single URL and routes to OpenAI, Anthropic, or Google models. Agencies no longer need to maintain separate integrations or client code changes when switching between providers.

Semantic caching

Detects when a new request matches a previously answered question (even with different wording) and returns the cached result in approximately 11 milliseconds without re-billing. Reduces redundant API spend for clients with repetitive query patterns.

Per-team budgets and audit logs

Team plan includes per-team spending limits, alerting, and audit trails so agencies can track which internal team or client account spent what and why. Enables board-ready cost reports by model, team, and savings layer.

Model family lock-in for Claude Code and Cursor

Works inside tools that restrict model selection to a single family. Agencies can point Claude Code or Cursor at IQ Routing and route to the right tier within that family without leaking to another vendor, maintaining conversation state across model switches.

Per-step cost and latency tracking

Dashboard shows cost, latency, and token count for each step in an agent session, so agencies can identify which steps are expensive and adjust routing rules or quality thresholds accordingly.

What Makes IQ Routing Different

Unique advantages vs similar tools in this niche

Purpose-built routing classifier that holds quality where naive cheapest-model routers drop it

vs Simple cost-based routers that sacrifice quality

The router weighs true difficulty against live cost and latency, with instant fallback if a model slips.

Semantic cache that catches repeats and paraphrases, cutting costs on repeated queries

vs Exact-match caching in other gateways

IQ catches exact repeats and ones that just mean the same thing, with per-team cache isolation.

Per-step routing for agent loops, assigning each step the cheapest model that can do it well

vs Pinning one frontier model for all steps

The resolver picks the cheapest variant per step, as shown in the LangChain example cutting cost 58%.

Investment ROI Calculator

Value equation analysis for IQ Routing, based on the Hormozi framework

What is the Hormozi framework? A four-factor score: (what the service delivers × how reliably it delivers) divided by (how long it takes × how much effort it requires). A higher Value Multiplier means a better return on the time and money invested: faster, easier, and more proven results.

Value MultiplierExcellent

2.7× value multiple: invest $70/mo and agencies typically charge $1K–$3K/project for the work it powers.

Outcome40
÷
Friction15

Why This Succeeds

Higher is better

Implementation Challenges

Lower is better

Strong ROI. IQ Routing at $70/mo supports market rates of $1K–$3K. Its 2.7× value-equation score weighs client outcome and likelihood against the time and effort to deliver, not cost.

Best if:You have clients running LangChain agent loops or multi-step RAG pipelines where different steps have different complexity requirements, since IQ Routing's per-step resolver can route planning to a reasoning model and verification to a cheaper variant.Your clients use Claude Code or Cursor and need cost control without switching model families, because IQ Routing maintains model family constraints while routing within tiers.You manage 5+ client accounts and need per-team budgets and audit logs to track spend by client, since the Team plan ($70/mo) includes per-team controls and cost reports by model and team.Your clients have predictable query patterns where semantic caching can reduce redundant API calls, since cache hits resolve in approximately 11 milliseconds without re-billing.

Pricing

IQ Routing platform cost to your agency

~63% margin

Team: $70/mo

Free

$0/mo
Free forever
  • OpenAI, Anthropic, and Google, with one URL for all three
  • Bring your own keys (BYOK)
  • Semantic cache
  • 240 requests/min usage limit

Team

$70/mo
  • Everything in Free
  • Per-team budgets
  • Audit log
  • Per-team alerting
Enterprise

Enterprise

Custom
  • Everything in Team
  • On-prem or VPC deployment
  • SOC2 evidence pack on request
  • ERP integrations (on the roadmap)

No verified white-label program for IQ Routing: client-facing delivery runs under the platform's native branding.

Market Intelligence

How agencies monetize IQ Routing: real offer economics and market positioning

Service Applications
Automation & IntegrationsReporting & AnalyticsDelivery & Production
Best For
  • AI product agencies
  • Agencies building chatbots
  • Agencies running RAG systems
Not Ideal For
  • Agencies not using LLM APIs
  • Agencies with minimal AI infrastructure spend

Project-Based

ai-tools

Agency charges per-project fee for implementation. Ongoing optimization as optional retainer.

Offer Economics: What You Charge vs. What It Costs

Margin includes platform cost + agency labor at $75/hr.

IQ Routing SMB Starterlocal smb

Local service businesses or solo practitioners using OpenAI/Anthropic APIs who want to cut LLM spend without rebuilding their stack

$2.5K
Tool: $70/mo (2 mo = $140)Labor: 20h setup × $75 = $1.5KMargin: 34%Benchmark: $1K–$3K/project
Configure IQ Routing gateway as a drop-in replacement for existing OpenAI or Anthropic SDK endpointSet up semantic cache rules to reduce redundant API calls for common queriesDeploy cost-band routing policy (cheap vs. frontier) matched to client's use-case quality requirementsDocument handoff guide with cost dashboard walkthrough and savings baseline report
IQ Routing Growth Deploymentgrowth smb

Funded startups or regional brands with multiple teams consuming LLM APIs who need budget controls and audit visibility

$5.5K
Tool: $70/mo (2 mo = $140)Labor: 48h setup × $75 = $3.6KMargin: 32%Benchmark: $3K–$8K/project
Configure per-team budget limits and alerting rules across all active product teamsIntegrate IQ Routing unified endpoint with existing CI/CD pipeline and staging environmentsBuild custom band-map routing logic aligned to each team's latency and quality tolerancesSet up audit log export and monthly cost-savings reporting dashboard for stakeholders
IQ Routing Mid-Market Optimizationmid marketHIGH MARGIN

Mid-size companies with 50–500 employees running multi-step AI agent loops or internal LLM tooling at scale who need governance and cost accountability

$14K
Tool: $70/mo (2 mo = $140)Labor: 80h setup × $75 = $6KMargin: 56%Benchmark: $8K–$20K/project
Deploy IQ Routing with org-scoped access controls and per-team budget enforcement across all business unitsConfigure per-step routing for existing agent loops to minimize frontier model usage on low-complexity stepsIntegrate semantic caching layer and tune cache TTL policies based on historical request pattern analysisBuild executive cost governance report template with projected vs. actual LLM spend tracking
IQ Routing Enterprise GatewayenterpriseHIGH MARGIN

Enterprise organizations with 500+ employees requiring on-prem or VPC deployment, SOC2 compliance evidence, and centralized LLM cost governance across divisions

$42K
Tool: $70/mo (2 mo = $140)Labor: 160h setup × $75 = $12KMargin: 71%Benchmark: $20K–$60K/project
Deploy IQ Routing in client's VPC or on-prem environment with SOC2 evidence pack configuration and security reviewIntegrate unified LLM gateway with existing ERP and internal tooling systems across all consuming teamsConfigure division-level budget controls, alerting thresholds, and audit log pipelines into SIEM or data warehouseOptimize per-step agent routing policies and deliver 90-day cost reduction roadmap with baseline benchmarks

Scale Economics: Based on Starter Offer

Using IQ Routing SMB Starter at $2.5K/client. Platform: $70/mo. Labor: 4h/client × $75/hr.

5 clients
$12.5K
MRR
$10.9K net (87%)
10 clients
$25K
MRR
$21.9K net (88%)
20 clients
$50K
MRR
$43.9K net (88%)

Net = MRR - platform cost - labor (4h/client × $75/hr).

Weighted Avg Margin
63%
Across all offer tiers, incl. labor at $75/hr
Run your agency audit

Investment Decision Framework

Strategic vetting analysis for IQ Routing

Vetting Verdict

Consider

Favorable fit, worth a closer look

Agency Fit(white-label + resell pathway)
58/100
0255075100
Resell Friction(WL + mode + complexity)
60/100
0255075100

Buy If

4
OPERATIONAL FIT

You have clients running LangChain agent loops or multi-step RAG pipelines where different steps have different complexity requirements, since IQ Routing's per-step resolver can route planning to a reasoning model and verification to a cheaper variant.

OPERATIONAL FIT

Your clients use Claude Code or Cursor and need cost control without switching model families, because IQ Routing maintains model family constraints while routing within tiers.

OPERATIONAL FIT

You manage 5+ client accounts and need per-team budgets and audit logs to track spend by client, since the Team plan ($70/mo) includes per-team controls and cost reports by model and team.

OPERATIONAL FIT

Your clients have predictable query patterns where semantic caching can reduce redundant API calls, since cache hits resolve in approximately 11 milliseconds without re-billing.

Skip If

4
DEAL BREAKER

You want to resell IQ Routing as a standalone managed service to non-technical clients, since the tool requires API key management and quality threshold tuning that demands technical setup.

CAUTION

Your clients require HIPAA or FedRAMP compliance, since IQ Routing only offers SOC2 evidence on request and does not publish healthcare or government compliance certifications.

CAUTION

You need white-label client portals or branded cost dashboards, because no verified white-label program exists in the provided content.

CAUTION

Your clients run fewer than 240 requests per minute on the Free plan and cannot justify the $70/mo Team plan, since the Free tier caps at 240 requests/min and lacks per-team budgets.

Bottom Line

IQ Routing sits between your LLM calls and OpenAI, Anthropic, or Google endpoints, automatically routing each request to the cheapest model capable of handling it while maintaining quality. Agencies building chatbots, RAG systems, or agent workflows can reduce LLM costs by 40-80% without changing application code. The tool works inside locked environments like Claude Code and Cursor, making it useful for agencies that need model consistency but want cost control. Best fit: AI product agencies and those running multi-step agent loops where per-step routing can compound savings.

Reality Check

Trade-offs & Gotchas

IQ Routing requires agencies to manage separate API keys for OpenAI, Anthropic, and Google upfront, and cost visibility depends on accurate quality thresholds being set per step. If a client's workload doesn't have clear cost-quality tradeoffs (e.g., all requests genuinely need frontier models), savings will be minimal.

Implementation Reality

Moderate effort: standard configuration with some customization needed

Effort: 3/10Time: 5/10

Academy for IQ Routing

Work through it in order: the course for this service first, then the modules behind it.

Course for this service

IQ Routing Agency Implementation, Cost-Optimized LLM Delivery

Learn how to deploy IQ Routing as a managed service for clients, reduce their LLM spend by 40-80% through intelligent model routing and semantic caching, and build recurring revenue by monitoring per-team budgets and cost optimization. This course covers gateway setup, per-step routing configuration for agent workflows, client onboarding, and cost reporting that justifies retainer fees.

Open the course

Core concepts

The mental model you need to price and scope the work.

  1. IQ Routing Cost Band StrategyConcept

    IQ Routing's core value for agencies is its per-step routing capability, which assigns each API call to the cheapest model that can handle it without sacrificing quality. This framework, the Cost Band Strategy, involves mapping your client's workflows into distinct quality tiers: 'cheap' for simple tasks like classification or extraction, 'auto' for balanced performance, and 'frontier' for complex reasoning or creative generation. By configuring custom band maps in IQ Routing, you can enforce these tiers per step, ensuring that a summarization call uses a cheaper model while a code generation step uses a frontier one. For an agency running a multi-step agent loop for a client, this compounds savings across thousands of calls, potentially cutting LLM costs by 40-80%. The Team plan at $70/month includes custom band maps, making it a low-cost investment that can be passed on as a managed service retainer, improving your margin on AI projects.

  2. Inference Cost Pass-Through CeilingConcept

    Inference Cost Pass-Through Ceiling is the point at which an agency can no longer absorb a model provider's price or latency change inside a fixed retainer, so the cost has to move to the client or the work has to shrink. The framework asks three questions per client engagement: what share of delivery cost is metered inference, how fast can that share be re-routed to a cheaper model, and what contract language lets you reprice. Forrester's 2027 predictions flag AI growth colliding with energy and infrastructure limits, which converts compute scarcity into API price movement on agency tools. A concrete case: an agency running document analysis on a frontier API can shift bulk classification to a smaller open-weight model served through Ollama or a gateway like Helicone, keeping the frontier model only for reasoning steps. That split is the ceiling defense.

  3. Provider Substitution WindowConcept

    Provider Substitution Window is the interval during which an agency can move a client workload from one model provider to another without rewriting prompts, evals, or integration code. The window is widest at the orchestration layer and narrowest at the fine-tuned weights layer: a gateway swap takes hours, a retrained model takes a quarter. Agencies that measure this window per client account know exactly when they hold pricing leverage and when a vendor holds it. Forrester's 2027 predictions flag compute and energy constraints pushing API pricing upward, which turns a wide substitution window into a margin defense rather than an engineering nicety. A concrete case: an agency routing Claude and GPT traffic through a gateway such as Helicone or Portkey can shift a client's summarization workload in an afternoon when one provider raises rates, while a competitor with hardcoded SDK calls absorbs the increase on a fixed retainer.

13 modules selected for IQ Routing

Frequently Asked Questions

Answers about pricing, setup, implementation

IQ Routing is an LLM gateway that intercepts API calls to OpenAI, Anthropic, or Google and routes each request to the cheapest model capable of handling it while maintaining quality. It provides semantic caching to avoid re-billing repeated requests, per-step routing for agent loops, and per-team budgets so agencies can track spend by client or internal team. The tool drops in front of existing OpenAI or Anthropic SDKs without requiring application code changes.

IQ Routing offers 3 pricing tiers, at $70/mo (Team). Agencies typically achieve 63% profit margins when reselling to clients.

No verified white-label program exists in the provided content. Client-facing surfaces display the IQ Routing brand. If white-label capabilities are planned, contact sales to confirm availability.

Yes. IQ Routing natively integrates with OpenAI, Anthropic, and Google. It also works with Claude Code, Cursor, and LangChain. Any OpenAI or Anthropic-shaped endpoint can point to IQ Routing's unified URL, and the tool routes to the appropriate model based on your quality thresholds.

IQ Routing goes live in approximately 30 seconds once you point your existing OpenAI or Anthropic SDK at the IQ Routing endpoint. No application code changes are required. Per-client setup depends on configuring quality thresholds and band maps for your specific workflows.

AI product agencies, agencies building chatbots, agencies running RAG systems, and agencies with agent workflows. Any client running multi-step LLM pipelines where different steps have different complexity requirements will see the largest cost savings.

Yes, on the Team plan and above. IQ Routing supports per-team budgets, per-team alerting, and org-scoped access controls. You can generate cost reports by model, team, and savings layer, so each client's spend is visible and auditable.

Requests will fail if they continue pointing to IQ Routing's endpoint without an active account. Agencies should migrate clients back to direct OpenAI or Anthropic SDK calls before cancellation, or maintain a fallback endpoint. IQ Routing does not publish a data export or retention policy in the provided content.