AI ToolAI Evaluation Observability

Referee

Referee.chat is a quality-assurance platform that runs user-defined prompts across multiple LLM models and judges the outputs against explicit quality standards (publication-grade, formal proof, decision-grade, or custom).

Referee is an AI evaluation observability platform, priced at $10/month on the Credit plan, integrating with Anthropic Claude, OpenAI GPT, Google Gemini, and Mistral. InnovaAI scores it 4.9/10 for agency adoption, best for Strategist, Project Manager, and Researcher roles handling 5+ client meetings per week.

Situational Fit4.9/10

Agency Audit

Referee.chat runs multiple LLM models against user-defined quality standards and judges outputs based on evidence rather than model confidence, enabling agencies to validate AI-generated work before publication or client delivery. Best suited for research teams, content verification workflows, and decision-support consultancies that need to formalize AI outputs to publication, formal-proof, or decision-grade standards. Integrates with 40+ model providers including Claude, GPT, Gemini, and Mistral, allowing teams to compare outputs across vendors without vendor lock-in. Agencies adopting Referee internally compress quality-assurance cycles for AI-generated research, analysis, and recommendations by replacing subjective review with evidence-based verdicts.

Situational FitNo WLTiered
Seats

5recommended

Est. Hours Saved

60/mo

Net Capacity

$4,490/mo

Friction

Low

Illustrative scenario. Not a guarantee. Net capacity is the value of reclaimed time at $75/hr, less the lowest verified paid base plan (flat plan cost is shared). Hours saved come from the service estimate; implementation, taxes, and unprovided usage charges are excluded.

Situational Fit
Fit49
Visit Referee
Best For Your Team
  • Strategist handling multi-model output comparison
  • Project Manager handling AI quality assurance before client delivery
  • Researcher handling formal proof validation
Not Ideal If
  • Your agency rarely generates AI outputs that require publication-grade or decision-grade validation. If most AI use is exploratory or internal brainstorming, Referee's overhead will not pay back.
  • Your team works with a single LLM provider and has no need to compare outputs across models. Referee's core value is multi-model evaluation; single-vendor shops get less ROI.
  • Your quality-assurance process is already lightweight and manual review takes less than 2 hours per week across the team. The cost per seat and setup friction will exceed the time savings.

Internal Adoption Path

Team Subscription

$10/mo

$10/mo flat plan

Time Saved Monthly

60 hr/mo

5 seats × 12 hr each

Value of Reclaimed Time

$4,500/mo

modeled at $75/hr labor rate

Net Capacity

$4,490/mo

value − subscription cost

In this model, 5 seats reclaim 60 hours of team time each month. Valued at $75/hr that is $4,500/mo, and after the $10/mo subscription it leaves $4,490/mo of capacity for billable client work.

Illustrative scenario. Not a guarantee. Uses the lowest verified paid base plan. Implementation, taxes, and unprovided usage charges are excluded.

Platform Features

Core capabilities of Referee

Multi-model panel evaluation

Run the same prompt across 40+ LLM providers (Claude, GPT, Gemini, Mistral, DeepSeek, and others) in a single match. Strategists and researchers compare outputs side-by-side without manually querying each vendor, saving 1-2 hours per evaluation cycle.

Evidence-based referee verdict

Referee judges outputs against your stated quality bar (publication standard, formal proof, decision-grade, or custom) and surfaces specific evidence gaps rather than declaring a winner based on confidence. Project managers and content leads get actionable feedback on what each model got right or wrong.

Custom quality standards

Define your own acceptance criteria (e.g., 'every claim must be sourced', 'all assumptions stated', 'risks named and costed'). Operations and founders lock in quality gates that persist across all future evaluations, removing ambiguity from QA workflows.

Formal proof integration (theorem.chat)

For mathematical or technical claims, Referee formalizes the panel's best argument in Lean against Mathlib, where the kernel verifies correctness. Research teams working on proofs or formal specifications eliminate manual verification steps.

Decision-grade recommendation output

Generate structured recommendations with costed options and named risks, ready for stakeholder sign-off. Strategists and account executives compress the time from analysis to client-ready recommendation by 2-3 hours per engagement.

Private or public match storage

Free tier publishes matches publicly; Credit plan keeps work private with up to 6 seats and refundable unused credits. Project managers and operations teams choose privacy and collaboration scope based on client confidentiality needs.

What Makes Referee Different

Unique advantages vs similar tools in this niche

Evidence-based judging over confidence-based

vs Traditional AI evaluation tools that rely on model confidence scores

The referee rules on evidence, not on who sounded surer, ensuring objective quality assessment.

Formal proof verification in Lean

vs Manual proof checking or informal verification

For mathematical claims, results are formalized in Lean against Mathlib, where the kernel decides correctness.

Customizable quality bars

vs One-size-fits-all evaluation criteria

Users can set standards from 'Careful expert' to 'Prize standard' or write their own, tailoring evaluation to specific needs.

Value Equation

Outcome-likelihood-time-effort assessment for Referee

Limited agency channel

Referee scored below the agency-resellability threshold (agency_fit_score < 50). The Value Equation projects agency-side outcomes, which don't apply to tools without a clear resell pathway.

Contact Referee

Pricing

Referee platform cost to your agency

Starts at $1 one-time (Free), scales to $10/mo (Credit)

Free

$1 one-time
  • Enough for a real match
  • Every model and every tool
  • Full transcript, evidence and verdict
  • Matches are published, not private

Credit

$10/mo
  • Private matches, your work stays yours
  • Any model, any panel size, up to 6 seats
  • Longer runs and higher spending limits
  • Resume and branch without limits

Your own key

Custom
  • Bring your own inference key
  • You pay your provider directly
  • The first $5.00 of usage costs you nothing here
  • Private matches included

No verified white-label program for Referee: client-facing delivery runs under the platform's native branding.

Market Intelligence

Offer + scale economics for Referee

Limited agency channel

Referee scored below the agency-resellability threshold (agency_fit_score < 50). It's a useful tool but not designed for white-labeled or retainer-based reselling, so we don't publish productized offer economics for it.

Contact Referee

Investment Decision Framework

Strategic vetting analysis for Referee

Vetting Verdict

Situational Fit

Fit depends on your client mix

Agency Fit(white-label + resell pathway)
49/100
0255075100
Resell Friction(WL + mode + complexity)
85/100
0255075100

Buy If

5
STRATEGIC DRIVER

Your strategists and researchers spend 3+ hours per week manually reviewing AI-generated analysis for factual accuracy, sourcing, and logical gaps before client delivery. Referee automates that review by running multiple models against a publication-standard bar and surfacing evidence-backed verdicts.

OPERATIONAL FIT

Your project managers need to formalize AI-generated recommendations (e.g., technical architecture, budget scenarios, risk assessments) into decision-grade outputs that stakeholders can act on with confidence. Referee costed options and named risks into a single verdict.

OPERATIONAL FIT

Your content and research teams work across multiple LLM providers and need a neutral comparison framework to select which model output to use for a given task. Referee runs the same prompt across your entire panel and ranks results by evidence, not by which vendor sounds most confident.

OPERATIONAL FIT

Your operations or founder role is building internal AI workflows (research summaries, proposal drafts, technical specs) and needs a quality gate that doesn't rely on human spot-checking every output. Referee holds a literal bar and flags when the panel falls short.

OPERATIONAL FIT

Your team formalizes mathematical proofs or technical arguments and needs to validate them against a formal standard. Referee integrates with theorem.chat to formalize results in Lean against Mathlib, where the kernel decides correctness.

Skip If

5
CAUTION

Your agency rarely generates AI outputs that require publication-grade or decision-grade validation. If most AI use is exploratory or internal brainstorming, Referee's overhead will not pay back.

CAUTION

Your team works with a single LLM provider and has no need to compare outputs across models. Referee's core value is multi-model evaluation; single-vendor shops get less ROI.

CAUTION

Your quality-assurance process is already lightweight and manual review takes less than 2 hours per week across the team. The cost per seat and setup friction will exceed the time savings.

CAUTION

Your agency cannot define explicit quality standards for the work being evaluated. Referee requires you to articulate what 'done' looks like (publication standard, formal proof, decision-grade) before running a match; teams without clear acceptance criteria will struggle.

CAUTION

Your team needs real-time feedback on AI outputs during client calls or live presentations. Referee is a batch-evaluation tool; it does not provide in-the-moment guidance or streaming verdicts.

Bottom Line

Referee.chat runs multiple LLM models against user-defined quality standards and judges outputs based on evidence rather than model confidence, enabling agencies to validate AI-generated work before publication or client delivery. Best suited for research teams, content verification workflows, and decision-support consultancies that need to formalize AI outputs to publication, formal-proof, or decision-grade standards. Integrates with 40+ model providers including Claude, GPT, Gemini, and Mistral, allowing teams to compare outputs across vendors without vendor lock-in. Agencies adopting Referee internally compress quality-assurance cycles for AI-generated research, analysis, and recommendations by replacing subjective review with evidence-based verdicts.

Reality Check

Trade-offs & Gotchas

Referee requires teams to define explicit quality bars upfront (publication standard, formal proof, decision-grade, or custom), which adds initial setup friction. Adoption ROI concentrates in agencies running 5+ AI evaluation cycles per week; smaller teams may find the overhead outweighs the payoff.

Implementation Reality

Moderate effort: standard configuration with some customization needed

Effort: 4/10Time: 4/10

Academy for Referee

Work through it in order: the course for this service first, then the modules behind it.

Course for this service

Referee Agency Implementation, Monetizing AI Quality Assurance

Learn how to package Referee's multi-model evaluation and evidence-based verdicts into retainer services for clients who need reliable AI outputs. This course covers setting up custom quality standards, running competitive model panels, delivering decision-grade recommendations, and building recurring revenue from AI quality assurance.

Open the course

Core concepts

The mental model you need to price and scope the work.

  1. Eval Debt CompoundingConcept

    Eval debt is the accumulated gap between what an AI agent does in production and what anyone on the agency team can actually prove it does. Like technical debt, it accrues quietly and charges interest: every untraced failure mode, every scoring rubric that lives in a Slack thread, every client demo that worked once and was never re-run. The interest payment arrives as a retainer conversation. Agencies that instrument early convert that debt into a premium line item, because "production-ready" is a claim only evidence can support. The cost curve is moving in their favor: OpenAI cut GPT-6 Sol and Luna API prices 50% versus GPT-5.6, and prompt caching now discounts up to 90% on reused prefixes, so high-volume agent pipelines are cheaper to run and cheaper to trace. Meanwhile Forrester's 2027 predictions flag compute and infrastructure constraints that will push API-dependent tool costs upward, compressing margins on AI-inclusive retainers. Tracing spend is the hedge.

  2. Silent Failure SurfaceConcept

    The Silent Failure Surface is the set of AI behaviors that pass every automated check yet still damage the client relationship: a voice agent that interrupts callers, a support bot that loops a user through three retries, a research agent that returns confident but stale answers. Standard evals score outputs against expected answers, so they miss friction that only appears in live sessions. Agencies that map this surface before launch can price a monitoring retainer against it; agencies that skip it discover failures when the client forwards a complaint. Cekura simulates thousands of personas to expose interruption and gibberish patterns before go-live, while Agnost AI ingests real conversations and flags repeated retries and broken workflows as actionable intents. Both approaches treat production traffic as the primary test set, not a post-launch afterthought. The surface shrinks only when someone owns the loop between detection and a shipped fix.

  3. Trace-to-Trust RatioConcept

    Trace-to-Trust Ratio is the proportion of an AI agent's production behavior that is actually instrumented, logged, and reviewable, measured against the trust a client extends to that system. Agencies that instrument every LLM call, tool invocation, and retrieval step can show clients exactly what happened when an output went wrong, which converts a vague reliability claim into a defensible audit trail. The ratio matters because trust is not granted by model choice; it is granted by evidence. A voice agent handling inbound calls with no tracing is a liability, while one instrumented through a platform like Cekura or Langfuse can surface interruption rates, gibberish detection, and latency per session. When a client asks why a response was wrong, the agency with trace coverage answers in minutes; the agency without it answers with a guess. That gap is where retainer renewals and premium pricing are decided.

Frequently Asked Questions

Answers about pricing, setup, implementation

Referee.chat runs multiple LLM models against a user-defined quality standard and judges the outputs based on evidence rather than model confidence. You set the bar (publication standard, formal proof, decision-grade, or custom), seat a panel of models, and Referee tells you which outputs meet the bar and why. For mathematical claims, it integrates with theorem.chat to formalize proofs in Lean against Mathlib.

The Free plan costs $1 one-time and includes every model and full transcripts, but matches are published publicly. The Credit plan costs $10 per month per seat (up to 6 seats), keeps matches private, and includes refundable unused credits and higher spending limits. You can also bring your own inference key and pay your LLM provider directly; Referee charges $0.05 per unit of infrastructure after the first $5.00 of usage, with no monthly seat fee.

Strategists and researchers compress quality-assurance cycles by running multi-model evaluations against publication or formal-proof standards. Project managers formalize AI-generated recommendations into decision-grade outputs with costed options and named risks. Content leads and operations teams use Referee as a quality gate to validate AI outputs before client delivery, replacing manual spot-checking with evidence-backed verdicts. Founders building internal AI workflows use Referee to lock in quality standards that persist across all future evaluations.

A strategist or researcher running 5+ AI evaluation cycles per week typically saves 3-5 hours per week by replacing manual multi-model comparison and quality review with a single Referee match. A project manager formalizing 2-3 AI-generated recommendations per week saves 2-3 hours by automating the evidence-gathering and risk-naming step. Payback depends on baseline QA time; teams spending less than 2 hours per week on AI review will see minimal ROI.

Referee integrates with 40+ model providers including Anthropic Claude, OpenAI GPT (including o1, o3, GPT-5 series), Google Gemini, Mistral, DeepSeek, Meta Llama, Cohere, AI21, Amazon Nova, and many others. You can also bring your own inference key and use Referee with your existing vendor relationships. The full model list is available on the Referee.chat platform.

Referee does not publish pricing or data-retention terms in the extracted content. Contact Referee.chat support for details on data export, deletion, and retention policies after cancellation.

Initial setup takes 30-60 minutes per team: define your quality standards (publication, formal proof, decision-grade, or custom), select your model panel, and run a test match. Rollout to the full team is low-friction; each team member can start running matches immediately after onboarding. No infrastructure changes or API integrations required.

Referee automates the evidence-gathering and verdict step of QA, but does not replace human judgment on whether the quality bar itself is appropriate for a given task. Use Referee to formalize and speed up QA workflows, not to eliminate human review entirely. Teams typically use Referee to flag outputs that fall short of the bar, then route those outputs back to the model panel or to human review.