Referee
Referee.chat is a quality-assurance platform that runs user-defined prompts across multiple LLM models and judges the outputs against explicit quality standards (publication-grade, formal proof, decision-grade, or custom). The referee verdict is based on evidence gaps and claim sourcing, not on which model sounds most confident. Agencies define the quality bar once, then run unlimited matches; each match compares outputs across 40+ model providers and surfaces which outputs meet the bar and why. For mathematical or formal arguments, Referee integrates with theorem.chat to formalize the best output in Lean and verify it against Mathlib. Private or public storage, bring-your-own-key option, and up to 6 seats per subscription.
Referee is an AI evaluation observability platform, priced at $10/month on the Credit plan, integrating with Anthropic Claude, OpenAI GPT, Google Gemini, and Mistral. InnovaAI scores it 4.9/10 for agency adoption, best for Strategist, Project Manager, and Researcher roles handling 5+ client meetings per week.
Agency Audit
Referee.chat runs multiple LLM models against user-defined quality standards and judges outputs based on evidence rather than model confidence, enabling agencies to validate AI-generated work before publication or client delivery. Best suited for research teams, content verification workflows, and decision-support consultancies that need to formalize AI outputs to publication, formal-proof, or decision-grade standards. Integrates with 40+ model providers including Claude, GPT, Gemini, and Mistral, allowing teams to compare outputs across vendors without vendor lock-in. Agencies adopting Referee internally compress quality-assurance cycles for AI-generated research, analysis, and recommendations by replacing subjective review with evidence-based verdicts.
5recommended
60/mo
$4,490/mo
Low
Illustrative scenario. Not a guarantee. Net capacity is the value of reclaimed time at $75/hr, less the lowest verified paid base plan (flat plan cost is shared). Hours saved come from the service estimate; implementation, taxes, and unprovided usage charges are excluded.
- Strategist handling multi-model output comparison
- Project Manager handling AI quality assurance before client delivery
- Researcher handling formal proof validation
- Your agency rarely generates AI outputs that require publication-grade or decision-grade validation. If most AI use is exploratory or internal brainstorming, Referee's overhead will not pay back.
- Your team works with a single LLM provider and has no need to compare outputs across models. Referee's core value is multi-model evaluation; single-vendor shops get less ROI.
- Your quality-assurance process is already lightweight and manual review takes less than 2 hours per week across the team. The cost per seat and setup friction will exceed the time savings.
Internal Adoption Path
$10/mo
$10/mo flat plan
60 hr/mo
5 seats × 12 hr each
$4,500/mo
modeled at $75/hr labor rate
$4,490/mo
value − subscription cost
In this model, 5 seats reclaim 60 hours of team time each month. Valued at $75/hr that is $4,500/mo, and after the $10/mo subscription it leaves $4,490/mo of capacity for billable client work.
Illustrative scenario. Not a guarantee. Uses the lowest verified paid base plan. Implementation, taxes, and unprovided usage charges are excluded.
Platform Features
Core capabilities of Referee
Multi-model panel evaluation
Run the same prompt across 40+ LLM providers (Claude, GPT, Gemini, Mistral, DeepSeek, and others) in a single match. Strategists and researchers compare outputs side-by-side without manually querying each vendor, saving 1-2 hours per evaluation cycle.
Evidence-based referee verdict
Referee judges outputs against your stated quality bar (publication standard, formal proof, decision-grade, or custom) and surfaces specific evidence gaps rather than declaring a winner based on confidence. Project managers and content leads get actionable feedback on what each model got right or wrong.
Custom quality standards
Define your own acceptance criteria (e.g., 'every claim must be sourced', 'all assumptions stated', 'risks named and costed'). Operations and founders lock in quality gates that persist across all future evaluations, removing ambiguity from QA workflows.
Formal proof integration (theorem.chat)
For mathematical or technical claims, Referee formalizes the panel's best argument in Lean against Mathlib, where the kernel verifies correctness. Research teams working on proofs or formal specifications eliminate manual verification steps.
Decision-grade recommendation output
Generate structured recommendations with costed options and named risks, ready for stakeholder sign-off. Strategists and account executives compress the time from analysis to client-ready recommendation by 2-3 hours per engagement.
Private or public match storage
Free tier publishes matches publicly; Credit plan keeps work private with up to 6 seats and refundable unused credits. Project managers and operations teams choose privacy and collaboration scope based on client confidentiality needs.
What Makes Referee Different
Unique advantages vs similar tools in this niche
Evidence-based judging over confidence-based
vs Traditional AI evaluation tools that rely on model confidence scoresThe referee rules on evidence, not on who sounded surer, ensuring objective quality assessment.
Formal proof verification in Lean
vs Manual proof checking or informal verificationFor mathematical claims, results are formalized in Lean against Mathlib, where the kernel decides correctness.
Customizable quality bars
vs One-size-fits-all evaluation criteriaUsers can set standards from 'Careful expert' to 'Prize standard' or write their own, tailoring evaluation to specific needs.
Value Equation
Outcome-likelihood-time-effort assessment for Referee
Limited agency channel
Referee scored below the agency-resellability threshold (agency_fit_score < 50). The Value Equation projects agency-side outcomes, which don't apply to tools without a clear resell pathway.
Contact RefereePricing
Referee platform cost to your agency
Starts at $1 one-time (Free), scales to $10/mo (Credit)
Free
- Enough for a real match
- Every model and every tool
- Full transcript, evidence and verdict
- Matches are published, not private
Credit
- Private matches, your work stays yours
- Any model, any panel size, up to 6 seats
- Longer runs and higher spending limits
- Resume and branch without limits
Your own key
- Bring your own inference key
- You pay your provider directly
- The first $5.00 of usage costs you nothing here
- Private matches included
No verified white-label program for Referee: client-facing delivery runs under the platform's native branding.
Market Intelligence
Offer + scale economics for Referee
Limited agency channel
Referee scored below the agency-resellability threshold (agency_fit_score < 50). It's a useful tool but not designed for white-labeled or retainer-based reselling, so we don't publish productized offer economics for it.
Contact RefereeInvestment Decision Framework
Strategic vetting analysis for Referee
Situational Fit
Fit depends on your client mix
Buy If
5Your strategists and researchers spend 3+ hours per week manually reviewing AI-generated analysis for factual accuracy, sourcing, and logical gaps before client delivery. Referee automates that review by running multiple models against a publication-standard bar and surfacing evidence-backed verdicts.
Your project managers need to formalize AI-generated recommendations (e.g., technical architecture, budget scenarios, risk assessments) into decision-grade outputs that stakeholders can act on with confidence. Referee costed options and named risks into a single verdict.
Your content and research teams work across multiple LLM providers and need a neutral comparison framework to select which model output to use for a given task. Referee runs the same prompt across your entire panel and ranks results by evidence, not by which vendor sounds most confident.
Your operations or founder role is building internal AI workflows (research summaries, proposal drafts, technical specs) and needs a quality gate that doesn't rely on human spot-checking every output. Referee holds a literal bar and flags when the panel falls short.
Your team formalizes mathematical proofs or technical arguments and needs to validate them against a formal standard. Referee integrates with theorem.chat to formalize results in Lean against Mathlib, where the kernel decides correctness.
Skip If
5Your agency rarely generates AI outputs that require publication-grade or decision-grade validation. If most AI use is exploratory or internal brainstorming, Referee's overhead will not pay back.
Your team works with a single LLM provider and has no need to compare outputs across models. Referee's core value is multi-model evaluation; single-vendor shops get less ROI.
Your quality-assurance process is already lightweight and manual review takes less than 2 hours per week across the team. The cost per seat and setup friction will exceed the time savings.
Your agency cannot define explicit quality standards for the work being evaluated. Referee requires you to articulate what 'done' looks like (publication standard, formal proof, decision-grade) before running a match; teams without clear acceptance criteria will struggle.
Your team needs real-time feedback on AI outputs during client calls or live presentations. Referee is a batch-evaluation tool; it does not provide in-the-moment guidance or streaming verdicts.
Bottom Line
Referee.chat runs multiple LLM models against user-defined quality standards and judges outputs based on evidence rather than model confidence, enabling agencies to validate AI-generated work before publication or client delivery. Best suited for research teams, content verification workflows, and decision-support consultancies that need to formalize AI outputs to publication, formal-proof, or decision-grade standards. Integrates with 40+ model providers including Claude, GPT, Gemini, and Mistral, allowing teams to compare outputs across vendors without vendor lock-in. Agencies adopting Referee internally compress quality-assurance cycles for AI-generated research, analysis, and recommendations by replacing subjective review with evidence-based verdicts.
Reality Check
Referee requires teams to define explicit quality bars upfront (publication standard, formal proof, decision-grade, or custom), which adds initial setup friction. Adoption ROI concentrates in agencies running 5+ AI evaluation cycles per week; smaller teams may find the overhead outweighs the payoff.
Moderate effort: standard configuration with some customization needed
Academy for Referee
Work through it in order: the course for this service first, then the modules behind it.
Course for this service
Referee Agency Implementation, Monetizing AI Quality Assurance
Learn how to package Referee's multi-model evaluation and evidence-based verdicts into retainer services for clients who need reliable AI outputs. This course covers setting up custom quality standards, running competitive model panels, delivering decision-grade recommendations, and building recurring revenue from AI quality assurance.
Open the courseNo Academy modules are published for this service yet. Browse the full Academy
Core concepts
The mental model you need to price and scope the work.
- Eval Debt CompoundingConcept
Eval debt is the accumulated gap between what an AI agent does in production and what anyone on the agency team can actually prove it does. Like technical debt, it accrues quietly and charges interest: every untraced failure mode, every scoring rubric that lives in a Slack thread, every client demo that worked once and was never re-run. The interest payment arrives as a retainer conversation. Agencies that instrument early convert that debt into a premium line item, because "production-ready" is a claim only evidence can support. The cost curve is moving in their favor: OpenAI cut GPT-6 Sol and Luna API prices 50% versus GPT-5.6, and prompt caching now discounts up to 90% on reused prefixes, so high-volume agent pipelines are cheaper to run and cheaper to trace. Meanwhile Forrester's 2027 predictions flag compute and infrastructure constraints that will push API-dependent tool costs upward, compressing margins on AI-inclusive retainers. Tracing spend is the hedge.
- Silent Failure SurfaceConcept
The Silent Failure Surface is the set of AI behaviors that pass every automated check yet still damage the client relationship: a voice agent that interrupts callers, a support bot that loops a user through three retries, a research agent that returns confident but stale answers. Standard evals score outputs against expected answers, so they miss friction that only appears in live sessions. Agencies that map this surface before launch can price a monitoring retainer against it; agencies that skip it discover failures when the client forwards a complaint. Cekura simulates thousands of personas to expose interruption and gibberish patterns before go-live, while Agnost AI ingests real conversations and flags repeated retries and broken workflows as actionable intents. Both approaches treat production traffic as the primary test set, not a post-launch afterthought. The surface shrinks only when someone owns the loop between detection and a shipped fix.
- Trace-to-Trust RatioConcept
Trace-to-Trust Ratio is the proportion of an AI agent's production behavior that is actually instrumented, logged, and reviewable, measured against the trust a client extends to that system. Agencies that instrument every LLM call, tool invocation, and retrieval step can show clients exactly what happened when an output went wrong, which converts a vague reliability claim into a defensible audit trail. The ratio matters because trust is not granted by model choice; it is granted by evidence. A voice agent handling inbound calls with no tracing is a liability, while one instrumented through a platform like Cekura or Langfuse can surface interruption rates, gibberish detection, and latency per session. When a client asks why a response was wrong, the agency with trace coverage answers in minutes; the agency without it answers with a guess. That gap is where retainer renewals and premium pricing are decided.
Decision and risk
How to judge the fit, and the ways it goes wrong.
- When AI Output Quality Is Contested, Instrument Before You ArgueEvaluation Rule
Instrument the AI workflow with tracing and scoring before you defend its output quality to a client.
- AI Evaluation Rule: Price the Model Swap Before You Ship ItEvaluation Rule
Re-run the client's own evaluation set against the candidate model before migrating, and only swap when quality holds at the same or better score and the cost delta is documented.
- Evaluation Pipeline Before Launch vs Observability Retrofitted After Client EscalationDecision Framework
IF an agency is shipping LLM features into a client retainer and cannot currently answer 'what did the agent do on turn 14 of last Tuesday's session', THEN build the tracing and scoring layer before the next release, not after the first incident. IF the AI work is still internal tooling with no client-facing output or contractual quality bar, THEN defer the spend and revisit when a client name attaches to the output.
- The Demo-Data Trap: Why AI Evaluation and Observability Stalls After the PilotFailure Pattern
- The Judge-Only Trap: Why AI Evaluation and Observability Collapses Under Client ScrutinyFailure Pattern
Delivery system
Blueprints and procedures for running it as a service.
- Production AI Readiness Audit (7-12 days)Implementation Blueprint
A fixed-scope diagnostic that instruments a client's live AI feature with tracing, scoring, and drift checks, then hands over a scored reliability report the agency can bill against. It converts an unmonitored pilot into a supportable retainer line.
- Eval Baseline Before Client AI Go-Live (Onboarding)Operating Procedure
- Trace Coverage Audit Before Retainer Renewal (Retention)Operating Procedure
- Production Failure Triage and Fix Loop (QA)Operating Procedure
13 modules selected for Referee
Frequently Asked Questions
Answers about pricing, setup, implementation
Referee.chat runs multiple LLM models against a user-defined quality standard and judges the outputs based on evidence rather than model confidence. You set the bar (publication standard, formal proof, decision-grade, or custom), seat a panel of models, and Referee tells you which outputs meet the bar and why. For mathematical claims, it integrates with theorem.chat to formalize proofs in Lean against Mathlib.
The Free plan costs $1 one-time and includes every model and full transcripts, but matches are published publicly. The Credit plan costs $10 per month per seat (up to 6 seats), keeps matches private, and includes refundable unused credits and higher spending limits. You can also bring your own inference key and pay your LLM provider directly; Referee charges $0.05 per unit of infrastructure after the first $5.00 of usage, with no monthly seat fee.
Strategists and researchers compress quality-assurance cycles by running multi-model evaluations against publication or formal-proof standards. Project managers formalize AI-generated recommendations into decision-grade outputs with costed options and named risks. Content leads and operations teams use Referee as a quality gate to validate AI outputs before client delivery, replacing manual spot-checking with evidence-backed verdicts. Founders building internal AI workflows use Referee to lock in quality standards that persist across all future evaluations.
A strategist or researcher running 5+ AI evaluation cycles per week typically saves 3-5 hours per week by replacing manual multi-model comparison and quality review with a single Referee match. A project manager formalizing 2-3 AI-generated recommendations per week saves 2-3 hours by automating the evidence-gathering and risk-naming step. Payback depends on baseline QA time; teams spending less than 2 hours per week on AI review will see minimal ROI.
Referee integrates with 40+ model providers including Anthropic Claude, OpenAI GPT (including o1, o3, GPT-5 series), Google Gemini, Mistral, DeepSeek, Meta Llama, Cohere, AI21, Amazon Nova, and many others. You can also bring your own inference key and use Referee with your existing vendor relationships. The full model list is available on the Referee.chat platform.
Referee does not publish pricing or data-retention terms in the extracted content. Contact Referee.chat support for details on data export, deletion, and retention policies after cancellation.
Initial setup takes 30-60 minutes per team: define your quality standards (publication, formal proof, decision-grade, or custom), select your model panel, and run a test match. Rollout to the full team is low-friction; each team member can start running matches immediately after onboarding. No infrastructure changes or API integrations required.
Referee automates the evidence-gathering and verdict step of QA, but does not replace human judgment on whether the quality bar itself is appropriate for a given task. Use Referee to formalize and speed up QA workflows, not to eliminate human review entirely. Teams typically use Referee to flag outputs that fall short of the bar, then route those outputs back to the model panel or to human review.