Leibler
Kullback is an open-source testing harness that reconstructs AI agent environments from execution traces. It reads logs your agent already writes, rebuilds the tools and data the agent used, replays the logs to verify the rebuild is accurate, and generates per-task verifiers that check final data state for pass or fail. Every reconstructed value is traceable back to the original trace line. The framework is designed for agencies building custom AI agents that need reproducible, deterministic validation before client deployment.
Leibler is an open-source testing harness. InnovaAI scores it 2.5/10 for agency adoption, best for AI Development Lead, Project Manager, and QA Engineer roles handling 5+ client meetings per week.
Agency Audit
Kullback is an open-source framework that reconstructs AI agent tools, data, and rules from execution traces, enabling agencies to verify agent behavior without manual transcript review. Teams building or deploying custom AI agents can use it to debug failures, validate task execution, and generate pass/fail reports tied to final data state rather than agent reasoning. Best suited for AI evaluation teams and agencies developing production AI agents where reproducible testing and deterministic verification are critical to deployment confidence.
3recommended
54/mo
No paid plan published
High
Illustrative scenario. Not a guarantee. Net capacity needs a verified paid base plan, and none is published for this service, so it is not modeled. Hours saved come from the service estimate; implementation, taxes, and unprovided usage charges are excluded.
- AI Development Lead handling agent execution validation
- Project Manager handling failure mode debugging
- QA Engineer handling pre-deployment verification
- Your agency does not build or deploy AI agents as a core service. Kullback is purpose-built for agent development and testing; it adds no value to traditional digital agency workflows like design, copywriting, or client strategy.
- Your team runs fewer than two agent projects per quarter. The overhead of maintaining execution traces and integrating Kullback into your CI/CD pipeline is not justified by infrequent deployments.
- Your AI agents are third-party tools (e.g., off-the-shelf LLM APIs or SaaS agent platforms) that you do not modify or control. Kullback requires access to execution traces from agents you build; it cannot reconstruct behavior from external black-box systems.
Internal Adoption Path
No paid plan published
54 hr/mo
3 seats × 18 hr each
$4,050/mo
modeled at $75/hr labor rate
No paid plan published
Illustrative scenario. Not a guarantee. No verified paid base plan is published for this service, so subscription cost and net capacity are not modeled. Implementation, taxes, and unprovided usage charges are excluded.
Platform Features
Core capabilities of Leibler
Trace-to-environment reconstruction
Kullback reads execution logs from your agent runs and automatically rebuilds the tools, data, and rules the agent used. Your QA or development team no longer manually transcribes agent behavior; the framework extracts it directly from traces.
Replay and verification
Once reconstructed, Kullback replays the logs against the rebuilt environment to confirm the rebuild matches the original execution. Your team gains confidence that the test environment is faithful before running new agent models or validating client deployments.
Per-task verifier generation
Kullback generates one verifier per task that checks final data state to determine pass or fail. Your Project Manager or QA lead no longer reads transcripts; code-driven verdicts replace subjective judgment.
Execution report generation
Kullback produces structured reports showing which runs passed or failed, with every value linked back to the trace line it came from. Your Operations or Founder team gets visibility into agent reliability without manual log review.
Simulated user context
Kullback creates a simulated user that knows only what the real user knew during the original run. Your development team can test whether agents behave correctly when user knowledge is limited, catching over-assumption bugs before client deployment.
Open-source transparency
All code, design decisions, and measured validation results are public on GitHub under Apache-2.0. Your team can audit the framework, contribute fixes, and avoid vendor lock-in on a critical testing tool.
What Makes Leibler Different
Unique advantages vs similar tools in this niche
Reconstructs the exact environment from traces rather than relying on manual transcript review
vs Manual transcript review or simple loggingKullback rebuilds tools, data, and rules from the logs the agent already writes, ensuring nothing is invented.
Uses final data for pass/fail decisions instead of transcript content
vs Transcript-based evaluation methodsCode decides pass or fail from the final data, not from the transcript, reducing subjectivity.
Provides full transparency with public code and design decisions
vs Closed-source evaluation toolsAll code, design philosophy, and decision logs are public, allowing for community review and contribution.
Value Equation
Outcome-likelihood-time-effort assessment for Leibler
Value math requires real pricing
The Value Equation (dream outcome × likelihood ÷ time × effort) feeds directly into ROI math. Leibler has no published pricing, so we hold this section until real numbers are available.
Contact LeiblerPricing
Pricing data not yet available for Leibler.
Reality Check
Kullback requires agencies to maintain detailed execution traces and integrate trace-reading into their agent development pipeline. Adoption payoff is highest for teams running 5+ agent deployments per month; smaller or one-off agent projects may not justify the infrastructure investment.
High effort: requires technical configuration and team training
How This Accelerates White-Label Services
Who It's For
- ✓ai-agent-development-agencies
- ✓ai-evaluation-and-testing-teams
- ✓agencies-building-custom-ai-agents
Acceleration Steps
- 1Schedule onboarding with the vendor
- 2Configure rebuild agent tools, data, and rules from execution traces
- 3Launch your first client project
Academy for Leibler
Work through it in order: the course for this service first, then the modules behind it.
Course for this service
Kullback Agency Implementation, Reproducible AI Agent Testing
Learn how to set up Kullback's trace-to-environment reconstruction and per-task verifiers to validate custom AI agents before client deployment. This course teaches agencies how to replace manual QA transcription with code-driven verdicts, reduce validation cycles, and deliver deterministic test reports that prove agent behavior matches production requirements.
Open the courseNo Academy modules are published for this service yet. Browse the full Academy
Core concepts
The mental model you need to price and scope the work.
- Eval Debt CompoundingConcept
Eval debt is the accumulated gap between what an AI agent does in production and what anyone on the agency team can actually prove it does. Like technical debt, it accrues quietly and charges interest: every untraced failure mode, every scoring rubric that lives in a Slack thread, every client demo that worked once and was never re-run. The interest payment arrives as a retainer conversation. Agencies that instrument early convert that debt into a premium line item, because "production-ready" is a claim only evidence can support. The cost curve is moving in their favor: OpenAI cut GPT-6 Sol and Luna API prices 50% versus GPT-5.6, and prompt caching now discounts up to 90% on reused prefixes, so high-volume agent pipelines are cheaper to run and cheaper to trace. Meanwhile Forrester's 2027 predictions flag compute and infrastructure constraints that will push API-dependent tool costs upward, compressing margins on AI-inclusive retainers. Tracing spend is the hedge.
- Silent Failure SurfaceConcept
The Silent Failure Surface is the set of AI behaviors that pass every automated check yet still damage the client relationship: a voice agent that interrupts callers, a support bot that loops a user through three retries, a research agent that returns confident but stale answers. Standard evals score outputs against expected answers, so they miss friction that only appears in live sessions. Agencies that map this surface before launch can price a monitoring retainer against it; agencies that skip it discover failures when the client forwards a complaint. Cekura simulates thousands of personas to expose interruption and gibberish patterns before go-live, while Agnost AI ingests real conversations and flags repeated retries and broken workflows as actionable intents. Both approaches treat production traffic as the primary test set, not a post-launch afterthought. The surface shrinks only when someone owns the loop between detection and a shipped fix.
- Trace-to-Trust RatioConcept
Trace-to-Trust Ratio is the proportion of an AI agent's production behavior that is actually instrumented, logged, and reviewable, measured against the trust a client extends to that system. Agencies that instrument every LLM call, tool invocation, and retrieval step can show clients exactly what happened when an output went wrong, which converts a vague reliability claim into a defensible audit trail. The ratio matters because trust is not granted by model choice; it is granted by evidence. A voice agent handling inbound calls with no tracing is a liability, while one instrumented through a platform like Cekura or Langfuse can surface interruption rates, gibberish detection, and latency per session. When a client asks why a response was wrong, the agency with trace coverage answers in minutes; the agency without it answers with a guess. That gap is where retainer renewals and premium pricing are decided.
Decision and risk
How to judge the fit, and the ways it goes wrong.
- When AI Output Quality Is Contested, Instrument Before You ArgueEvaluation Rule
Instrument the AI workflow with tracing and scoring before you defend its output quality to a client.
- AI Evaluation Rule: Price the Model Swap Before You Ship ItEvaluation Rule
Re-run the client's own evaluation set against the candidate model before migrating, and only swap when quality holds at the same or better score and the cost delta is documented.
- Evaluation Pipeline Before Launch vs Observability Retrofitted After Client EscalationDecision Framework
IF an agency is shipping LLM features into a client retainer and cannot currently answer 'what did the agent do on turn 14 of last Tuesday's session', THEN build the tracing and scoring layer before the next release, not after the first incident. IF the AI work is still internal tooling with no client-facing output or contractual quality bar, THEN defer the spend and revisit when a client name attaches to the output.
- The Demo-Data Trap: Why AI Evaluation and Observability Stalls After the PilotFailure Pattern
- The Judge-Only Trap: Why AI Evaluation and Observability Collapses Under Client ScrutinyFailure Pattern
Delivery system
Blueprints and procedures for running it as a service.
- Production AI Readiness Audit (7-12 days)Implementation Blueprint
A fixed-scope diagnostic that instruments a client's live AI feature with tracing, scoring, and drift checks, then hands over a scored reliability report the agency can bill against. It converts an unmonitored pilot into a supportable retainer line.
- Eval Baseline Before Client AI Go-Live (Onboarding)Operating Procedure
- Trace Coverage Audit Before Retainer Renewal (Retention)Operating Procedure
- Production Failure Triage and Fix Loop (QA)Operating Procedure
13 modules selected for Leibler
Frequently Asked Questions
Answers about pricing, setup
Kullback reads execution traces from your AI agents and reconstructs the tools, data, and rules they used during real runs. It then replays those traces to verify the rebuild is accurate, generates per-task verifiers that check final data for pass/fail status, and produces reports showing which runs succeeded. Your team validates agent behavior without manual transcript review.
Kullback is open-source and free to use. There is no per-seat pricing or subscription fee. Your team hosts and runs it on your own infrastructure.
AI development leads and engineers use Kullback to debug agent failures and validate task execution. Project Managers and QA leads use it to generate pass/fail reports and track agent reliability across deployments. Founders and Operations leads use it to verify that deployed agents behave consistently before client handoff. Best suited for agencies building custom AI agents as a core service.
For an AI development team running 5+ agent deployments per month, Kullback saves approximately 4-6 hours per week by eliminating manual trace review and transcript-based debugging. Actual savings depend on the number of agent runs per week and the complexity of your verification logic. Teams with fewer deployments see lower absolute time savings but higher per-project ROI.
Kullback is framework-agnostic and works with any agent system that produces execution traces. Your team must integrate trace collection into your agent codebase, then point Kullback at those traces. Setup time is typically 2-4 weeks for a team new to trace-based testing.
No. Kullback requires access to detailed execution traces from agents you control. It cannot reconstruct behavior from third-party SaaS agent platforms or black-box LLM APIs that do not expose their internal logs.