Running Leibler as a service, AI Evaluation Observability
Kullback Agency Implementation, Reproducible AI Agent Testing
Learn how to set up Kullback's trace-to-environment reconstruction and per-task verifiers to validate custom AI agents before client deployment. This course teaches agencies how to replace manual QA transcription with code-driven verdicts, reduce validation cycles, and deliver deterministic test reports that prove agent behavior matches production requirements.
Open the decision record for LeiblerWhat does running Leibler for clients commit you to?
Published figures for this service. Blank fields are not published.
- Monthly tool cost
- Not published. Kullback is open-source (pricing_model: open-source) and no pricing tiers were extracted in the Level 1 data, so vendor cost basis and setup costs cannot be stated.
- Time to first value
- Not published
- Payback
- Not modeled
- Guided implementation
- 8 hours
Is Leibler worth running as a client service?
Kullback is an open-source, high-setup framework for validating AI agent behavior from execution traces, suited to agencies with technical staff. Because no pricing tiers, time-to-value duration, or market rates were published in the extracted data, vendor cost basis and ROI cannot be evidenced.
An agency-fit judgement for reselling this service. It is separate from the tool description on the decision record.
Before you start
What has to be in place before the first client engagement.
Tools and subscriptions
- Access to AI agent execution traces (required input per Level 1 data)
- A technical environment capable of running an open-source framework, since setup_complexity is high
- Judge model access for verifier generation and pass/fail verdicts
- Agent tools and data sources that can be reconstructed from traces
- Reporting capability to deliver pass/fail reports and flag runs where a model stood in for a missing tool
People and inputs
- Engineering time to rebuild tools, data, and rules from traces (core function)
- Process for replaying logs and confirming rebuilt environment matches original (core function)
- Verifier generation workflow using final data rather than transcripts (platform feature)
- Client trace intake and privacy handling plan, since traces are required inputs
Included with the course
6 working documents for delivering this service.
- Trace Log Collection Checklist for AI Agentschecklist
- Environment Reconstruction SOPsop
- Per-Task Verifier Templatetemplate
- Agent Validation Report Worksheetworksheet
- Replay Verification Runbookguide
- Client Deployment Readiness Checklistchecklist
Listed by name. These documents are not yet published as individual downloads.