Running Kyutai as a service, AI Infrastructure

Kyutai Agency Implementation, Building Voice-Native Client Products

Learn how to integrate Kyutai's open-source speech models into client deliverables, from deploying Moshi for real-time dialogue to adding voice layers with Unmute and scaling multilingual projects with Hibiki-Zero. This course teaches agencies how to architect voice-first solutions, manage model dependencies in production, and price voice-native services as retainers.

Open the decision record for Kyutai

What does running Kyutai for clients commit you to?

Published figures for this service. Blank fields are not published.

Monthly tool cost
Not published, Kyutai is open-source with no paid tiers. Vendor cost basis is $0 for model licenses; agency investment is engineering labor and hosting, which must be supplied by your agency.
Time to first value
Not published
Payback
Not modeled, client price, labor, usage, overhead, and expected volume are not supplied
Guided implementation
8 hours

Is Kyutai worth running as a client service?

Kyutai offers free open-source multimodal models that agencies with engineering capability can integrate as a managed service, but it is not a turnkey or white-label product. Without published pricing or ROI inputs, agencies must gather their own cost-to-serve and client-pricing data before modeling returns.

An agency-fit judgement for reselling this service. It is separate from the tool description on the decision record.

Before you start

What has to be in place before the first client engagement.

Tools and subscriptions

  • GitHub access for Kyutai open-source repositories
  • Hugging Face account for model weights and datasets
  • API/SDK integration capability for speech and vision models
  • Development environment for real-time audio streaming (Moshi, Unmute, Hibiki-Zero)
  • Version control and CI/CD for maintaining forked or integrated model code

People and inputs

  • Engineers experienced with open-source model deployment and inference optimization
  • Audio and video streaming infrastructure for testing MoshiVis, CASA, and speech-native models
  • Access to Kyutai scientific publications and tutorials for model-specific implementation guidance
  • GPU or CPU inference resources appropriate for chosen models (e.g., Pocket TTS runs on CPU faster than real-time)

Included with the course

7 working documents for delivering this service.

  • Kyutai Model Selection Worksheet for Client Scopingworksheet
  • Voice-Native Project Delivery SOP: Moshi to Productionsop
  • Kyutai Infrastructure Cost Calculator and Margin Templatetemplate
  • Speech Model Integration Checklist: Unmute, Hibiki-Zero, Pocket TTSchecklist
  • Multilingual Voice Project Scope Guide: Hibiki-Zero Editionguide
  • Client Handoff Documentation Template for Self-Hosted Modelstemplate
  • Voice Feature Pricing Playbook: Retainer Models for Speech Servicesguide

Listed by name. These documents are not yet published as individual downloads.