Verbara.Sdk.VoiceAi 2.8.0

dotnet add package Verbara.Sdk.VoiceAi --version 2.8.0
                    
NuGet\Install-Package Verbara.Sdk.VoiceAi -Version 2.8.0
                    
This command is intended to be used within the Package Manager Console in Visual Studio, as it uses the NuGet module's version of Install-Package.
<PackageReference Include="Verbara.Sdk.VoiceAi" Version="2.8.0" />
                    
For projects that support PackageReference, copy this XML node into the project file to reference the package.
<PackageVersion Include="Verbara.Sdk.VoiceAi" Version="2.8.0" />
                    
Directory.Packages.props
<PackageReference Include="Verbara.Sdk.VoiceAi" />
                    
Project file
For projects that support Central Package Management (CPM), copy this XML node into the solution Directory.Packages.props file to version the package.
paket add Verbara.Sdk.VoiceAi --version 2.8.0
                    
#r "nuget: Verbara.Sdk.VoiceAi, 2.8.0"
                    
#r directive can be used in F# Interactive and Polyglot Notebooks. Copy this into the interactive tool or source code of the script to reference the package.
#:package Verbara.Sdk.VoiceAi@2.8.0
                    
#:package directive can be used in C# file-based apps starting in .NET 10 preview 4. Copy this into a .cs file before any lines of code to reference the package.
#addin nuget:?package=Verbara.Sdk.VoiceAi&version=2.8.0
                    
Install as a Cake Addin
#tool nuget:?package=Verbara.Sdk.VoiceAi&version=2.8.0
                    
Install as a Cake Tool

Verbara.Sdk.VoiceAi

Voice AI pipeline for Verbara.Sdk — orchestration layer for STT, TTS, and conversation with turn-taking and barge-in detection.

Installation

dotnet add package Verbara.Sdk.VoiceAi

Quick Start

// Implement your conversation handler (called once per user utterance)
public class MyHandler : IConversationHandler
{
    public ValueTask<string> HandleAsync(
        string transcript, ConversationContext context, CancellationToken ct)
    {
        return ValueTask.FromResult($"You said: {transcript}");
    }
}

// Register in DI
services.AddAudioSocketServer(opts => opts.Port = 9092);
services.AddVoiceAiPipeline<MyHandler>(opts =>
{
    opts.InputFormat = AudioFormat.Slin16Mono8kHz;
    opts.OutputFormat = AudioFormat.Slin16Mono8kHz;
    opts.EndOfUtteranceSilence = TimeSpan.FromMilliseconds(600);
});

// Subscribe to pipeline events
var pipeline = app.Services.GetRequiredService<VoiceAiPipeline>();
pipeline.Events.Subscribe(evt => Console.WriteLine(evt));

Features

  • VoiceAiPipeline — full VAD → STT → IConversationHandler → TTS loop per AudioSocket session
  • Barge-in detection: cancels TTS playback when the caller speaks
  • IConversationHandler — scoped per session; implement to plug in any LLM or business logic
  • ISessionHandler — low-level interface; implement for fully custom session handling
  • VoiceAiSessionBroker — hosted service that hands each AudioSocketSession to the active ISessionHandler, once
    • The session ends when its handler does. As soon as HandleSessionAsync returns, throws or is cancelled, the broker calls the session's HangupAsync: one hangup frame if the line is still live, then the close. On Asterisk 20 and later the call goes on in the dialplan after AudioSocket(); an ARI externalMedia channel leaves Stasis with the caller still in your bridge. A handler keeps its line exactly as long as it runs, so await the conversation rather than leaving it to a background task. A session the caller or the handler already ended gets nothing more, and one the broker ends while still live is logged once at Information.
    • A graceful stop ends the calls, then waits. While the token passed to StopAsync is not cancelled, the stop ends every live session it handed out with a hangup frame and waits for the handlers to return; it does not cancel their token. Once that token is cancelled (the host's shutdown budget ran out), the stop cancels the handlers' token and returns without waiting further. A direct StopAsync(CancellationToken.None) therefore waits until every handler has returned: pass a token you cancel after a bound of your own.
    • Disposal does not wait. Dispose cancels the handlers' token and returns.
    • The handler's CancellationToken belongs to the broker: only a stop that is no longer graceful, or Dispose, cancels it. Once its stop has been called, or once it is disposed, the broker hands no new session on; the AudioSocket server keeps such a session and releases it when it stops. The broker is not restartable: a start after a stop subscribes nothing, a start after Dispose throws ObjectDisposedException, and a stop after Dispose does nothing.
    • Register AddAudioSocketServer before AddVoiceAiPipeline, as above: the host stops services in reverse order, so the broker then stops first and every call ends with a hangup frame. A server stopped first closes its sessions without one.
  • Observable Events stream (SpeechStartedEvent, TranscriptReceivedEvent, BargInDetectedEvent, etc.)
  • Native AOT compatible

Custom STT / TTS Providers

When writing your own SpeechRecognizer or SpeechSynthesizer subclass, override ProviderName with a stable literal to avoid the per-utterance GetType().Name allocation on the pipeline hot path (used as a tag on STT/TTS activities):

public sealed class MyCustomRecognizer : SpeechRecognizer
{
    public override string ProviderName => "MyCustom";

    public override IAsyncEnumerable<SpeechRecognitionResult> StreamAsync(
        IAsyncEnumerable<ReadOnlyMemory<byte>> audioFrames,
        AudioFormat format,
        CancellationToken ct = default)
    {
        // ...
    }
}

If you don't override ProviderName the default falls back to GetType().Name — functional, but incurs one reflection call per utterance.

Observability

  • Metrics: VoiceAiMetrics (sessions started/completed/failed, session duration), SpeechRecognitionMetrics (transcriptions started/completed/failed/cancelled, latency), SpeechSynthesisMetrics (syntheses started/completed/failed/cancelled/silent, latency, characters). Every recognition and synthesis the pipeline starts ends in exactly one of completed, failed or cancelled; cancelled is tagged voiceai.ending with what cut it short (session-cancelled, and for syntheses also barge-in, disposal, far-end), and tts.syntheses.completed counts only a synthesis whose audio was all written.
  • Tracing: VoiceAiActivitySource — session / recognition / synthesis spans.
  • Health: VoiceAiHealthCheck, SttHealthCheck, TtsHealthCheck auto-registered by AddVoiceAiPipeline<THandler>().
  • Discover names via VerbaraTelemetry.ActivitySourceNames / MeterNames from Verbara.Sdk.Hosting.

Documentation

See the main README for full documentation. Upgrading from 2.6.1: voice-session-ending-migration.md.

Product Compatible and additional computed target framework versions.
.NET net10.0 is compatible.  net10.0-android was computed.  net10.0-browser was computed.  net10.0-ios was computed.  net10.0-maccatalyst was computed.  net10.0-macos was computed.  net10.0-tvos was computed.  net10.0-windows was computed. 
Compatible target framework(s)
Included target framework(s) (in package)
Learn more about Target Frameworks and .NET Standard.

NuGet packages (5)

Showing the top 5 NuGet packages that depend on Verbara.Sdk.VoiceAi:

Package Downloads
Verbara.Sdk.VoiceAi.Stt

STT providers for Verbara.Sdk.VoiceAi — Deepgram, Cartesia, AssemblyAI (WebSocket streaming), OpenAI Whisper, Azure Whisper, Google Speech (REST batch). Zero third-party dependencies.

Verbara.Sdk.VoiceAi.Tts

TTS providers for Verbara.Sdk.VoiceAi — ElevenLabs (WebSocket streaming) and Azure TTS (REST). Zero third-party dependencies.

Verbara.Sdk.VoiceAi.Testing

Test fakes for Verbara.Sdk.VoiceAi -- FakeSpeechRecognizer, FakeSpeechSynthesizer, FakeConversationHandler for testing Voice AI apps without API keys.

Verbara.Sdk.VoiceAi.OpenAiRealtime

OpenAI Realtime API bridge for Verbara.Sdk.VoiceAi — connects Asterisk AudioSocket directly to an OpenAI Realtime model, with function calling and observability. Zero third-party dependencies.

Verbara.Sdk.VoiceAi.TurnDetection

ML-based turn detection for Verbara.Sdk VoiceAi pipeline using the Pipecat smart-turn-v3.2-cpu ONNX model.

GitHub repositories

This package is not used by any popular GitHub repositories.

Version Downloads Last Updated
2.8.0 0 10/9/2026
2.7.0 549 10/3/2026
2.6.1 498 9/30/2026
2.6.0 741 9/24/2026
2.5.3 623 9/13/2026
2.5.2 298 9/13/2026
2.5.1 308 9/12/2026
2.5.0 379 8/25/2026
2.4.0 1,978 7/27/2026
2.3.2 679 7/20/2026
2.3.1 306 7/14/2026
2.3.0 428 7/6/2026
2.2.1 881 5/23/2026
2.2.0 188 5/20/2026
2.1.2 192 5/8/2026
2.1.1 171 5/7/2026
2.1.0 3,442 5/7/2026
2.0.0 170 5/6/2026