Microsoft SAPI TTS Engine example for Govornik. https://govornik.eu/prenos/#SAPI
  • C++ 93.8%
  • Batchfile 6.2%
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
2026-08-23 17:58:36 +02:00
src Update GovornikEngine.cpp 2026-08-23 17:58:36 +02:00
.gitignore First commit 2026-08-23 17:56:57 +02:00
build.bat First commit 2026-08-23 17:56:57 +02:00
LICENSE First commit 2026-08-23 17:56:57 +02:00
README.md First commit 2026-08-23 17:56:57 +02:00

Minimal SAPI 5 TTS voice — a worked example

A tiny, complete example of how to implement a Windows SAPI 5 text-to-speech voice in C++. When I set out to build one I couldn't find a simple, end-to-end sample anywhere — so this is that sample. It's backed by the free Slovenian Govornik synthesizer, but the SAPI parts are the point and apply to any backend.

The whole thing is ~300 lines across three source files.

What a SAPI voice actually is

A "voice" in Windows is not a file — it's a COM object plus some registry entries:

  1. An in-process COM class (a DLL) that implements two interfaces:
    • ISpTTSEngine — SAPI calls Speak() to turn text into audio, and GetOutputFormat() to learn your sample format.
    • ISpObjectWithToken — SAPI hands you your voice token so you can read per-voice settings (here: which Govornik voice to use).
  2. Registry entries that tell SAPI the class exists:
    • a CLSID under HKCR\CLSID\{…}\InprocServer32 → your DLL,
    • one voice token per voice under HKLM\SOFTWARE\Microsoft\Speech\Voices\Tokens\…, pointing at that CLSID.

When an app speaks with your voice, SAPI creates the object, calls SetObjectToken(), GetOutputFormat(), then Speak() with the text already parsed into plain-text fragments. Your job in Speak() is simply to produce PCM audio and hand it to the "site" SAPI gives you. That's it.

Files

src/GovornikEngine.cpp   the engine: ISpTTSEngine + ISpObjectWithToken,
                         the class factory, and the DLL exports  <- the lesson
src/GovornikEngine.def   exports DllGetClassObject / DllCanUnloadNow
src/Http.h / Http.cpp    tiny WinHTTP client for the Govornik API + WAV parsing
src/GovornikRegister.cpp the registrar exe: writes the CLSID + voice tokens,
                         and removes them again with --remove
build.bat                builds both, x64

No third-party dependencies — just the Windows SDK (WinHTTP + SAPI). The CRT is linked statically (/MT), so the outputs are self-contained.

Build

Requires Visual Studio 2022 (C++ workload) + the Windows 10/11 SDK. Then:

build.bat

Produces build\GovornikTTS.dll (the engine) and build\GovornikRegister.exe (the installer). Keep them together.

Install the voices

Run the registrar as Administrator (it prompts automatically — it writes to HKLM):

build\GovornikRegister.exe

It registers the engine's CLSID and fetches the current voice list from s1.govornik.eu, adding one voice per entry (named Govornik … (primer)).

Remove everything again:

build\GovornikRegister.exe --remove

Try it

Any SAPI app now lists the voices (Balabolka, NVDA's SAPI5 driver, Word's classic Read Aloud, Windows Narrator, …). Quick test from PowerShell:

Add-Type -AssemblyName System.Speech
$s = New-Object System.Speech.Synthesis.SpeechSynthesizer
$s.SelectVoice("Govornik Lars (primer)")
$s.Speak("Pozdravljen iz preprostega primera.")

How the code maps to the flow

  • CEngine::GetOutputFormat → declares 16 kHz / 16-bit / mono PCM (what Govornik returns; SAPI resamples to the device).
  • CEngine::SetObjectToken → stores the token; Speak reads GovornikVoice from it to know which voice to request.
  • CEngine::Speak → for each text fragment: GovornikSynth() POSTs to the server, ExtractWavePcm() finds the PCM, and site->Write() streams it to SAPI (checking GetActions() for an abort).
  • GovornikRegister.cpp → the registry side: Register() writes the CLSID and tokens; Remove() deletes them.

Deliberately left out (so the example stays readable)

This is the minimum that works. A production engine would typically add:

  • Word-boundary / bookmark events (site->AddEvents) so hosts highlight text as it's spoken.
  • Rate / pitch / volume handling (site->GetRate/GetVolume, or backend effects).
  • Server failover, retries, and audio caching for robustness/latency.
  • A 32-bit build. This sample is x64 and registers the 64-bit classic (…\Speech\Voices\Tokens) and OneCore (…\Speech_OneCore\Voices\Tokens) roots — enough for 64-bit classic SAPI apps and Windows Narrator. For 32-bit apps you'd also build an x86 DLL and mirror the tokens under …\WOW6432Node\….
  • An installer instead of a console registrar.

Each is a small addition on top of this skeleton.

Notes

  • Govornik is a free public API for Slovenian TTS. For real projects, register a free source id at https://govornik.eu/prijava-projekta/ and set it in Http.cpp (kSource). API docs: https://govornik.eu/docs/.
  • Modern WinRT (Windows.Media.SpeechSynthesis) / UWP apps do not load third-party SAPI engines — this works with classic SAPI apps and Narrator.

License

MIT — see LICENSE. Use it however you like.