- Implemented the /speak command in Discord bot to synthesize speech using the TTS gateway. - Added voice handling logic to join voice channels and play synthesized audio. - Created tests for the new command and voice functionalities. - Introduced TTSGateway interface for TTS service communication. - Updated configuration to include TTS gateway address. - Documented the TTS gateway integration and model artifact distribution process.
16 KiB
tts-gateway: plan & feasibility
Proposal for turning tmp/reference/uma-tts-api (Python/Flask, VITS voice model) into a fourth
home-services module, tts-gateway, written in Go, following the same hexagonal pattern as
ha-gateway/ai-gateway/discord-bot, deployed on nik-gpu (Nvidia RTX 2080 Super).
See tmp/reference/uma-tts-api/FINDINGS.md for how the
reference implementation works today; this document is about porting it.
TL;DR feasibility
| Piece | Verdict | Why |
|---|---|---|
| Serving/API layer (gRPC, config, health, telemetry) | Straightforward | Identical shape to the other three services; pure Go, no new tech. |
| Audio postprocessing (WAV → AAC/Opus) | Straightforward | Write PCM/WAV by hand, shell out to ffmpeg (already a required system dep). |
| Model inference (VITS forward pass) | Feasible, one open risk | No mature native-Go path to run a .pth checkpoint directly. The realistic route is: export the inference-only graph to ONNX once (offline, in Python), then run it in Go via onnxruntime-go with the CUDA execution provider. The 2080 Super (Turing, CC 7.5) is well inside ONNX Runtime's supported GPU range. The risk isn't "can Go do GPU inference" (yes), it's "does this specific checkpoint's graph export cleanly to ONNX" — see Risk 1. |
Japanese text front-end (pyopenjtalk g2p) |
Feasible, one open risk | No usable native-Go OpenJTalk/mecab g2p implementation exists. Shelling out to the open_jtalk CLI (a real C++ binary, not Python) from Go is the likely path, but its default output format needs to be checked against what pyopenjtalk.g2p() actually returns — see Risk 2. |
| Overall "100% Go, no Python at runtime" | Likely, not guaranteed | Contingent on Risks 1 and 2 resolving cleanly. Both have a concrete fallback (below) that still ships a real Go service, just with a small non-Go component for one stage. |
Bottom line: this is buildable and the GPU is not a blocker. The two things that could force a hybrid design instead of pure Go are named explicitly below, and Phase 0 exists specifically to answer them before any Go code is written.
What "parity" means
The new service should support what uma-tts-api supports today:
- Synthesize speech for one of 92 Umamusume speakers given Japanese text, with the same
noise_scale/noise_scale_w/length_scaleknobs. - List/search speakers.
- Health check.
Not required to carry over: training code, the monotonic_align Cython extension (only used by
the training loss, confirmed below — never called from infer()), or CPU fallback as a
first-class target (nik-gpu has a GPU; CPU-only would defeat the point of moving it there).
Why ONNX Runtime instead of a Go libtorch binding
Two ways exist to run a PyTorch model's math from Go:
- cgo bindings to libtorch (e.g. community "gotorch"-style projects). Rejected: these bind
directly to PyTorch's C++ ABI, which is unstable across versions, poorly maintained as Go
packages, and would require hand-porting the whole
SynthesizerTrnforward pass into Go/C++ rather than reusing a serialized graph. High effort, fragile, no real upside here. - Export to ONNX once, run via
onnxruntime-go(e.g.github.com/yalue/onnxruntime_go), using Microsoft's prebuiltonnxruntime-linux-x64-gpushared library with the CUDA execution provider. This is the standard, maintained path for "run a PyTorch model from a non-Python language," and it converts the problem from "port a neural net to Go" into "export a graph once, then call a stable C API." Recommended.
Risk 1: ONNX export of the stochastic duration predictor
Checked in the reference code: configs/uma.json's model section doesn't set use_sdp, and
SynthesizerTrn.__init__ defaults use_sdp=True (models/models.py:413), so this checkpoint
does use the flow-based StochasticDurationPredictor, not the simpler deterministic one.
That module's reverse (inference) path chains several ConvFlows, which call into
transforms.py's rational_quadratic_spline / unconstrained_rational_quadratic_spline. Those
functions use boolean-mask indexed assignment (outputs[outside_interval_mask] = ...,
transforms.py:68-95) to handle spline boundary conditions. This pattern is a known rough edge
for torch.onnx.export's tracer — it's data-dependent control flow that doesn't always lower to
a single static ONNX graph cleanly, unlike the rest of the model (convolutions, attention,
LayerNorm, ConvTranspose1d) which are all completely standard and export fine.
The custom torch.searchsorted-alike in the same file (transforms.py:47) is actually
implemented with plain torch.sum/comparison ops rather than the real torch.searchsorted op —
that's good news, since it sidesteps an op with historically inconsistent ONNX opset support.
What this means practically: export is very likely possible (nothing here is fundamentally
unexportable), but it may need one of: PyTorch's newer dynamo-based exporter instead of the
legacy tracer, a higher opset version, or a small rewrite of the masked-assignment lines to
torch.where(...) equivalents (behavior-preserving, just export-friendlier). This is a
half-day-to-a-few-days spike, not a redesign, but it's not knowable in advance without trying it
— hence Phase 0.
Fallback if export genuinely can't be made to match: run inference in a small, minimal
Python (or C++/libtorch) sidecar container that does only net_g.infer(...) — no Flask, no
Japanese text handling — and have the Go tts-gateway call it over a loopback gRPC/HTTP call
from its onnxengine-equivalent secondary adapter. The service is still a real Go hexagonal
service at the architecture and API level; only the tensor math lives elsewhere, hidden behind a
port like any other outbound dependency (same shape as how ai-gateway calls out to Ollama).
Risk 2: text front-end parity
text/cleaners.py::japanese_cleaners calls pyopenjtalk.g2p(text, kana=False), which wraps the
OpenJTalk C++ library (NAIST Japanese dictionary + internal mecab) and returns a
space-stripped phoneme string. There is no maintained native-Go equivalent of this — it's not
just "run mecab," it's OpenJTalk's full NJD → JPCommon phoneme pipeline.
Two realistic options, in order of preference:
- Shell out to the
open_jtalkCLI binary from Go (os/exec), which is a real, independently-installable C++ binary (not Python) — same idea as this repo already shelling out toffmpeg. Needs verification that the CLI's phoneme/label output can be mapped 1:1 to whatpyopenjtalk.g2p(..., kana=False)produces (the CLI's default output is full HTS-style context labels, not the bare phoneme stringpyopenjtalkreturns — extracting the phoneme string from labels is a known, small parsing task, but it needs to be checked against real output before relying on it). - A tiny, persistent Python process that does only text normalization (just
cleaners.py+pyopenjtalk, notorch, no Flask) called over loopback — smaller and lower risk than option 1 if the CLI mapping turns out to be awkward, but reintroduces a Python runtime dependency.
The unidecode call in the same cleaner only ever runs on the small set of punctuation
characters matched by _japanese_marks (text/cleaners.py:9), not arbitrary Unicode — that part
is trivial to reimplement as a small Go lookup table, no library needed. The symbol vocabulary
(text/symbols.py) is a straight port to a Go slice/string constant.
Proposed architecture
New module tts-gateway, added to go.work, mirroring ha-gateway's layout:
tts-gateway/
cmd/gateway/ entrypoint: load .env, wire adapters, start gRPC
internal/core/domain/ Speaker, SynthesisParams, AudioClip
internal/core/ports/driven/ TTSEngine (Synthesize), TextNormalizer (Normalize),
AudioEncoder (Encode)
internal/core/ports/driving/ SynthesizeUseCase-shaped interface for the gRPC adapter
internal/app/ orchestration: normalize -> engine.Synthesize -> encode
internal/adapters/primary/grpc/ tts.v1.TTSService server (Synthesize, ListSpeakers)
internal/adapters/secondary/onnxengine/ onnxruntime-go + CUDA EP, loads the exported .onnx
internal/adapters/secondary/jtalk/ open_jtalk CLI subprocess adapter (or sidecar client)
internal/adapters/secondary/ffmpeg/ WAV write + ffmpeg transcode to AAC/Opus
internal/config/, internal/logger/, internal/telemetry/ copied pattern from ha-gateway
New proto package proto/tts/v1/tts.proto (buf module, same convention as proto/ha,
proto/ai):
service TTSService {
rpc Synthesize(SynthesizeRequest) returns (SynthesizeResponse);
rpc ListSpeakers(ListSpeakersRequest) returns (ListSpeakersResponse);
}
Suggested port: 50053 (next free after 50051/50052). Same .env conventions as the other
three services: TLS_DIR for optional mTLS, OTEL_ENDPOINT, LOG_FORMAT, plus new
ONNX_MODEL_PATH, CUDA_DEVICE_ID, OPEN_JTALK_BIN (or sidecar address).
Phased plan
Phase 0 — Spike & de-risk (Python, inside tmp/reference/uma-tts-api, throwaway code)
Export net_g's inference path to ONNX and run it standalone with the onnxruntime-gpu Python
package against a CUDA GPU; diff the resulting audio against app.py's current output for a
handful of speaker/text/parameter combinations. Separately, run the open_jtalk CLI on the same
sample texts and check whether its output can be turned into the same phoneme string
pyopenjtalk.g2p() produces. This phase's outcome decides whether Phases 1+ build "full Go +
ONNX" or "Go gateway + inference/text sidecar." Don't start the Go work before this answers
both questions — everything downstream depends on it.
Phase 1 — Proto + service skeleton
Add proto/tts/v1/tts.proto, buf generate, scaffold the tts-gateway module (go.mod,
go.work entry), hexagonal skeleton, health check + reflection, config/logger/telemetry copied
from ha-gateway. No inference yet — Synthesize can return a canned tone to prove the
plumbing.
Phase 2 — Text front-end adapter Port the symbol table and punctuation handling; wire the g2p adapter chosen in Phase 0.
Phase 3 — Inference adapter
Wire onnxruntime-go with the CUDA execution provider, load the Phase-0-exported .onnx file,
implement TTSEngine.Synthesize. Validate against the Phase 0 reference outputs — expect small
numeric drift between PyTorch-CPU/GPU and ONNX Runtime-GPU (different kernels), so compare on
waveform length / gross spectral similarity rather than bit-exact equality, consistent with this
repo's hand-written-mock testing convention rather than golden-byte comparison.
Phase 4 — Audio + gRPC wiring
WAV writer, ffmpeg transcode adapter, full Synthesize RPC wiring, speaker-not-found →
INVALID_ARGUMENT (mirroring app.py's current 400 behavior), ListSpeakers RPC.
Phase 5 — Containerize & deploy to nik-gpu — done. Deployed and confirmed working in
production; see handoff.md for status. Model artifact distribution is resolved as a small
versioned image (tts-gateway/model/Dockerfile) copied into a shared volume by a k8s
initContainer, built manually on nik-gpu (the checkpoint never existed anywhere CI can reach) —
see tts-gateway/README.md's "Model artifact distribution (production)" section.
Phase 6 — Client integration — done. discord-bot registers a /speak command calling
this gateway via the same secondary-adapter + gRPC-client pattern used for its
ha-gateway/ai-gateway clients (discord-bot/internal/adapters/secondary/ttsgateway/), then
transcodes the AAC response to Opus locally (ffmpeg + github.com/jonas747/dca) and streams it
into the invoking user's current voice channel
(discord-bot/internal/adapters/primary/discord/voice.go).
GPU and deployment notes
This would be the first GPU-bound, host-specific service in the repo, which breaks an assumption the other three quietly rely on (any image can run on any host):
- Dockerfile can't reuse the existing pattern.
ha-gateway/ai-gateway/discord-botall build withCGO_ENABLED=0ontogcr.io/distroless/static:nonroot— deliberately tiny, fully static, portable.onnxruntime-goneeds cgo (it links againstlibonnxruntime.so), and the CUDA execution provider needs the CUDA/cuDNN runtime present in the image. This service needs its own Dockerfile based on something likenvidia/cuda:<version>-runtime(or a slim CUDA runtime variant), withCGO_ENABLED=1,ffmpeg, and the matchingonnxruntime-linux-x64-gpurelease unpacked into it. - Version matching matters.
onnxruntime-gpuprebuilt releases are pinned to specific CUDA/cuDNN major versions. Check nik-gpu's installed Nvidia driver/CUDA version before picking an ONNX Runtime release, not after — this is the one place a Turing-generation card (CC 7.5, i.e. no Ampere-only op requirements) plus older driver could bite, not because the 2080 Super is unsupported, but because of a driver/toolkit version mismatch. --gpus all/ nvidia-container-toolkit required. nik-gpu needs the Nvidia Container Toolkit installed and Docker's nvidia runtime configured; the compose file needs adeploy.resources.reservations.devices(or--gpus) entry, unlike any existingdocker-compose.ymlin this repo.- Model artifact distribution. The exported
.onnxfile (likely similar order of magnitude to the current 455 MB.pth) shouldn't be baked into a git-tracked Dockerfile context. Decide between a build-time download step, a bind-mounted volume on nik-gpu, or an artifact registry — this is a real decision to make in Phase 5, not a detail to skip. - CI implications.
.gitea/workflows/ci.yamlcurrently builds/pushes all three service images generically. This service's image is only meaningful on a CUDA host — worth deciding whether it belongs in that same generic build/push step at all, or gets its own workflow that only runs when explicitly triggered, since a successful build here says nothing about whether it actually runs correctly on nik-gpu's GPU.
Working conventions on nik-gpu
Any command that installs or otherwise changes the nik-gpu host itself (apt packages, a
Python/pip install outside a container, Docker daemon config, driver/CUDA toolkit updates, or
anything needing sudo) must never be run automatically — print the exact command and ask the
user to run it themselves. This mirrors the hard rule already built into the nik-gpu-sync and
nik-gpu-docker-build skills, restated here because Phases 0 and 5 are the two points in this
plan most likely to want something installed on the host directly (e.g. if the Phase 0 spike is
run bare-metal instead of inside a container). Prefer pushing installs into a Dockerfile/build
step wherever possible (open-jtalk, ffmpeg, onnxruntime-gpu, CUDA runtime libs) — that keeps
nik-gpu's host state untouched and the setup reproducible from the Dockerfile alone.
Open questions (product/design, not technical blockers)
- Output codec:
keep AAC (parity with today) or switch to Opus?Resolved: kept AAC.tts-gatewayitself is unchanged;discord-bottranscodes AAC → Opus locally for voice playback (see Phase 6 above), rather than pushing an Opus-specific output format onto every other current/future consumer of this gateway. - Concurrency: the reference implementation serializes all inference behind one lock. Same starting point is reasonable here — revisit only if real latency/throughput needs show up.
- Auth: consistent with every other service in this repo, no app-layer authorization —
acceptable only if nik-gpu stays on the same trusted internal network (or gets mTLS via
TLS_DIR, matching the other three).
Non-goals
- No training/fine-tuning support — this is inference-only, same as the reference.
- No multi-GPU or batched-request scheduling in the first version.
- No change to the voice roster or model weights — same 92 speakers, same checkpoint.