home-services/TTS_GATEWAY_PLAN.md
Nik Afiq 0d58e46740
All checks were successful
CI / changes (push) Successful in 1s
CI / test (push) Successful in 6s
CI / build-ai-gateway (push) Has been skipped
CI / build-ha-gateway (push) Has been skipped
CI / build-discord-bot (push) Successful in 1m33s
CI / build-tts-gateway (push) Successful in 37s
CI / build-tts-sidecar (push) Has been skipped
feat: add /speak command for TTS integration
- Implemented the /speak command in Discord bot to synthesize speech using the TTS gateway.
- Added voice handling logic to join voice channels and play synthesized audio.
- Created tests for the new command and voice functionalities.
- Introduced TTSGateway interface for TTS service communication.
- Updated configuration to include TTS gateway address.
- Documented the TTS gateway integration and model artifact distribution process.
2026-07-25 02:09:02 +09:00

246 lines
16 KiB
Markdown

# tts-gateway: plan & feasibility
Proposal for turning `tmp/reference/uma-tts-api` (Python/Flask, VITS voice model) into a fourth
`home-services` module, **`tts-gateway`**, written in Go, following the same hexagonal pattern as
`ha-gateway`/`ai-gateway`/`discord-bot`, deployed on **nik-gpu** (Nvidia RTX 2080 Super).
See [`tmp/reference/uma-tts-api/FINDINGS.md`](tmp/reference/uma-tts-api/FINDINGS.md) for how the
reference implementation works today; this document is about porting it.
## TL;DR feasibility
| Piece | Verdict | Why |
|---|---|---|
| Serving/API layer (gRPC, config, health, telemetry) | **Straightforward** | Identical shape to the other three services; pure Go, no new tech. |
| Audio postprocessing (WAV → AAC/Opus) | **Straightforward** | Write PCM/WAV by hand, shell out to `ffmpeg` (already a required system dep). |
| Model inference (VITS forward pass) | **Feasible, one open risk** | No mature native-Go path to run a `.pth` checkpoint directly. The realistic route is: export the *inference-only* graph to **ONNX** once (offline, in Python), then run it in Go via `onnxruntime-go` with the CUDA execution provider. The 2080 Super (Turing, CC 7.5) is well inside ONNX Runtime's supported GPU range. The risk isn't "can Go do GPU inference" (yes), it's "does this specific checkpoint's graph export cleanly to ONNX" — see [Risk 1](#risk-1-onnx-export-of-the-stochastic-duration-predictor). |
| Japanese text front-end (`pyopenjtalk` g2p) | **Feasible, one open risk** | No usable native-Go OpenJTalk/mecab g2p implementation exists. Shelling out to the `open_jtalk` CLI (a real C++ binary, not Python) from Go is the likely path, but its default output format needs to be checked against what `pyopenjtalk.g2p()` actually returns — see [Risk 2](#risk-2-text-front-end-parity). |
| Overall "100% Go, no Python at runtime" | **Likely, not guaranteed** | Contingent on Risks 1 and 2 resolving cleanly. Both have a concrete fallback (below) that still ships a real Go service, just with a small non-Go component for one stage. |
Bottom line: this is buildable and the GPU is not a blocker. The two things that could force a
hybrid design instead of pure Go are named explicitly below, and Phase 0 exists specifically to
answer them *before* any Go code is written.
## What "parity" means
The new service should support what `uma-tts-api` supports today:
- Synthesize speech for one of 92 Umamusume speakers given Japanese text, with the same
`noise_scale` / `noise_scale_w` / `length_scale` knobs.
- List/search speakers.
- Health check.
Not required to carry over: training code, the `monotonic_align` Cython extension (only used by
the *training* loss, confirmed below — never called from `infer()`), or CPU fallback as a
first-class target (nik-gpu has a GPU; CPU-only would defeat the point of moving it there).
## Why ONNX Runtime instead of a Go libtorch binding
Two ways exist to run a PyTorch model's math from Go:
1. **cgo bindings to libtorch** (e.g. community "gotorch"-style projects). Rejected: these bind
directly to PyTorch's C++ ABI, which is unstable across versions, poorly maintained as Go
packages, and would require hand-porting the whole `SynthesizerTrn` forward pass into
Go/C++ rather than reusing a serialized graph. High effort, fragile, no real upside here.
2. **Export to ONNX once, run via `onnxruntime-go`** (e.g. `github.com/yalue/onnxruntime_go`),
using Microsoft's prebuilt `onnxruntime-linux-x64-gpu` shared library with the CUDA execution
provider. This is the standard, maintained path for "run a PyTorch model from a non-Python
language," and it converts the problem from "port a neural net to Go" into "export a graph
once, then call a stable C API." **Recommended.**
## Risk 1: ONNX export of the stochastic duration predictor
Checked in the reference code: `configs/uma.json`'s `model` section doesn't set `use_sdp`, and
`SynthesizerTrn.__init__` defaults `use_sdp=True` (`models/models.py:413`), so this checkpoint
*does* use the flow-based `StochasticDurationPredictor`, not the simpler deterministic one.
That module's reverse (inference) path chains several `ConvFlow`s, which call into
`transforms.py`'s `rational_quadratic_spline` / `unconstrained_rational_quadratic_spline`. Those
functions use boolean-mask indexed assignment (`outputs[outside_interval_mask] = ...`,
`transforms.py:68-95`) to handle spline boundary conditions. This pattern is a known rough edge
for `torch.onnx.export`'s tracer — it's data-dependent control flow that doesn't always lower to
a single static ONNX graph cleanly, unlike the rest of the model (convolutions, attention,
LayerNorm, ConvTranspose1d) which are all completely standard and export fine.
The custom `torch.searchsorted`-alike in the same file (`transforms.py:47`) is actually
implemented with plain `torch.sum`/comparison ops rather than the real `torch.searchsorted` op —
that's good news, since it sidesteps an op with historically inconsistent ONNX opset support.
**What this means practically:** export is very likely possible (nothing here is fundamentally
unexportable), but it may need one of: PyTorch's newer `dynamo`-based exporter instead of the
legacy tracer, a higher opset version, or a small rewrite of the masked-assignment lines to
`torch.where(...)` equivalents (behavior-preserving, just export-friendlier). This is a
half-day-to-a-few-days spike, not a redesign, but it's not knowable in advance without trying it
— hence Phase 0.
**Fallback if export genuinely can't be made to match:** run inference in a small, minimal
Python (or C++/libtorch) sidecar container that does *only* `net_g.infer(...)` — no Flask, no
Japanese text handling — and have the Go `tts-gateway` call it over a loopback gRPC/HTTP call
from its `onnxengine`-equivalent secondary adapter. The service is still a real Go hexagonal
service at the architecture and API level; only the tensor math lives elsewhere, hidden behind a
port like any other outbound dependency (same shape as how `ai-gateway` calls out to Ollama).
## Risk 2: text front-end parity
`text/cleaners.py::japanese_cleaners` calls `pyopenjtalk.g2p(text, kana=False)`, which wraps the
OpenJTalk C++ library (NAIST Japanese dictionary + internal mecab) and returns a
space-stripped phoneme string. There is no maintained native-Go equivalent of this — it's not
just "run mecab," it's OpenJTalk's full NJD → JPCommon phoneme pipeline.
Two realistic options, in order of preference:
1. **Shell out to the `open_jtalk` CLI binary** from Go (`os/exec`), which is a real,
independently-installable C++ binary (not Python) — same idea as this repo already shelling
out to `ffmpeg`. Needs verification that the CLI's phoneme/label output can be mapped 1:1 to
what `pyopenjtalk.g2p(..., kana=False)` produces (the CLI's default output is full
HTS-style context labels, not the bare phoneme string `pyopenjtalk` returns — extracting the
phoneme string from labels is a known, small parsing task, but it needs to be checked against
real output before relying on it).
2. **A tiny, persistent Python process that does only text normalization** (just
`cleaners.py` + `pyopenjtalk`, no `torch`, no Flask) called over loopback — smaller and lower
risk than option 1 if the CLI mapping turns out to be awkward, but reintroduces a Python
runtime dependency.
The `unidecode` call in the same cleaner only ever runs on the small set of punctuation
characters matched by `_japanese_marks` (`text/cleaners.py:9`), not arbitrary Unicode — that part
is trivial to reimplement as a small Go lookup table, no library needed. The symbol vocabulary
(`text/symbols.py`) is a straight port to a Go slice/string constant.
## Proposed architecture
New module `tts-gateway`, added to `go.work`, mirroring `ha-gateway`'s layout:
```text
tts-gateway/
cmd/gateway/ entrypoint: load .env, wire adapters, start gRPC
internal/core/domain/ Speaker, SynthesisParams, AudioClip
internal/core/ports/driven/ TTSEngine (Synthesize), TextNormalizer (Normalize),
AudioEncoder (Encode)
internal/core/ports/driving/ SynthesizeUseCase-shaped interface for the gRPC adapter
internal/app/ orchestration: normalize -> engine.Synthesize -> encode
internal/adapters/primary/grpc/ tts.v1.TTSService server (Synthesize, ListSpeakers)
internal/adapters/secondary/onnxengine/ onnxruntime-go + CUDA EP, loads the exported .onnx
internal/adapters/secondary/jtalk/ open_jtalk CLI subprocess adapter (or sidecar client)
internal/adapters/secondary/ffmpeg/ WAV write + ffmpeg transcode to AAC/Opus
internal/config/, internal/logger/, internal/telemetry/ copied pattern from ha-gateway
```
New proto package `proto/tts/v1/tts.proto` (buf module, same convention as `proto/ha`,
`proto/ai`):
```protobuf
service TTSService {
rpc Synthesize(SynthesizeRequest) returns (SynthesizeResponse);
rpc ListSpeakers(ListSpeakersRequest) returns (ListSpeakersResponse);
}
```
Suggested port: `50053` (next free after `50051`/`50052`). Same `.env` conventions as the other
three services: `TLS_DIR` for optional mTLS, `OTEL_ENDPOINT`, `LOG_FORMAT`, plus new
`ONNX_MODEL_PATH`, `CUDA_DEVICE_ID`, `OPEN_JTALK_BIN` (or sidecar address).
## Phased plan
**Phase 0 — Spike & de-risk (Python, inside `tmp/reference/uma-tts-api`, throwaway code)**
Export `net_g`'s inference path to ONNX and run it standalone with the `onnxruntime-gpu` Python
package against a CUDA GPU; diff the resulting audio against `app.py`'s current output for a
handful of speaker/text/parameter combinations. Separately, run the `open_jtalk` CLI on the same
sample texts and check whether its output can be turned into the same phoneme string
`pyopenjtalk.g2p()` produces. **This phase's outcome decides whether Phases 1+ build "full Go +
ONNX" or "Go gateway + inference/text sidecar."** Don't start the Go work before this answers
both questions — everything downstream depends on it.
**Phase 1 — Proto + service skeleton**
Add `proto/tts/v1/tts.proto`, `buf generate`, scaffold the `tts-gateway` module (`go.mod`,
`go.work` entry), hexagonal skeleton, health check + reflection, config/logger/telemetry copied
from `ha-gateway`. No inference yet — `Synthesize` can return a canned tone to prove the
plumbing.
**Phase 2 — Text front-end adapter**
Port the symbol table and punctuation handling; wire the g2p adapter chosen in Phase 0.
**Phase 3 — Inference adapter**
Wire `onnxruntime-go` with the CUDA execution provider, load the Phase-0-exported `.onnx` file,
implement `TTSEngine.Synthesize`. Validate against the Phase 0 reference outputs — expect small
numeric drift between PyTorch-CPU/GPU and ONNX Runtime-GPU (different kernels), so compare on
waveform length / gross spectral similarity rather than bit-exact equality, consistent with this
repo's hand-written-mock testing convention rather than golden-byte comparison.
**Phase 4 — Audio + gRPC wiring**
WAV writer, `ffmpeg` transcode adapter, full `Synthesize` RPC wiring, speaker-not-found →
`INVALID_ARGUMENT` (mirroring `app.py`'s current 400 behavior), `ListSpeakers` RPC.
**Phase 5 — Containerize & deploy to nik-gpu****done.** Deployed and confirmed working in
production; see `handoff.md` for status. Model artifact distribution is resolved as a small
versioned image (`tts-gateway/model/Dockerfile`) copied into a shared volume by a k8s
`initContainer`, built manually on nik-gpu (the checkpoint never existed anywhere CI can reach) —
see `tts-gateway/README.md`'s "Model artifact distribution (production)" section.
**Phase 6 — Client integration****done.** `discord-bot` registers a `/speak` command calling
this gateway via the same secondary-adapter + gRPC-client pattern used for its
`ha-gateway`/`ai-gateway` clients (`discord-bot/internal/adapters/secondary/ttsgateway/`), then
transcodes the AAC response to Opus locally (`ffmpeg` + `github.com/jonas747/dca`) and streams it
into the invoking user's current voice channel
(`discord-bot/internal/adapters/primary/discord/voice.go`).
## GPU and deployment notes
This would be the first GPU-bound, host-specific service in the repo, which breaks an assumption
the other three quietly rely on (any image can run on any host):
- **Dockerfile can't reuse the existing pattern.** `ha-gateway`/`ai-gateway`/`discord-bot` all
build with `CGO_ENABLED=0` onto `gcr.io/distroless/static:nonroot` — deliberately tiny, fully
static, portable. `onnxruntime-go` needs cgo (it links against `libonnxruntime.so`), and the
CUDA execution provider needs the CUDA/cuDNN runtime present in the image. This service needs
its own Dockerfile based on something like `nvidia/cuda:<version>-runtime` (or a slim CUDA
runtime variant), with `CGO_ENABLED=1`, `ffmpeg`, and the matching `onnxruntime-linux-x64-gpu`
release unpacked into it.
- **Version matching matters.** `onnxruntime-gpu` prebuilt releases are pinned to specific
CUDA/cuDNN major versions. Check nik-gpu's installed Nvidia driver/CUDA version *before*
picking an ONNX Runtime release, not after — this is the one place a Turing-generation card
(CC 7.5, i.e. no Ampere-only op requirements) plus older driver could bite, not because the
2080 Super is unsupported, but because of a driver/toolkit version mismatch.
- **`--gpus all` / nvidia-container-toolkit required.** nik-gpu needs the Nvidia Container
Toolkit installed and Docker's nvidia runtime configured; the compose file needs a
`deploy.resources.reservations.devices` (or `--gpus`) entry, unlike any existing
`docker-compose.yml` in this repo.
- **Model artifact distribution.** The exported `.onnx` file (likely similar order of magnitude
to the current 455 MB `.pth`) shouldn't be baked into a git-tracked Dockerfile context. Decide
between a build-time download step, a bind-mounted volume on nik-gpu, or an artifact registry
— this is a real decision to make in Phase 5, not a detail to skip.
- **CI implications.** `.gitea/workflows/ci.yaml` currently builds/pushes all three service
images generically. This service's image is only meaningful on a CUDA host — worth deciding
whether it belongs in that same generic build/push step at all, or gets its own workflow that
only runs when explicitly triggered, since a successful build here says nothing about whether
it actually runs correctly on nik-gpu's GPU.
## Working conventions on nik-gpu
Any command that installs or otherwise changes the nik-gpu **host** itself (apt packages, a
Python/pip install outside a container, Docker daemon config, driver/CUDA toolkit updates, or
anything needing `sudo`) must never be run automatically — print the exact command and ask the
user to run it themselves. This mirrors the hard rule already built into the `nik-gpu-sync` and
`nik-gpu-docker-build` skills, restated here because Phases 0 and 5 are the two points in this
plan most likely to want something installed on the host directly (e.g. if the Phase 0 spike is
run bare-metal instead of inside a container). Prefer pushing installs into a Dockerfile/build
step wherever possible (`open-jtalk`, `ffmpeg`, `onnxruntime-gpu`, CUDA runtime libs) — that keeps
nik-gpu's host state untouched and the setup reproducible from the Dockerfile alone.
## Open questions (product/design, not technical blockers)
- **Output codec:** ~~keep AAC (parity with today) or switch to Opus?~~ **Resolved: kept AAC.**
`tts-gateway` itself is unchanged; `discord-bot` transcodes AAC → Opus locally for voice
playback (see Phase 6 above), rather than pushing an Opus-specific output format onto every
other current/future consumer of this gateway.
- **Concurrency:** the reference implementation serializes all inference behind one lock. Same
starting point is reasonable here — revisit only if real latency/throughput needs show up.
- **Auth:** consistent with every other service in this repo, no app-layer authorization —
acceptable only if nik-gpu stays on the same trusted internal network (or gets mTLS via
`TLS_DIR`, matching the other three).
## Non-goals
- No training/fine-tuning support — this is inference-only, same as the reference.
- No multi-GPU or batched-request scheduling in the first version.
- No change to the voice roster or model weights — same 92 speakers, same checkpoint.