- Implemented the /speak command in Discord bot to synthesize speech using the TTS gateway. - Added voice handling logic to join voice channels and play synthesized audio. - Created tests for the new command and voice functionalities. - Introduced TTSGateway interface for TTS service communication. - Updated configuration to include TTS gateway address. - Documented the TTS gateway integration and model artifact distribution process.
14 KiB
CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
Repo Overview
home-services is a Go workspace of four internal services for home control, connected by gRPC and sharing committed protobuf-generated code:
Discord users
|
v
discord-bot -----> ha-gateway -----> Home Assistant REST API
|
v
ai-gateway ------> Ollama
|
v
ha-gateway
discord-bot -----> tts-gateway -----> tts-sidecar (GPU inference)
- ha-gateway (port
50051) — gRPC boundary for Home Assistant. Talks to HA's REST API; implements entity state, light control/discovery, switch control/discovery, and climate (HVAC) control/discovery; also relays SwitchBot Cloud remote commands (RemoteApp) whenSWITCHBOT_TOKEN/SWITCHBOT_SECRETare set. Event streaming is stubbed. - ai-gateway (port
50052) — gRPC service that turns free-form text into home actions. Calls Ollama for intent extraction, resolves intents against a cached light list fromha-gateway, and callsha-gatewayto execute approved actions. - discord-bot — registers
/light,/switch,/ac,/ai,/speakslash commands and callsha-gateway/ai-gateway/tts-gatewayvia gRPC clients./speaksynthesizes speech viatts-gateway(which returns AAC) and transcodes it to Opus locally (viaffmpeg/github.com/jonas747/dca) to stream into the invoking user's current voice channel — seeinternal/adapters/primary/discord/voice.go. - tts-gateway (port
50053) — gRPC text-to-speech service (VITS voice model, 92 Umamusume voices), paired with a Python/libtorch inference sidecar (tts-gateway/sidecar/, deployed as a second container in the same pod) that needs an Nvidia GPU — seeTTS_GATEWAY_PLAN.mdfor why. Called bydiscord-bot's/speakcommand. Seetts-gateway/README.mdfor its API/config/local-run details.
Each service is a separate Go module (own go.mod) joined by go.work at the root, plus a gen module for shared generated code. Module paths are gitea.nik4nao.com/nik/home-services/{ha-gateway,ai-gateway,discord-bot,tts-gateway,gen}.
Architecture (hexagonal, per service)
Every service follows the same internal layout, with dependencies pointing inward:
cmd/<entrypoint>/ # process entrypoint and wiring (loads .env, builds adapters, starts gRPC)
internal/adapters/primary/ # inbound edges: gRPC servers, Discord handlers
internal/adapters/secondary/ # outbound edges: HA REST client, Ollama client, ha-gateway/ai-gateway gRPC clients
internal/app/ # use-case orchestration
internal/core/domain/ # domain types
internal/core/ports/ # driving (inbound) and driven (outbound) interfaces
internal/config/ # environment loading
internal/logger/ # slog setup
internal/telemetry/ # OpenTelemetry setup
internal/core has no dependency on adapters — ports are interfaces that adapters implement (driven) or call into (driving). When adding a capability, the usual path is: define/extend a port in core/ports, implement orchestration in app, then wire an adapter in adapters/primary or adapters/secondary.
Protobuf contracts live in proto/ (buf module, ai/v1, ha/v1, and tts/v1 packages). Generated Go code is committed under gen/ and consumed by all four services through the Go workspace — do not hand-edit files in gen/.
tmp/ is reference-only, never a dependency
tmp/ is gitignored — nothing under it is pushed to git, and it should be treated as temporary scratch space for reference material (e.g. tmp/reference/switchbot-control-reference/, a standalone CLI copied in for local discovery/testing against an external API). Rules:
- It's fine to read code under
tmp/for patterns, to run its scripts/tools locally (a discovery script, a CLI, etc.), or to use it as a manual testing aid. - Never make any of the four services (
ha-gateway,ai-gateway,discord-bot,tts-gateway) import,go.work use, or otherwise depend on anything undertmp/at build or runtime. Sincetmp/isn't committed, that dependency would silently break for every other clone of the repo (including CI).tts-gateway/sidecar/is a real example of doing this correctly — its model code was copied out oftmp/reference/uma-tts-api/into a committed location rather than referenced in place. - If a reference tool under
tmp/lives in its own Go module nested inside this repo'sgo.workworkspace, invoke it withGOWORK=offrather than adding it to the rootgo.work— e.g.GOWORK=off ./scripts/some-tool ...from that tool's own directory. - If something under
tmp/turns out to be genuinely needed at runtime, port the actual logic into the relevant service'sinternal/tree (following the hexagonal layout above) instead of reaching intotmp/from committed code.
Common Commands
Regenerate protobuf code after changing anything under proto/ (requires buf):
buf generate
Run tests / vet (from repo root, or cd into a service and drop the prefix):
go test ./ha-gateway/... ./ai-gateway/... ./discord-bot/... ./tts-gateway/...
go vet ./ha-gateway/... ./ai-gateway/... ./discord-bot/... ./tts-gateway/...
Run a single test:
cd ha-gateway && go test ./internal/app/... -run TestEntityAppGetState
Build binaries:
go build ./ha-gateway/... ./ai-gateway/... ./discord-bot/... ./tts-gateway/...
Run a service locally (each loads .env from its own working directory via godotenv, so cd into the service dir first):
cd ha-gateway && cp .env.example .env && go run ./cmd/gateway
cd ai-gateway && cp .env.example .env && go run ./cmd/gateway
cd discord-bot && cp .env.example .env && go run ./cmd/bot
tts-gateway additionally needs the open_jtalk CLI/dictionary/voice, ffmpeg, and a running
inference sidecar to actually synthesize anything — see tts-gateway/README.md for the full
local/nik-gpu run instructions rather than a short snippet here.
For local plaintext dev, point gateways at each other with TLS_DIR empty, e.g. HA_GATEWAY_ADDR=localhost:50051, AI_GATEWAY_ADDR=localhost:50052.
Build container images (from repo root, since Dockerfiles reference the whole workspace):
docker build -f ha-gateway/Dockerfile -t ha-gateway:dev .
docker build -f ai-gateway/Dockerfile -t ai-gateway:dev .
docker build -f discord-bot/Dockerfile -t discord-bot:dev .
docker build -f tts-gateway/Dockerfile -t tts-gateway:dev .
docker build -f tts-gateway/sidecar/Dockerfile -t tts-sidecar:dev tts-gateway/sidecar
gRPC smoke checks (with the relevant service running):
grpcurl -plaintext -d '{"domain":"light"}' localhost:50051 ha.v1.EntityService/ListStates
grpcurl -plaintext -d '{"entity_id":"light.living_room","brightness_pct":80}' localhost:50051 ha.v1.LightService/TurnOn
grpcurl -plaintext -d '{"text":"turn on the desk lamp","source":"local"}' localhost:50052 ai.v1.AIService/Query
grpcurl -plaintext -d '{}' localhost:50052 ai.v1.AIService/ListModels
tts-gateway's smoke check needs the sidecar (and, for real inference, a GPU) up first — see
tts-gateway/README.md rather than duplicating the fuller sequence here.
Note: ha-gateway always registers gRPC reflection; ai-gateway only registers reflection when LOG_LEVEL=debug.
Testing Conventions
Tests use the standard library testing package only (no testify). Mocks are hand-written structs implementing the relevant core/ports/driven interface with function fields (see ha-gateway/internal/app/entity_test.go for the pattern).
Configuration Notes
- Every service reads
TLS_DIRto enable optional mTLS; when set, the directory must containtls.crt,tls.key, andca.crt. OTEL_ENDPOINTenables OTLP gRPC traces/metrics; leave empty for local no-op telemetry.LOG_FORMAT=jsonis the production default;textis easier to read locally.ha-gatewayreadsSWITCHBOT_TOKEN/SWITCHBOT_SECRET(optional) to enable SwitchBot Cloud remote commands viaRemoteApp; leave empty to disable that path.- None of the services implement app-layer authorization — they rely on being kept on a trusted internal network or on mTLS. Keep this in mind before adding any endpoint that wasn't previously reachable.
nik-gpu deployment target
tts-gateway's inference sidecar (see TTS_GATEWAY_PLAN.md for the full history) runs on
nik-gpu, a remote Nvidia GPU host reachable via ssh nik-gpu — both for local dev/smoke-testing
(via the Claude Code skills nik-gpu-status (read-only check), nik-gpu-sync (rsync this repo to
~/repo/home-service/ there), nik-gpu-docker-build (build/smoke-test via a nik-gpu Docker
context), and tts-gateway/README.md's run instructions) and as the actual GPU node backing the
production Kubernetes deployment — see "Infrastructure / Deployment" below.
Never automatically run installation or other host-system-altering commands on nik-gpu —
apt/apt-get, pip install outside a container, Docker daemon config changes, driver/toolkit
updates, or anything requiring sudo. Always print the exact command and ask the user to run it
themselves (their own terminal, or ! <command> in a Claude Code session). This applies to the
bare nik-gpu host specifically; installing packages inside a Dockerfile build (e.g.
apt-get install open-jtalk as a build step) is a normal container build action, not a host
mutation, and is fine to run.
Infrastructure / Deployment
This repo builds and pushes Docker images (see CI below) but does not deploy them. Actual
deployment — Kubernetes manifests, Secrets, GPU scheduling — lives in a separate repo:
~/repo/homelab, specifically manifests/home-services/*.yaml. That covers Deployments/Services
for all four services here (namespace home-services), plus cluster-level pieces like
nvidia-device-plugin.yaml and sealed secrets (*-sealed.yaml manifests generated from
*-secret.sh scripts, one pair per service that needs one — tts-gateway doesn't have one yet,
since unlike ha-gateway/discord-bot it currently needs no external tokens/credentials).
When a question is about how something actually behaves in production — not just what the code
does — check that repo before guessing or assuming something isn't deployed; this repo alone
doesn't show the full picture. Specifics of what's actually in tts-gateway.yaml there (current
as of 2026-07-25 — it's maintained independently of this repo, so verify against the live file
rather than trusting this description to stay accurate):
tts-gatewayandtts-sidecarrun as two containers in one Pod, not separate Deployments — the Go gateway reaches the sidecar overlocalhost:50054, mirroring the--network hostpattern documented intts-gateway/README.mdfor local Docker testing.- GPU scheduling:
runtimeClassName: nvidia,nodeSelector: nik4nao.com/gpu: "true", andnvidia.com/gpu: 1request/limit on the sidecar container.strategy: Recreateinstead of the defaultRollingUpdate— nik-gpu only has one allocatable GPU, so a rolling update would deadlock waiting for a GPU still held by the pod it's replacing. - Model artifact distribution (checkpoint + hparams) is an
emptyDirvolume populated at pod start by aninitContainersentry (model-init) that copies from a small, versioned image (gitea.nik4nao.com/nik/tts-model:<tag>, built fromtts-gateway/model/Dockerfile) — not a rawhostPathinto the node's disk anymore. That model image still can't be built by CI (the checkpoint isn't committed to git, ~455MB, and only ever existed on nik-gpu's local disk) — it's built and pushed manually, on nik-gpu, whenever the checkpoint/config change. Seetts-gateway/README.mdfor the exact commands. - mTLS (
TLS_DIR) is currently commented out ontts-gateway, for plaintextgrpcurltesting from outside the cluster during initial rollout — a live TODO, not a permanent decision, unlike every other service here which assumes mTLS-or-trusted-network as its only access boundary (see "Configuration Notes" above).
CI
.gitea/workflows/ci.yaml always runs go vet/go test for all five modules (gen, ai-gateway, ha-gateway, discord-bot, tts-gateway) on every push/PR. On pushes to main, a changes job (dorny/paths-filter) determines which of the five images actually need rebuilding based on which paths changed, so an edit scoped to one service doesn't rebuild (and re-push) all of them — gen/, go.work, and go.work.sum count as shared and mark every Go-based image as changed, since a dependency bump there can affect all of them. Each build-* job is gated on that output and, when it runs, uses docker/build-push-action's registry-based cache (cache-from/cache-to: type=registry,ref=.../<image>:buildcache) so unchanged Docker layers (e.g. go mod download, apt-get/pip install) don't get redone on every run — the first run after adding this has nothing to pull from and builds fully fresh, subsequent ones should be much faster. tts-sidecar (the Python/CUDA inference sidecar under tts-gateway/sidecar/) is a large ~13GB image; it builds fine on a generic runner since only running it needs a GPU, not building it.
Note that ai-gateway/Dockerfile, ha-gateway/Dockerfile, discord-bot/Dockerfile, and tts-gateway/Dockerfile each COPY every other service directory (not just their own) because go.work lists all five Go modules as workspace members — Go's workspace-mode module resolution needs every listed directory present in the build context, even ones a given service doesn't otherwise depend on. Adding a new module to go.work means adding a matching COPY line (both the manifest-only and full-source copies, see below) to the other Dockerfiles too, or their builds break. Each of those four Dockerfiles also copies every module's go.mod/go.sum first and runs go mod download before copying full source, so that layer's cache survives source-only edits instead of being invalidated by every commit.