home-services/CLAUDE.md
Nik Afiq 0d58e46740
All checks were successful
CI / changes (push) Successful in 1s
CI / test (push) Successful in 6s
CI / build-ai-gateway (push) Has been skipped
CI / build-ha-gateway (push) Has been skipped
CI / build-discord-bot (push) Successful in 1m33s
CI / build-tts-gateway (push) Successful in 37s
CI / build-tts-sidecar (push) Has been skipped
feat: add /speak command for TTS integration
- Implemented the /speak command in Discord bot to synthesize speech using the TTS gateway.
- Added voice handling logic to join voice channels and play synthesized audio.
- Created tests for the new command and voice functionalities.
- Introduced TTSGateway interface for TTS service communication.
- Updated configuration to include TTS gateway address.
- Documented the TTS gateway integration and model artifact distribution process.
2026-07-25 02:09:02 +09:00

14 KiB

CLAUDE.md

This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.

Repo Overview

home-services is a Go workspace of four internal services for home control, connected by gRPC and sharing committed protobuf-generated code:

Discord users
     |
     v
discord-bot -----> ha-gateway -----> Home Assistant REST API
     |
     v
ai-gateway ------> Ollama
     |
     v
ha-gateway

discord-bot -----> tts-gateway -----> tts-sidecar (GPU inference)
  • ha-gateway (port 50051) — gRPC boundary for Home Assistant. Talks to HA's REST API; implements entity state, light control/discovery, switch control/discovery, and climate (HVAC) control/discovery; also relays SwitchBot Cloud remote commands (RemoteApp) when SWITCHBOT_TOKEN/SWITCHBOT_SECRET are set. Event streaming is stubbed.
  • ai-gateway (port 50052) — gRPC service that turns free-form text into home actions. Calls Ollama for intent extraction, resolves intents against a cached light list from ha-gateway, and calls ha-gateway to execute approved actions.
  • discord-bot — registers /light, /switch, /ac, /ai, /speak slash commands and calls ha-gateway/ai-gateway/tts-gateway via gRPC clients. /speak synthesizes speech via tts-gateway (which returns AAC) and transcodes it to Opus locally (via ffmpeg/github.com/jonas747/dca) to stream into the invoking user's current voice channel — see internal/adapters/primary/discord/voice.go.
  • tts-gateway (port 50053) — gRPC text-to-speech service (VITS voice model, 92 Umamusume voices), paired with a Python/libtorch inference sidecar (tts-gateway/sidecar/, deployed as a second container in the same pod) that needs an Nvidia GPU — see TTS_GATEWAY_PLAN.md for why. Called by discord-bot's /speak command. See tts-gateway/README.md for its API/config/local-run details.

Each service is a separate Go module (own go.mod) joined by go.work at the root, plus a gen module for shared generated code. Module paths are gitea.nik4nao.com/nik/home-services/{ha-gateway,ai-gateway,discord-bot,tts-gateway,gen}.

Architecture (hexagonal, per service)

Every service follows the same internal layout, with dependencies pointing inward:

cmd/<entrypoint>/        # process entrypoint and wiring (loads .env, builds adapters, starts gRPC)
internal/adapters/primary/   # inbound edges: gRPC servers, Discord handlers
internal/adapters/secondary/ # outbound edges: HA REST client, Ollama client, ha-gateway/ai-gateway gRPC clients
internal/app/            # use-case orchestration
internal/core/domain/    # domain types
internal/core/ports/     # driving (inbound) and driven (outbound) interfaces
internal/config/         # environment loading
internal/logger/         # slog setup
internal/telemetry/      # OpenTelemetry setup

internal/core has no dependency on adapters — ports are interfaces that adapters implement (driven) or call into (driving). When adding a capability, the usual path is: define/extend a port in core/ports, implement orchestration in app, then wire an adapter in adapters/primary or adapters/secondary.

Protobuf contracts live in proto/ (buf module, ai/v1, ha/v1, and tts/v1 packages). Generated Go code is committed under gen/ and consumed by all four services through the Go workspace — do not hand-edit files in gen/.

tmp/ is reference-only, never a dependency

tmp/ is gitignored — nothing under it is pushed to git, and it should be treated as temporary scratch space for reference material (e.g. tmp/reference/switchbot-control-reference/, a standalone CLI copied in for local discovery/testing against an external API). Rules:

  • It's fine to read code under tmp/ for patterns, to run its scripts/tools locally (a discovery script, a CLI, etc.), or to use it as a manual testing aid.
  • Never make any of the four services (ha-gateway, ai-gateway, discord-bot, tts-gateway) import, go.work use, or otherwise depend on anything under tmp/ at build or runtime. Since tmp/ isn't committed, that dependency would silently break for every other clone of the repo (including CI). tts-gateway/sidecar/ is a real example of doing this correctly — its model code was copied out of tmp/reference/uma-tts-api/ into a committed location rather than referenced in place.
  • If a reference tool under tmp/ lives in its own Go module nested inside this repo's go.work workspace, invoke it with GOWORK=off rather than adding it to the root go.work — e.g. GOWORK=off ./scripts/some-tool ... from that tool's own directory.
  • If something under tmp/ turns out to be genuinely needed at runtime, port the actual logic into the relevant service's internal/ tree (following the hexagonal layout above) instead of reaching into tmp/ from committed code.

Common Commands

Regenerate protobuf code after changing anything under proto/ (requires buf):

buf generate

Run tests / vet (from repo root, or cd into a service and drop the prefix):

go test ./ha-gateway/... ./ai-gateway/... ./discord-bot/... ./tts-gateway/...
go vet ./ha-gateway/... ./ai-gateway/... ./discord-bot/... ./tts-gateway/...

Run a single test:

cd ha-gateway && go test ./internal/app/... -run TestEntityAppGetState

Build binaries:

go build ./ha-gateway/... ./ai-gateway/... ./discord-bot/... ./tts-gateway/...

Run a service locally (each loads .env from its own working directory via godotenv, so cd into the service dir first):

cd ha-gateway && cp .env.example .env && go run ./cmd/gateway
cd ai-gateway && cp .env.example .env && go run ./cmd/gateway
cd discord-bot && cp .env.example .env && go run ./cmd/bot

tts-gateway additionally needs the open_jtalk CLI/dictionary/voice, ffmpeg, and a running inference sidecar to actually synthesize anything — see tts-gateway/README.md for the full local/nik-gpu run instructions rather than a short snippet here.

For local plaintext dev, point gateways at each other with TLS_DIR empty, e.g. HA_GATEWAY_ADDR=localhost:50051, AI_GATEWAY_ADDR=localhost:50052.

Build container images (from repo root, since Dockerfiles reference the whole workspace):

docker build -f ha-gateway/Dockerfile -t ha-gateway:dev .
docker build -f ai-gateway/Dockerfile -t ai-gateway:dev .
docker build -f discord-bot/Dockerfile -t discord-bot:dev .
docker build -f tts-gateway/Dockerfile -t tts-gateway:dev .
docker build -f tts-gateway/sidecar/Dockerfile -t tts-sidecar:dev tts-gateway/sidecar

gRPC smoke checks (with the relevant service running):

grpcurl -plaintext -d '{"domain":"light"}' localhost:50051 ha.v1.EntityService/ListStates
grpcurl -plaintext -d '{"entity_id":"light.living_room","brightness_pct":80}' localhost:50051 ha.v1.LightService/TurnOn
grpcurl -plaintext -d '{"text":"turn on the desk lamp","source":"local"}' localhost:50052 ai.v1.AIService/Query
grpcurl -plaintext -d '{}' localhost:50052 ai.v1.AIService/ListModels

tts-gateway's smoke check needs the sidecar (and, for real inference, a GPU) up first — see tts-gateway/README.md rather than duplicating the fuller sequence here.

Note: ha-gateway always registers gRPC reflection; ai-gateway only registers reflection when LOG_LEVEL=debug.

Testing Conventions

Tests use the standard library testing package only (no testify). Mocks are hand-written structs implementing the relevant core/ports/driven interface with function fields (see ha-gateway/internal/app/entity_test.go for the pattern).

Configuration Notes

  • Every service reads TLS_DIR to enable optional mTLS; when set, the directory must contain tls.crt, tls.key, and ca.crt.
  • OTEL_ENDPOINT enables OTLP gRPC traces/metrics; leave empty for local no-op telemetry.
  • LOG_FORMAT=json is the production default; text is easier to read locally.
  • ha-gateway reads SWITCHBOT_TOKEN/SWITCHBOT_SECRET (optional) to enable SwitchBot Cloud remote commands via RemoteApp; leave empty to disable that path.
  • None of the services implement app-layer authorization — they rely on being kept on a trusted internal network or on mTLS. Keep this in mind before adding any endpoint that wasn't previously reachable.

nik-gpu deployment target

tts-gateway's inference sidecar (see TTS_GATEWAY_PLAN.md for the full history) runs on nik-gpu, a remote Nvidia GPU host reachable via ssh nik-gpu — both for local dev/smoke-testing (via the Claude Code skills nik-gpu-status (read-only check), nik-gpu-sync (rsync this repo to ~/repo/home-service/ there), nik-gpu-docker-build (build/smoke-test via a nik-gpu Docker context), and tts-gateway/README.md's run instructions) and as the actual GPU node backing the production Kubernetes deployment — see "Infrastructure / Deployment" below.

Never automatically run installation or other host-system-altering commands on nik-gpuapt/apt-get, pip install outside a container, Docker daemon config changes, driver/toolkit updates, or anything requiring sudo. Always print the exact command and ask the user to run it themselves (their own terminal, or ! <command> in a Claude Code session). This applies to the bare nik-gpu host specifically; installing packages inside a Dockerfile build (e.g. apt-get install open-jtalk as a build step) is a normal container build action, not a host mutation, and is fine to run.

Infrastructure / Deployment

This repo builds and pushes Docker images (see CI below) but does not deploy them. Actual deployment — Kubernetes manifests, Secrets, GPU scheduling — lives in a separate repo: ~/repo/homelab, specifically manifests/home-services/*.yaml. That covers Deployments/Services for all four services here (namespace home-services), plus cluster-level pieces like nvidia-device-plugin.yaml and sealed secrets (*-sealed.yaml manifests generated from *-secret.sh scripts, one pair per service that needs one — tts-gateway doesn't have one yet, since unlike ha-gateway/discord-bot it currently needs no external tokens/credentials).

When a question is about how something actually behaves in production — not just what the code does — check that repo before guessing or assuming something isn't deployed; this repo alone doesn't show the full picture. Specifics of what's actually in tts-gateway.yaml there (current as of 2026-07-25 — it's maintained independently of this repo, so verify against the live file rather than trusting this description to stay accurate):

  • tts-gateway and tts-sidecar run as two containers in one Pod, not separate Deployments — the Go gateway reaches the sidecar over localhost:50054, mirroring the --network host pattern documented in tts-gateway/README.md for local Docker testing.
  • GPU scheduling: runtimeClassName: nvidia, nodeSelector: nik4nao.com/gpu: "true", and nvidia.com/gpu: 1 request/limit on the sidecar container. strategy: Recreate instead of the default RollingUpdate — nik-gpu only has one allocatable GPU, so a rolling update would deadlock waiting for a GPU still held by the pod it's replacing.
  • Model artifact distribution (checkpoint + hparams) is an emptyDir volume populated at pod start by an initContainers entry (model-init) that copies from a small, versioned image (gitea.nik4nao.com/nik/tts-model:<tag>, built from tts-gateway/model/Dockerfile) — not a raw hostPath into the node's disk anymore. That model image still can't be built by CI (the checkpoint isn't committed to git, ~455MB, and only ever existed on nik-gpu's local disk) — it's built and pushed manually, on nik-gpu, whenever the checkpoint/config change. See tts-gateway/README.md for the exact commands.
  • mTLS (TLS_DIR) is currently commented out on tts-gateway, for plaintext grpcurl testing from outside the cluster during initial rollout — a live TODO, not a permanent decision, unlike every other service here which assumes mTLS-or-trusted-network as its only access boundary (see "Configuration Notes" above).

CI

.gitea/workflows/ci.yaml always runs go vet/go test for all five modules (gen, ai-gateway, ha-gateway, discord-bot, tts-gateway) on every push/PR. On pushes to main, a changes job (dorny/paths-filter) determines which of the five images actually need rebuilding based on which paths changed, so an edit scoped to one service doesn't rebuild (and re-push) all of them — gen/, go.work, and go.work.sum count as shared and mark every Go-based image as changed, since a dependency bump there can affect all of them. Each build-* job is gated on that output and, when it runs, uses docker/build-push-action's registry-based cache (cache-from/cache-to: type=registry,ref=.../<image>:buildcache) so unchanged Docker layers (e.g. go mod download, apt-get/pip install) don't get redone on every run — the first run after adding this has nothing to pull from and builds fully fresh, subsequent ones should be much faster. tts-sidecar (the Python/CUDA inference sidecar under tts-gateway/sidecar/) is a large ~13GB image; it builds fine on a generic runner since only running it needs a GPU, not building it.

Note that ai-gateway/Dockerfile, ha-gateway/Dockerfile, discord-bot/Dockerfile, and tts-gateway/Dockerfile each COPY every other service directory (not just their own) because go.work lists all five Go modules as workspace members — Go's workspace-mode module resolution needs every listed directory present in the build context, even ones a given service doesn't otherwise depend on. Adding a new module to go.work means adding a matching COPY line (both the manifest-only and full-source copies, see below) to the other Dockerfiles too, or their builds break. Each of those four Dockerfiles also copies every module's go.mod/go.sum first and runs go mod download before copying full source, so that layer's cache survives source-only edits instead of being invalidated by every commit.