home-services/CLAUDE.md
Nik Afiq 7d9a75a3e7
All checks were successful
CI / changes (push) Successful in 19s
CI / test (push) Successful in 23s
CI / build-ai-gateway (push) Has been skipped
CI / build-ha-gateway (push) Has been skipped
CI / build-discord-bot (push) Has been skipped
CI / build-tts-gateway (push) Successful in 54s
CI / build-tts-sidecar (push) Has been skipped
CI / build-tts-model (push) Successful in 39s
feat: add tts-model to CI workflow and Dockerfile, include model files
2026-07-25 02:31:44 +09:00

193 lines
14 KiB
Markdown

# CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
## Repo Overview
`home-services` is a Go workspace of four internal services for home control, connected by gRPC and sharing committed protobuf-generated code:
```text
Discord users
|
v
discord-bot -----> ha-gateway -----> Home Assistant REST API
|
v
ai-gateway ------> Ollama
|
v
ha-gateway
discord-bot -----> tts-gateway -----> tts-sidecar (GPU inference)
```
- **ha-gateway** (port `50051`) — gRPC boundary for Home Assistant. Talks to HA's REST API; implements entity state, light control/discovery, switch control/discovery, and climate (HVAC) control/discovery; also relays SwitchBot Cloud remote commands (`RemoteApp`) when `SWITCHBOT_TOKEN`/`SWITCHBOT_SECRET` are set. Event streaming is stubbed.
- **ai-gateway** (port `50052`) — gRPC service that turns free-form text into home actions. Calls Ollama for intent extraction, resolves intents against a cached light list from `ha-gateway`, and calls `ha-gateway` to execute approved actions.
- **discord-bot** — registers `/light`, `/switch`, `/ac`, `/ai`, `/speak` slash commands and calls `ha-gateway`/`ai-gateway`/`tts-gateway` via gRPC clients. `/speak` synthesizes speech via `tts-gateway` (which returns AAC) and transcodes it to Opus locally (via `ffmpeg`/`github.com/jonas747/dca`) to stream into the invoking user's current voice channel — see `internal/adapters/primary/discord/voice.go`.
- **tts-gateway** (port `50053`) — gRPC text-to-speech service (VITS voice model, 92 Umamusume voices), paired with a Python/libtorch inference sidecar (`tts-gateway/sidecar/`, deployed as a second container in the same pod) that needs an Nvidia GPU — see `TTS_GATEWAY_PLAN.md` for why. Called by `discord-bot`'s `/speak` command. See `tts-gateway/README.md` for its API/config/local-run details.
Each service is a separate Go module (own `go.mod`) joined by `go.work` at the root, plus a `gen` module for shared generated code. Module paths are `gitea.nik4nao.com/nik/home-services/{ha-gateway,ai-gateway,discord-bot,tts-gateway,gen}`.
## Architecture (hexagonal, per service)
Every service follows the same internal layout, with dependencies pointing inward:
```text
cmd/<entrypoint>/ # process entrypoint and wiring (loads .env, builds adapters, starts gRPC)
internal/adapters/primary/ # inbound edges: gRPC servers, Discord handlers
internal/adapters/secondary/ # outbound edges: HA REST client, Ollama client, ha-gateway/ai-gateway gRPC clients
internal/app/ # use-case orchestration
internal/core/domain/ # domain types
internal/core/ports/ # driving (inbound) and driven (outbound) interfaces
internal/config/ # environment loading
internal/logger/ # slog setup
internal/telemetry/ # OpenTelemetry setup
```
`internal/core` has no dependency on adapters — ports are interfaces that adapters implement (driven) or call into (driving). When adding a capability, the usual path is: define/extend a port in `core/ports`, implement orchestration in `app`, then wire an adapter in `adapters/primary` or `adapters/secondary`.
Protobuf contracts live in `proto/` (buf module, `ai/v1`, `ha/v1`, and `tts/v1` packages). Generated Go code is committed under `gen/` and consumed by all four services through the Go workspace — do not hand-edit files in `gen/`.
## `tmp/` is reference-only, never a dependency
`tmp/` is gitignored — nothing under it is pushed to git, and it should be treated as temporary scratch space for reference material (e.g. `tmp/reference/switchbot-control-reference/`, a standalone CLI copied in for local discovery/testing against an external API). Rules:
- It's fine to read code under `tmp/` for patterns, to run its scripts/tools locally (a discovery script, a CLI, etc.), or to use it as a manual testing aid.
- Never make any of the four services (`ha-gateway`, `ai-gateway`, `discord-bot`, `tts-gateway`) import, `go.work use`, or otherwise depend on anything under `tmp/` at build or runtime. Since `tmp/` isn't committed, that dependency would silently break for every other clone of the repo (including CI). `tts-gateway/sidecar/` is a real example of doing this correctly — its model code was copied out of `tmp/reference/uma-tts-api/` into a committed location rather than referenced in place.
- If a reference tool under `tmp/` lives in its own Go module nested inside this repo's `go.work` workspace, invoke it with `GOWORK=off` rather than adding it to the root `go.work` — e.g. `GOWORK=off ./scripts/some-tool ...` from that tool's own directory.
- If something under `tmp/` turns out to be genuinely needed at runtime, port the actual logic into the relevant service's `internal/` tree (following the hexagonal layout above) instead of reaching into `tmp/` from committed code.
## Common Commands
Regenerate protobuf code after changing anything under `proto/` (requires `buf`):
```bash
buf generate
```
Run tests / vet (from repo root, or `cd` into a service and drop the prefix):
```bash
go test ./ha-gateway/... ./ai-gateway/... ./discord-bot/... ./tts-gateway/...
go vet ./ha-gateway/... ./ai-gateway/... ./discord-bot/... ./tts-gateway/...
```
Run a single test:
```bash
cd ha-gateway && go test ./internal/app/... -run TestEntityAppGetState
```
Build binaries:
```bash
go build ./ha-gateway/... ./ai-gateway/... ./discord-bot/... ./tts-gateway/...
```
Run a service locally (each loads `.env` from its own working directory via `godotenv`, so `cd` into the service dir first):
```bash
cd ha-gateway && cp .env.example .env && go run ./cmd/gateway
cd ai-gateway && cp .env.example .env && go run ./cmd/gateway
cd discord-bot && cp .env.example .env && go run ./cmd/bot
```
`tts-gateway` additionally needs the `open_jtalk` CLI/dictionary/voice, `ffmpeg`, and a running
inference sidecar to actually synthesize anything — see `tts-gateway/README.md` for the full
local/nik-gpu run instructions rather than a short snippet here.
For local plaintext dev, point gateways at each other with `TLS_DIR` empty, e.g. `HA_GATEWAY_ADDR=localhost:50051`, `AI_GATEWAY_ADDR=localhost:50052`.
Build container images (from repo root, since Dockerfiles reference the whole workspace):
```bash
docker build -f ha-gateway/Dockerfile -t ha-gateway:dev .
docker build -f ai-gateway/Dockerfile -t ai-gateway:dev .
docker build -f discord-bot/Dockerfile -t discord-bot:dev .
docker build -f tts-gateway/Dockerfile -t tts-gateway:dev .
docker build -f tts-gateway/sidecar/Dockerfile -t tts-sidecar:dev tts-gateway/sidecar
```
gRPC smoke checks (with the relevant service running):
```bash
grpcurl -plaintext -d '{"domain":"light"}' localhost:50051 ha.v1.EntityService/ListStates
grpcurl -plaintext -d '{"entity_id":"light.living_room","brightness_pct":80}' localhost:50051 ha.v1.LightService/TurnOn
grpcurl -plaintext -d '{"text":"turn on the desk lamp","source":"local"}' localhost:50052 ai.v1.AIService/Query
grpcurl -plaintext -d '{}' localhost:50052 ai.v1.AIService/ListModels
```
`tts-gateway`'s smoke check needs the sidecar (and, for real inference, a GPU) up first — see
`tts-gateway/README.md` rather than duplicating the fuller sequence here.
Note: `ha-gateway` always registers gRPC reflection; `ai-gateway` only registers reflection when `LOG_LEVEL=debug`.
## Testing Conventions
Tests use the standard library `testing` package only (no testify). Mocks are hand-written structs implementing the relevant `core/ports/driven` interface with function fields (see `ha-gateway/internal/app/entity_test.go` for the pattern).
## Configuration Notes
- Every service reads `TLS_DIR` to enable optional mTLS; when set, the directory must contain `tls.crt`, `tls.key`, and `ca.crt`.
- `OTEL_ENDPOINT` enables OTLP gRPC traces/metrics; leave empty for local no-op telemetry.
- `LOG_FORMAT=json` is the production default; `text` is easier to read locally.
- `ha-gateway` reads `SWITCHBOT_TOKEN`/`SWITCHBOT_SECRET` (optional) to enable SwitchBot Cloud remote commands via `RemoteApp`; leave empty to disable that path.
- None of the services implement app-layer authorization — they rely on being kept on a trusted internal network or on mTLS. Keep this in mind before adding any endpoint that wasn't previously reachable.
## nik-gpu deployment target
`tts-gateway`'s inference sidecar (see `TTS_GATEWAY_PLAN.md` for the full history) runs on
`nik-gpu`, a remote Nvidia GPU host reachable via `ssh nik-gpu` — both for local dev/smoke-testing
(via the Claude Code skills `nik-gpu-status` (read-only check), `nik-gpu-sync` (rsync this repo to
`~/repo/home-service/` there), `nik-gpu-docker-build` (build/smoke-test via a `nik-gpu` Docker
context), and `tts-gateway/README.md`'s run instructions) and as the actual GPU node backing the
production Kubernetes deployment — see "Infrastructure / Deployment" below.
**Never automatically run installation or other host-system-altering commands on nik-gpu**
`apt`/`apt-get`, `pip install` outside a container, Docker daemon config changes, driver/toolkit
updates, or anything requiring `sudo`. Always print the exact command and ask the user to run it
themselves (their own terminal, or `! <command>` in a Claude Code session). This applies to the
bare nik-gpu host specifically; installing packages *inside* a Dockerfile build (e.g.
`apt-get install open-jtalk` as a build step) is a normal container build action, not a host
mutation, and is fine to run.
## Infrastructure / Deployment
This repo builds and pushes Docker images (see CI below) but does not deploy them. Actual
deployment — Kubernetes manifests, Secrets, GPU scheduling — lives in a **separate** repo:
`~/repo/homelab`, specifically `manifests/home-services/*.yaml`. That covers Deployments/Services
for all four services here (namespace `home-services`), plus cluster-level pieces like
`nvidia-device-plugin.yaml` and sealed secrets (`*-sealed.yaml` manifests generated from
`*-secret.sh` scripts, one pair per service that needs one — `tts-gateway` doesn't have one yet,
since unlike `ha-gateway`/`discord-bot` it currently needs no external tokens/credentials).
When a question is about how something actually behaves *in production* — not just what the code
does — check that repo before guessing or assuming something isn't deployed; this repo alone
doesn't show the full picture. Specifics of what's actually in `tts-gateway.yaml` there (current
as of 2026-07-25 — it's maintained independently of this repo, so verify against the live file
rather than trusting this description to stay accurate):
- `tts-gateway` and `tts-sidecar` run as **two containers in one Pod**, not separate Deployments —
the Go gateway reaches the sidecar over `localhost:50054`, mirroring the `--network host`
pattern documented in `tts-gateway/README.md` for local Docker testing.
- GPU scheduling: `runtimeClassName: nvidia`, `nodeSelector: nik4nao.com/gpu: "true"`, and
`nvidia.com/gpu: 1` request/limit on the sidecar container. `strategy: Recreate` instead of the
default `RollingUpdate` — nik-gpu only has one allocatable GPU, so a rolling update would
deadlock waiting for a GPU still held by the pod it's replacing.
- Model artifact distribution (checkpoint + hparams) is an `emptyDir` volume populated at pod
start by an `initContainers` entry (`model-init`) that copies from `gitea.nik4nao.com/nik/tts-model:latest`
— not a raw `hostPath` into the node's disk anymore. That image is built from
`tts-gateway/model/Dockerfile`, whose checkpoint/hparams are committed directly to git (a
deliberate exception to not committing large binaries — see that directory) and built/pushed
by CI's `build-tts-model` job like the other four images, path-filtered on `tts-gateway/model/**`.
- mTLS (`TLS_DIR`) is currently commented out on `tts-gateway`, for plaintext `grpcurl` testing
from outside the cluster during initial rollout — a live TODO, not a permanent decision, unlike
every other service here which assumes mTLS-or-trusted-network as its only access boundary (see
"Configuration Notes" above).
## CI
`.gitea/workflows/ci.yaml` always runs `go vet`/`go test` for all five Go modules (`gen`, `ai-gateway`, `ha-gateway`, `discord-bot`, `tts-gateway`) on every push/PR. On pushes to `main`, a `changes` job (`dorny/paths-filter`) determines which of the six images (`ai-gateway`, `ha-gateway`, `discord-bot`, `tts-gateway`, `tts-sidecar`, `tts-model`) actually need rebuilding based on which paths changed, so an edit scoped to one service doesn't rebuild (and re-push) all of them — `gen/`, `go.work`, and `go.work.sum` count as shared and mark every Go-based image as changed, since a dependency bump there can affect all of them (`tts-model` isn't a Go module, so it's untouched by that rule — only `tts-gateway/model/**` triggers it). Each `build-*` job is gated on that output and, when it runs, uses `docker/build-push-action`'s registry-based cache (`cache-from`/`cache-to: type=registry,ref=.../<image>:buildcache`) so unchanged Docker layers (e.g. `go mod download`, `apt-get`/`pip install`) don't get redone on every run — the *first* run after adding this has nothing to pull from and builds fully fresh, subsequent ones should be much faster (`build-tts-model` skips this, since a busybox+`COPY` image has no build steps worth caching). `tts-sidecar` (the Python/CUDA inference sidecar under `tts-gateway/sidecar/`) is a large ~13GB image; it builds fine on a generic runner since only *running* it needs a GPU, not building it. `tts-model`'s own build context (`tts-gateway/model/`) is ~455MB (the committed checkpoint) — a deliberate, one-time exception to this repo otherwise keeping large binaries out of git.
Note that `ai-gateway/Dockerfile`, `ha-gateway/Dockerfile`, `discord-bot/Dockerfile`, and `tts-gateway/Dockerfile` each `COPY` every other service directory (not just their own) because `go.work` lists all five Go modules as workspace members — Go's workspace-mode module resolution needs every listed directory present in the build context, even ones a given service doesn't otherwise depend on. Adding a new module to `go.work` means adding a matching `COPY` line (both the manifest-only and full-source copies, see below) to the other Dockerfiles too, or their builds break. Each of those four Dockerfiles also copies every module's `go.mod`/`go.sum` first and runs `go mod download` *before* copying full source, so that layer's cache survives source-only edits instead of being invalidated by every commit.