Both Deployments' rolling update briefly ran an old+new pod pair on the same
node (node-role: storage), sharing the same config storage (qbittorrent's PVC,
jdownloader's hostPath) -- both apps are effectively singletons that lock
their config directory, so the new instance conflicted with the still-running
old one:
- qbittorrent: hit a known qbittorrent:5.2.0 image bug (linuxserver/
docker-qbittorrent#432) where a stale WebUI lockfile prevents the server
from ever binding its port while a second instance is present -- surfaced
as the new pod's readiness probe getting "connection refused" indefinitely.
On top of that, my livenessProbe (30s/30s) was killing the container
(exitCode 137) before qbittorrent had any chance to come up at all. Dropped
the livenessProbe (readiness alone can never kill a container, only mark it
not-ready) and loosened the readinessProbe timing.
- jdownloader: the new pod's app process detected the old instance's lock in
the shared /data/jdownloader hostPath and exited cleanly (exitCode 0) rather
than run a duplicate copy -- nothing to do with probe timing.
Both Deployments now use `strategy: Recreate` instead of the RollingUpdate
default, so any future rollout fully stops the old pod before starting the
new one -- this is the actual fix (matches the pattern immich.yaml already
used). Applying this will briefly restart both currently-running pods to
verify it.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Stage 7 of REFACTOR_PLAN.md (findings #16). Covers dashy, glances,
ca-installer, authentik-proxy-outpost, jellyfin, qbittorrent/jdownloader main
containers (their gluetun sidecars already had probes), and all 4 Immich
Deployments -- previously none of these had any protection against one
workload starving another on this fixed-capacity cluster, nor automatic
restart on hang.
Values are sized from live `kubectl top pod` baselines gathered this session
(not guessed): e.g. Jellyfin/Immich-server were observed at ~3.1-3.3Gi
resident, so their limits give headroom above that (4Gi) rather than an
arbitrary round number. Used tcpSocket probes instead of httpGet wherever I
wasn't certain of an app's exact health-check path (Immich, Postgres/Redis),
to avoid a wrong path causing false probe failures on a live service.
This is Kubernetes-native and takes effect on next pod restart, but should
still be rolled out watching `kubectl top`/restart counts rather than pushed
and forgotten -- limits set too low can OOMKill under real load.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Stage 5 + part of Stage 6/7 of REFACTOR_PLAN.md. This is the highest-risk
stage per the plan -- these Applications are NOT to be pushed/synced blindly.
Each needs `kubectl diff` against live state one at a time before enabling.
New Applications (previously-live resources with zero GitOps coverage):
- cert-manager-config.yaml (manifests/cert-manager: both ClusterIssuers + the
internal CA Certificate -- every TLS cert in the cluster depends on these,
and nothing currently restores them on a cold rebuild).
- authentik-config.yaml (manifests/authentik: ingress, proxy outpost,
middleware -- raw manifests only, low risk).
- authentik.yaml (the Authentik Helm chart itself): sync is deliberately left
MANUAL and targetRevision is a REPLACE_ME placeholder -- I don't have a safe
way to read the live chart version (`helm list -n authentik`), and guessing
wrong risks an unwanted upgrade/downgrade of the SSO IdP gating Argo CD/
Grafana/Gitea logins. Needs your input before this one goes anywhere.
- network.yaml: widens coverage to the 4 non-sealed files in manifests/network
(ddns-cronjob, glances-debian-ingress, traefik-dashboard-ingress,
watch-party-ingress) that were previously invisible to Argo CD; keeps
network-secrets.yaml scoped to *-sealed.yaml only.
Fixes:
- homeassistant.yaml: destination.namespace was "homeassistant" (empty,
unused) while the actual resources are hardcoded to "default" -- corrected,
dropped CreateNamespace=true. The old empty namespace isn't auto-deleted
(prune: false); safe to remove by hand if desired.
- gitea-backup.yaml: added the missing Namespace object (nothing created
"gitea-backup" before); replaced a cluster-wide ClusterRole/ClusterRoleBinding
granting pods/exec everywhere with a Role/RoleBinding scoped to the `gitea`
namespace, matching what the backup script actually execs into. NOTE: this
is already under active sync via gitea-secrets.yaml (selfHeal: true,
prune: false) -- once pushed, the old ClusterRole/ClusterRoleBinding will
need manual `kubectl delete` since Argo CD won't prune them.
- Added sync-wave "-2" to cert-manager/sealed-secrets Applications so their
CRDs land before consumers (matches the existing -1/0 wave pattern).
- Normalized targetRevision HEAD -> main on home-services/otel-collector/tempo.
- Normalized sync policy per your decision: home-services/otel-collector/tempo
prune true -> false; pihole/pihole-debian selfHeal false -> true (repo-wide
consistency, per your call on finding #18).
Verified: kubeconform valid across all manifests + Argo CD Application objects.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Stage 1 of REFACTOR_PLAN.md. values/gitea.yaml and config/dashy/conf.yaml now
reference secrets injected at apply-time (gitea-postgres-secret.sh, .env) instead
of hardcoding a live DB password and weather API key in git. Both values must be
treated as compromised and rotated by the operator (see .env.example).
Also fixes authentik-ingress.yaml and traefik-dashboard-ingress.yaml, which
pointed at the internal-ca root ClusterIssuer instead of internal-ca-issuer,
the chained issuer every other internal Certificate uses -- causing untrusted-cert
warnings on the SSO login and Traefik dashboard.
Extends .gitignore for *.retry, .vault_pass*, kubeconfig patterns, and editor
swap files.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
feat: update qBittorrent deployment to expose gluetun API on port 8000 and add TLS certificate for secure access
feat: add gluetun DNS entry to Pi-hole configuration for improved network management