dumpnet-argo/.goosehints

218 lines
12 KiB
Text
Raw Normal View History

2026-09-11 11:03:24 -04:00
# .goosehints for dumpnet-argo
This is a personal single-node Talos/Kubernetes cluster on AWS EC2, managed via
ArgoCD GitOps, migrating services off a legacy NixOS box (dumpnet.chat /
git.keane.sh). Below are conventions and hard-won lessons — follow them
before improvising a new pattern.
## Core architecture
- Single EC2 node (t3.medium), Talos Linux, one EIP, one cluster — this will
**never** be multi-node. Don't suggest multi-node solutions (NLB, EFS for
RWX, etc.) as the default; hostPath/local-disk patterns are correct here.
- Terraform (`terraform/`) owns all AWS infra: EC2, EIP, subnet, security
group, Talos bootstrap, Route53, S3, IAM, Secrets Manager, SES.
- ArgoCD App-of-Apps owns all Kubernetes-side resources. Root app is
`apps/apps.yaml`, which points at the `apps/` directory itself (self-managing,
only needs one manual `kubectl apply` ever, via `make bootstrap`).
- Groups under `apps/`: `cluster.yaml` (platform infra), `data.yaml` (shared
stateful stuff like Postgres), `services.yaml` (user-facing apps), `mcp.yaml`
(MCP servers). Manifests live in matching `manifests/<group>/` dirs. Add new
groups here rather than growing one flat directory.
- One global `values.yaml` at repo root holds `clusterName`, `domain`,
`repoURL`, `certEmail`, `awsRegion`, `registry.*`. Charts read it via the
multi-source `$values/values.yaml` pattern. `repoURL` inside ArgoCD
`Application` specs themselves can NOT be templated (Helm doesn't touch
those fields) — document this when it matters, don't try to fix it.
## Secrets
- **Exactly one** AWS Secrets Manager secret: `dumpnet` (nested JSON by
service, e.g. `cluster`, `postgres`, `kan`, `forgejo`, `ses`,
`tailscale`, `mcp_auth_proxy`). Never create a second Secrets Manager
secret — Secrets Manager billing is per-secret and that was an explicit
decision to avoid.
- External Secrets Operator (ESO) is the only way secrets get into the
cluster. `ClusterSecretStore` auths via the EC2 instance IAM role — no
static credentials anywhere.
- Sensitive local files (`controlplane.yaml`, `worker.yaml`, `talosconfig`,
`terraform.tfstate`) are gitignored entirely — never SOPS-encrypted in-repo.
Decrypt-to-`/tmp` was rejected in favor of "never committed at all."
`terraform.tfvars` is intentionally NOT gitignored (values aren't
sensitive) — don't re-add it to `.gitignore`.
- When adding a new service's secrets: add the key(s) under the service's
namespace in the `dumpnet` JSON blob (via `aws secretsmanager
get-secret-value` → merge in Python → `put-secret-value`), then add an
`ExternalSecret` template in that chart's `templates/` dir referencing
`property: <service>.<key>`.
## ArgoCD / Helm patterns (learned the hard way)
- **Use charts as-is.** Don't split an app into multiple ArgoCD Applications
just because another unrelated chart happens to be split that way (this
was explicitly called out as "cargo culting" — cert-manager's
ClusterIssuer split is justified due to CRD ordering; ESO's original split
was not and got merged back into one app).
- Extra resources (ClusterSecretStore, Issuers, Namespaces, ExternalSecrets)
belong in that chart's own `templates/` directory, not a second
Application, unless there's a *proven* cross-app ordering problem.
- Any namespace that needs `hostNetwork`, `hostPath`, or other privileged
pod behavior (ingress-nginx, fluent-bit, tailscale, mcp-auth-proxy) needs a
`Namespace` manifest in that chart's `templates/` with:
```yaml
metadata:
labels:
pod-security.kubernetes.io/enforce: privileged
pod-security.kubernetes.io/audit: privileged
pod-security.kubernetes.io/warn: privileged
```
Do **not** create these namespaces via Terraform `kubernetes_namespace` —
that was tried and reverted; it causes import/ordering pain on rebuilds.
ArgoCD-managed namespace manifests are the correct pattern.
- ESO `ExternalSecret` `.target.template` uses its own Go-template syntax
that Helm will try to consume if you're not careful. To keep a value
literal for ESO to interpolate at runtime while still letting Helm
interpolate `.Values.*`, use the `{{ "{{" }}` / backtick escaping pattern
already present in `charts/*/templates/*.yaml` — copy that pattern rather
than reinventing it.
- Watch out for fields ESO/operators add as defaults after creation
(`mergePolicy`, `engineVersion`, etc.) causing perpetual ArgoCD
OutOfSync — either match them explicitly in the manifest or use
`ignoreDifferences` in the Application spec.
## Storage
- Single shared hostPath-backed PV (`appdata`, ReadWriteMany, `Retain` reclaim
policy) at `/var/local/appdata` on the node. Services claim subpaths via
their own PVCs bound to that PV — this is the deliberate "Unraid appdata"
equivalent for a single-node cluster. Don't provision per-service EBS
volumes or suggest EFS/CSI drivers unless the user explicitly wants to
outgrow single-node.
- PV `capacity` on hostPath volumes is a label, not an enforced quota — don't
be alarmed if it's set higher than the actual disk size, but don't blow
past real disk space either.
## Networking / TLS
- ingress-nginx runs as a DaemonSet in `hostNetwork` mode (binds 80/443
directly on the node) — no LoadBalancer/NLB. This was a deliberate
cost-saving choice for single-node.
- cert-manager + Let's Encrypt (`letsencrypt-prod` ClusterIssuer, HTTP01 via
ingress-nginx) for all public hostnames. A separate `selfsigned` Issuer
handles internal webhook certs (e.g. ESO) — never try to get Let's Encrypt
certs for internal `*.svc.cluster.local` names, it will always fail with
"domain needs a public suffix."
- Tailscale operator is used for private/internal access only (e.g. DBeaver
→ Postgres), annotated per-Service with `tailscale.com/expose: "true"` and
`tailscale.com/hostname: dumpnet-<service>`. `ingressClass` is disabled on
the operator so it doesn't try to manage public ingresses too.
- DNS records are Terraform-managed (`dns_records` list in
`terraform.tfvars`) — add new subdomains there, not manually in Route53
console.
## Container images
- Prefer official upstream images. When a custom image must be built (e.g.
`mcp-auth-proxy`), it lives under `images/<name>/Dockerfile` in this same
repo — no separate repo for a thin wrapper Dockerfile.
- Use multi-stage builds; final stage should be `scratch` or
`distroless` for Go binaries — keep images small, self-hosted Forgejo has a
registry size limit (`nginx.clientMaxBodySize` + Forgejo's own
`packages.MAX_BLOB_SIZE`, both had to be raised from defaults on the
NixOS box).
- Custom images push to the **Forgejo container registry**
(`forge.keane.sh/ian/<image>`), not ECR. ECR was set up once and
deliberately removed — an EC2/Talos node can't easily pull from ECR without
baking a system extension into a custom AMI via Image Factory, which is
more infra than it's worth for a personal cluster. Forgejo + the node's
default pull path works fine and stays consistent with "no unnecessary
centralized services."
- `registry.host` / `registry.user` are values in the root `values.yaml` —
reference them, don't hardcode `forge.keane.sh`/`ian` in new charts.
- For any service whose image you personally build/push (e.g. `zoitestream`,
`mcp-auth-proxy`, `repertory-api`) — as opposed to an official upstream
image — set `imagePullPolicy: Always` on that container. These use
floating `:latest` tags with no digest pinning, and this is a single-node
cluster where Kubernetes will otherwise happily reuse a stale cached
image after you push a new one, requiring a manual `kubectl delete pod`
to force a repull. `imagePullPolicy: Always` makes every pod
restart/reschedule actually check the registry.
2026-09-11 11:03:24 -04:00
## Databases
- One shared Postgres in the `data` group/namespace (`postgres`), used by
multiple services (kan, future services). Don't spin up a dedicated
Postgres per service.
- New service databases: connect via DBeaver (Tailscale-exposed
`dumpnet-postgres:5432`) and `CREATE DATABASE <name>;` manually — there's
no automated database-per-service provisioning yet. Password comes from
`postgres.password` in the `dumpnet` secret (same user/password works for
every DB on the instance, it's one Postgres server).
- Migration Jobs (e.g. `kan-migrate`) should NOT use ArgoCD `PreSync` hooks —
hooks and sync-wave ordering don't compose the way you'd expect ExternalSecret
creation vs. hook timing. Use plain `Job` resources with `sync-wave`
annotations instead (ExternalSecret at wave `-1`, migration Job at wave
`1`), and set `ttlSecondsAfterFinished` so completed jobs clean themselves up.
## Email
- Outbound email (magic links etc.) goes through SES SMTP
(`email-smtp.us-east-1.amazonaws.com`), domain `dumpnet.chat` (verified via
Terraform-managed DKIM/TXT records). `keane.sh` is reserved for the
Fastmail alias — don't touch its DNS or try to send from `@keane.sh`.
- SES starts in sandbox mode — verify individual recipient addresses via
`aws ses verify-email-identity` until production access is granted.
- Check upstream app source (not assumptions) for the actual expected env
var names before wiring up SMTP — e.g. kan uses `SMTP_HOST` /
`SMTP_PORT` / `SMTP_USER` / `SMTP_PASSWORD` / `SMTP_SECURE` /
`EMAIL_FROM`, not `EMAIL_SERVER_*`. Wrong var names fail silently/weirdly
(e.g. connects to `127.0.0.1:465` when nothing is configured) rather than
erroring clearly.
## Operational workflow
- `make apply` is a **two-phase** terraform apply (first targets
`talos_cluster_kubeconfig.this` and its deps, then a full apply) — the
Terraform `helm`/`kubernetes` providers can't initialize before the
cluster exists. Don't try to collapse this into one `terraform apply`.
- After any fresh `apply`: `make post-apply` (waits for node, imports
ingress-nginx namespace state if needed, runs `make bootstrap`, builds/pushes
any custom images). Then DNS/cert issuance takes a few minutes.
- `make clean` before `make destroy` when doing a full from-zero rebuild —
force-deletes the Secrets Manager secret (avoids the 30-day recovery
window blocking recreation) and clears stale kubectl/talosctl contexts.
It does NOT touch Route53 (Terraform-managed, fine to destroy/recreate).
- After any pod-affecting change to an `ExternalSecret`, the **existing**
pod does not pick up new secret data automatically — force ESO to
re-sync (`kubectl annotate externalsecret <name> -n <ns> force-sync=$(date
+%s) --overwrite`) and then delete/restart the pod. `kubectl rollout
restart` alone is not reliably enough — prefer `kubectl delete pod` for a
clean restart, especially if the deployment spec itself changed and
ArgoCD hasn't synced yet (check `kubectl get deployment ... -o
jsonpath='{.spec.template.spec.containers[0].args}'` to confirm before
assuming a restart will fix anything).
- When something is "OutOfSync" in ArgoCD for no apparent reason, actually
look at the diff (`kubectl describe app <name> -n argocd`) rather than
guessing — several past sessions burned significant time on guessed fixes
(mergePolicy, engineVersion, type placement) before checking the real
diff.
- Prefer diagnosing root cause over reflexive retries. This repo's history
includes multiple incidents where the same command was rerun repeatedly
without new information — check logs/state first, form a hypothesis,
test it once.
## Documentation
- Keep `README.md` up to date for: first-time setup, secrets approach,
day-to-day ops, Makefile reference, forking/multi-environment notes. If a
new operational gotcha is discovered (like the two-phase apply, or
`make clean`), add it to the README, not just this file.
2026-09-14 16:23:55 -04:00
## Git
- **Never run `git commit` or `git push`** (in this repo, `zoitestream`,
`repertory`, `repertory-api`, or any other repo) unless the user
explicitly asks for it in that specific message. Staging/diffing is
fine; committing/pushing is the user's call, always. This applies even
after making a series of edits the user clearly wants kept — stop and
let them commit.