Guilgo Blog

Notes from my daily work with technology.

In February we published how to audit Kubernetes with Wazuh: the API server posts events to a webhook, the webhook injects them into analysisd’s socket, and Wazuh fires rules plus Telegram. Months later, before switching the pipeline back on in the homelab (k3s + Wazuh-lite in Docker), the base guide was still the map—but the terrain had moved.

TL;DR: Kubernetes audit into Wazuh does not fail only because “the webhook was never configured.” It fails when the policy records control-plane noise, when the tutorial port is already taken, or when the listener treats an EventList as a single datagram.

useful audit ≠ log everything

This post is that evolution: what we had, what broke, what we changed, and how to roll back without drama. It complements the February guide; it does not replace it.

Wazuh-lite series: Parental control · Grafana · Kubernetes audit (base guide) · Proactive SOC · AUR malware · urlscan · Docker listener · This post

How the pipeline aged

Stage When What we had What we learned
1. Base guide Feb 2026 Flask webhook, catch-all Metadata policy, port 8080, rules 110003–110006, Telegram API server → Wazuh works; create/delete can alert
2. k3s homelab Mar–Aug 2026 k3s via kube-apiserver-arg, Docker overlay, more services on the same host Port 8080 stops being “the free tutorial port”
3. Pre-reactivation review Oct 2026 Root disk at 72%, noisy policy, webhook sending whole EventLists Turning it on “as documented” risked DiskPressure and dropped events
4. Hardening Oct 2026 First-match policy, port 8088, one event per message + heavy-field truncation Useful volume (thousands/day), not hundreds of thousands; checklist + rollback

The through-line is not more features. It is less noise, more signal, and a k3s restart that does not leave you without a cluster.

Architecture (quick reminder)

k3s (API server)
  → audit-policy.yaml (what gets recorded)
  → audit-webhook.yaml (HTTPS to the listener)
    → custom-webhook.py (Flask, port 8088)
      → analysisd socket (location "k8s")
        → rules 110003–110006
          → custom-k8s-audit-telegram.py → Telegram

Our setup: single-node k3s (raam), Wazuh-lite in Docker, webhook as a service in the docker-compose.k8s-audit.yml overlay.

1. The catch-all that eats the disk

The original policy ended like this:

# Catch-all
- level: Metadata
  omitStages:
    - RequestReceived

Looks harmless. On a real cluster it records, among other things:

Resource / verb Why it hurts
leases (leader election) hundreds of events per minute
events the cluster audits itself
endpoints / endpointslices continuous service reconciliation
get / list / watch constant controller reads

In the tests and estimates for this homelab (a small node with ~30 pods and active CronJobs), that catch-all lands on hundreds of thousands of events per day. At ~1 KB of JSON each, that is easily hundreds of MB/day if audit goes to disk—or a flood of the socket/webhook if you only use the webhook. Not a universal figure for “any 30-pod cluster”: it depends on controllers, CronJobs, and load. Here it was enough not to enable the policy as-is.

With root at 72% (operational threshold 75%), switching that policy on without filtering was not a security win. It was a DiskPressure risk.

Hardened policy (first-match)

apiVersion: audit.k8s.io/v1
kind: Policy
rules:
  - level: None
    nonResourceURLs: ['/healthz*', '/logs', '/metrics', '/swagger*', '/version']

  - level: None
    resources:
      - group: coordination.k8s.io
        resources: [leases]
      - group: ''
        resources: [events, endpoints]
      - group: discovery.k8s.io
        resources: [endpointslices]

  - level: None
    verbs: [get, list, watch]

  - level: Metadata
    omitStages: [RequestReceived]
    resources:
      - group: authentication.k8s.io
        resources: [tokenreviews]

  - level: RequestResponse
    omitStages: [RequestReceived]
    resources:
      - group: authorization.k8s.io
        resources: [subjectaccessreviews]

  - level: RequestResponse
    omitStages: [RequestReceived]
    resources:
      - group: ''
        resources: [pods]
        verbs: [create, patch, update, delete]

  # Catch-all: mutations only (get/list/watch already dropped)
  - level: Metadata
    omitStages: [RequestReceived]

In this homelab, useful volume drops to the order of thousands of events/day (mutations + auth/authz), not hundreds of thousands.

Rule order matters

Kubernetes evaluates the policy first-match: the first matching rule wins; later rules are ignored. The YAML is not a preference list—it is a funnel:

  1. Drop noise first (None: healthz, leases, events, endpoints*, get/list/watch).
  2. Then raise what you care about with an explicit level (tokenreviews, subjectaccessreviews, pod mutations).
  3. Finally the catch-all Metadata only sees what did not match earlier—in practice, mutations of everything else.

Put the catch-all on top, or bury the verb None under a more specific rule that never gets reached, and the whole design collapses.

Why these levels (not “the correct policy”)

Rule Level Why
tokenreviews Metadata Knowing who authenticated is enough; the full body is mostly noise and volume.
subjectaccessreviews RequestResponse The signal is what authorization was attempted and what the decision was (allowed: true/false in the response); without the body you lose forensic value.
pod mutations RequestResponse create/patch/update/delete on pods is the usual operational impact; the object helps triage.
catch-all Metadata Other mutations: who/what/where without bloating the payload.

Globally dropping get / list / watch is deliberate: it prioritizes mutations and cuts noise. It also means a potentially sensitive read (Secrets, ConfigMaps, and so on) is not audited under this profile. If you need traceability for specific reads, add targeted rules before that verb None. This is not “the correct policy” in the abstract—it is a policy tuned for a homelab whose goal is to alert on changes, not to record every controller list.

Before vs after (policy)

Before (base guide) After (hardening)
Catch-all Metadata with no noise exclusions leases / events / endpoints* / get|list|watch at level: None
Every read and mutation hits the webhook Mutations only (+ tokenreviews / SAR at an explicit level)
Risk of filling the disk or saturating analysisd Volume capped to forensic interest

2. Port 8080 was already taken

The guide uses https://IP:8080. On this host that port belonged to another container (qbittorrent in our case). The API server would never have delivered events even with a perfect policy.

What we changed:

  • WEBHOOK_PORT=8088
  • audit-webhook.yaml → https://10.5.5.3:8088
  • compose overlay mapped to the same port

Before enabling: ss -tlnp | grep 8088 should be empty, and webhook certs generated (gen-certs.sh).

3. A whole EventList ≠ one event for analysisd

Kubernetes posts JSON like:

{
  "apiVersion": "audit.k8s.io/v1",
  "kind": "EventList",
  "items": [ { "...event1..." }, { "...event2..." } ]
}

The original webhook did a single send() of the entire object to Wazuh’s Unix socket (SOCK_DGRAM). As the batch grows, the message can exceed the effective limits of the socket or of Wazuh (OS, buffers, and config all matter) and get dropped or fail on send.

Fix: iterate items[] and send one event per message. That is the important part: each audit event reaches analysisd as its own datagram. If a single event is still huge (e.g. a Secret’s requestObject), truncate heavy fields and keep enough metadata for the rules (verb, requestURI, user, and so on).

items = data.get("items") or [data]
for item in items:
    send_to_wazuh_socket(item)  # "1:k8s:{json}"

Local rules (110003 create, 110004 delete, 110005 patch, 110006 update) still match on "verb": "create" (etc.) with PCRE2, without depending on field order inside a full EventList.

Before vs after (webhook)

Before After
One send() of the whole EventList One message per items[] element
Default port 8080 Port 8088 (via WEBHOOK_PORT)
No explicit size cap Truncate requestObject / responseObject above ~30 KB

4. Restarting k3s and rolling back

Enabling audit means editing /etc/rancher/k3s/config.yaml and restarting k3s. On a single node that takes every pod down for a few minutes.

Typical block:

kube-apiserver-arg:
  - audit-policy-file=/var/lib/rancher/k3s/server/audit-policy.yaml
  - audit-webhook-config-file=/var/lib/rancher/k3s/server/audit-webhook.yaml
  - audit-webhook-batch-max-size=1
  # Optional if you ever use audit-log-path:
  # - audit-log-maxsize=100
  # - audit-log-maxbackup=3
  # - audit-log-maxage=7

Rollback if the API server will not start:

  1. Remove the audit kube-apiserver-arg block (or the bad flags).
  2. sudo systemctl restart k3s
  3. kubectl get nodes → Ready

Comment the block clearly (# Kubernetes audit -> Wazuh) so you can find it in two seconds at 3 a.m.

5. Checklist before you enable

Step Pass criteria
Policy without a noisy catch-all leases / events / endpoints* / get|list|watch at level: None
Free port e.g. 8088, not the tutorial’s default 8080
Webhook TLS certs gen-certs.sh run; server.crt / server.key present
Webhook up and reachable HTTPS POST to the listener responds
Root disk Reasonable headroom above your ops threshold
Smoke test kubectl create deploy audit-test --image=nginx + delete → alerts 110003/110004

6. What this post does not cover

  • Filtering controller service accounts in n8n/IRIS (SOC triage, not audit policy).
  • Replacing Telegram with a full SOAR pipeline.
  • Local audit log files on k3s: if you enable them, configure rotation (maxsize / maxbackup / maxage) from day one.

Verifying in the homelab

# After policy + webhook + k3s are live
kubectl get nodes
kubectl create deployment audit-test --image=nginx
sleep 5
kubectl delete deployment audit-test

# In Wazuh: alerts rule_id 110003 / 110004
# In Telegram (if custom-k8s-audit-telegram.py is on): create/delete message

Reference code: wazuh-lite/k8s-audit/ (audit-policy.yaml, custom-webhook.py, setup-k3s-audit.sh, overlay docker-compose.k8s-audit.yml).

Closing

useful audit ≠ log everything

The February guide laid the path. Eight months later, the same homelab wrote that line on the wall. Harden the policy, measure volume, send one audit event per datagram, and document the k3s rollback. After that, February is still the map; this post is what to check before you hit restart.