In February we published how to audit Kubernetes with Wazuh: the API server posts events to a webhook, the webhook injects them into analysisd’s socket, and Wazuh fires rules plus Telegram. Months later, before switching the pipeline back on in the homelab (k3s + Wazuh-lite in Docker), the base guide was still the map—but the terrain had moved.
TL;DR: Kubernetes audit into Wazuh does not fail only because “the webhook was never configured.” It fails when the policy records control-plane noise, when the tutorial port is already taken, or when the listener treats an EventList as a single datagram.
useful audit ≠ log everything
This post is that evolution: what we had, what broke, what we changed, and how to roll back without drama. It complements the February guide; it does not replace it.
Wazuh-lite series: Parental control · Grafana · Kubernetes audit (base guide) · Proactive SOC · AUR malware · urlscan · Docker listener · This post
How the pipeline aged
| Stage | When | What we had | What we learned |
|---|---|---|---|
| 1. Base guide | Feb 2026 | Flask webhook, catch-all Metadata policy, port 8080, rules 110003–110006, Telegram |
API server → Wazuh works; create/delete can alert |
| 2. k3s homelab | Mar–Aug 2026 | k3s via kube-apiserver-arg, Docker overlay, more services on the same host |
Port 8080 stops being “the free tutorial port” |
| 3. Pre-reactivation review | Oct 2026 | Root disk at 72%, noisy policy, webhook sending whole EventLists |
Turning it on “as documented” risked DiskPressure and dropped events |
| 4. Hardening | Oct 2026 | First-match policy, port 8088, one event per message + heavy-field truncation | Useful volume (thousands/day), not hundreds of thousands; checklist + rollback |
The through-line is not more features. It is less noise, more signal, and a k3s restart that does not leave you without a cluster.
Architecture (quick reminder)
k3s (API server)
→ audit-policy.yaml (what gets recorded)
→ audit-webhook.yaml (HTTPS to the listener)
→ custom-webhook.py (Flask, port 8088)
→ analysisd socket (location "k8s")
→ rules 110003–110006
→ custom-k8s-audit-telegram.py → Telegram
Our setup: single-node k3s (raam), Wazuh-lite in Docker, webhook as a service in the docker-compose.k8s-audit.yml overlay.
1. The catch-all that eats the disk
The original policy ended like this:
# Catch-all
- level: Metadata
omitStages:
- RequestReceived
Looks harmless. On a real cluster it records, among other things:
| Resource / verb | Why it hurts |
|---|---|
leases (leader election) |
hundreds of events per minute |
events |
the cluster audits itself |
endpoints / endpointslices |
continuous service reconciliation |
get / list / watch |
constant controller reads |
In the tests and estimates for this homelab (a small node with ~30 pods and active CronJobs), that catch-all lands on hundreds of thousands of events per day. At ~1 KB of JSON each, that is easily hundreds of MB/day if audit goes to disk—or a flood of the socket/webhook if you only use the webhook. Not a universal figure for “any 30-pod cluster”: it depends on controllers, CronJobs, and load. Here it was enough not to enable the policy as-is.
With root at 72% (operational threshold 75%), switching that policy on without filtering was not a security win. It was a DiskPressure risk.
Hardened policy (first-match)
apiVersion: audit.k8s.io/v1
kind: Policy
rules:
- level: None
nonResourceURLs: ['/healthz*', '/logs', '/metrics', '/swagger*', '/version']
- level: None
resources:
- group: coordination.k8s.io
resources: [leases]
- group: ''
resources: [events, endpoints]
- group: discovery.k8s.io
resources: [endpointslices]
- level: None
verbs: [get, list, watch]
- level: Metadata
omitStages: [RequestReceived]
resources:
- group: authentication.k8s.io
resources: [tokenreviews]
- level: RequestResponse
omitStages: [RequestReceived]
resources:
- group: authorization.k8s.io
resources: [subjectaccessreviews]
- level: RequestResponse
omitStages: [RequestReceived]
resources:
- group: ''
resources: [pods]
verbs: [create, patch, update, delete]
# Catch-all: mutations only (get/list/watch already dropped)
- level: Metadata
omitStages: [RequestReceived]
In this homelab, useful volume drops to the order of thousands of events/day (mutations + auth/authz), not hundreds of thousands.
Rule order matters
Kubernetes evaluates the policy first-match: the first matching rule wins; later rules are ignored. The YAML is not a preference list—it is a funnel:
- Drop noise first (
None: healthz, leases, events, endpoints*,get/list/watch). - Then raise what you care about with an explicit level (
tokenreviews,subjectaccessreviews, pod mutations). - Finally the catch-all
Metadataonly sees what did not match earlier—in practice, mutations of everything else.
Put the catch-all on top, or bury the verb None under a more specific rule that never gets reached, and the whole design collapses.
Why these levels (not “the correct policy”)
| Rule | Level | Why |
|---|---|---|
tokenreviews |
Metadata |
Knowing who authenticated is enough; the full body is mostly noise and volume. |
subjectaccessreviews |
RequestResponse |
The signal is what authorization was attempted and what the decision was (allowed: true/false in the response); without the body you lose forensic value. |
| pod mutations | RequestResponse |
create/patch/update/delete on pods is the usual operational impact; the object helps triage. |
| catch-all | Metadata |
Other mutations: who/what/where without bloating the payload. |
Globally dropping get / list / watch is deliberate: it prioritizes mutations and cuts noise. It also means a potentially sensitive read (Secrets, ConfigMaps, and so on) is not audited under this profile. If you need traceability for specific reads, add targeted rules before that verb None. This is not “the correct policy” in the abstract—it is a policy tuned for a homelab whose goal is to alert on changes, not to record every controller list.
Before vs after (policy)
| Before (base guide) | After (hardening) |
|---|---|
Catch-all Metadata with no noise exclusions |
leases / events / endpoints* / get|list|watch at level: None |
| Every read and mutation hits the webhook | Mutations only (+ tokenreviews / SAR at an explicit level) |
| Risk of filling the disk or saturating analysisd | Volume capped to forensic interest |
2. Port 8080 was already taken
The guide uses https://IP:8080. On this host that port belonged to another container (qbittorrent in our case). The API server would never have delivered events even with a perfect policy.
What we changed:
WEBHOOK_PORT=8088audit-webhook.yaml→https://10.5.5.3:8088- compose overlay mapped to the same port
Before enabling: ss -tlnp | grep 8088 should be empty, and webhook certs generated (gen-certs.sh).
3. A whole EventList ≠ one event for analysisd
Kubernetes posts JSON like:
{
"apiVersion": "audit.k8s.io/v1",
"kind": "EventList",
"items": [ { "...event1..." }, { "...event2..." } ]
}
The original webhook did a single send() of the entire object to Wazuh’s Unix socket (SOCK_DGRAM). As the batch grows, the message can exceed the effective limits of the socket or of Wazuh (OS, buffers, and config all matter) and get dropped or fail on send.
Fix: iterate items[] and send one event per message. That is the important part: each audit event reaches analysisd as its own datagram. If a single event is still huge (e.g. a Secret’s requestObject), truncate heavy fields and keep enough metadata for the rules (verb, requestURI, user, and so on).
items = data.get("items") or [data]
for item in items:
send_to_wazuh_socket(item) # "1:k8s:{json}"
Local rules (110003 create, 110004 delete, 110005 patch, 110006 update) still match on "verb": "create" (etc.) with PCRE2, without depending on field order inside a full EventList.
Before vs after (webhook)
| Before | After |
|---|---|
One send() of the whole EventList |
One message per items[] element |
| Default port 8080 | Port 8088 (via WEBHOOK_PORT) |
| No explicit size cap | Truncate requestObject / responseObject above ~30 KB |
4. Restarting k3s and rolling back
Enabling audit means editing /etc/rancher/k3s/config.yaml and restarting k3s. On a single node that takes every pod down for a few minutes.
Typical block:
kube-apiserver-arg:
- audit-policy-file=/var/lib/rancher/k3s/server/audit-policy.yaml
- audit-webhook-config-file=/var/lib/rancher/k3s/server/audit-webhook.yaml
- audit-webhook-batch-max-size=1
# Optional if you ever use audit-log-path:
# - audit-log-maxsize=100
# - audit-log-maxbackup=3
# - audit-log-maxage=7
Rollback if the API server will not start:
- Remove the audit
kube-apiserver-argblock (or the bad flags). sudo systemctl restart k3skubectl get nodes→ Ready
Comment the block clearly (# Kubernetes audit -> Wazuh) so you can find it in two seconds at 3 a.m.
5. Checklist before you enable
| Step | Pass criteria |
|---|---|
| Policy without a noisy catch-all | leases / events / endpoints* / get|list|watch at level: None |
| Free port | e.g. 8088, not the tutorial’s default 8080 |
| Webhook TLS certs | gen-certs.sh run; server.crt / server.key present |
| Webhook up and reachable | HTTPS POST to the listener responds |
| Root disk | Reasonable headroom above your ops threshold |
| Smoke test | kubectl create deploy audit-test --image=nginx + delete → alerts 110003/110004 |
6. What this post does not cover
- Filtering controller service accounts in n8n/IRIS (SOC triage, not audit policy).
- Replacing Telegram with a full SOAR pipeline.
- Local audit log files on k3s: if you enable them, configure rotation (
maxsize/maxbackup/maxage) from day one.
Verifying in the homelab
# After policy + webhook + k3s are live
kubectl get nodes
kubectl create deployment audit-test --image=nginx
sleep 5
kubectl delete deployment audit-test
# In Wazuh: alerts rule_id 110003 / 110004
# In Telegram (if custom-k8s-audit-telegram.py is on): create/delete message
Reference code: wazuh-lite/k8s-audit/ (audit-policy.yaml, custom-webhook.py, setup-k3s-audit.sh, overlay docker-compose.k8s-audit.yml).
Closing
useful audit ≠ log everything
The February guide laid the path. Eight months later, the same homelab wrote that line on the wall. Harden the policy, measure volume, send one audit event per datagram, and document the k3s rollback. After that, February is still the map; this post is what to check before you hit restart.