Wazuh Detection Engineering, Part 4: Tuning Out Noise Without Going Blind
Measure Wazuh alert volume by rule, then cut it with level overwrites, level-0 child rules, agent-group scoping, and alert-level thresholds, while keeping the detections that matter.
Copy for your AI agent. A condensed version of this entry written as a prompt, so Claude, Cursor, or Copilot can apply it to your own codebase.
View prompt · Wazuh noise tuning playbook
Task: Reduce Wazuh alert volume to an actionable level without losing real detections.
Measure first (indexer query, last 7 days, top rules):
curl -sk -u admin:$PASS "https://indexer:9200/wazuh-alerts-*/_search" -H 'Content-Type: application/json' -d '{
"size": 0,
"query": {"range": {"timestamp": {"gte": "now-7d"}}},
"aggs": {"by_rule": {"terms": {"field": "rule.id", "size": 20},
"aggs": {"desc": {"terms": {"field": "rule.description", "size": 1}}}}}
}' | jq '.aggregations.by_rule.buckets[] | {rule: .key, count: .doc_count, desc: .desc.buckets[0].key}'Tuning tools, in order of preference:
- Level-0 child for a specific false positive:
<rule id="100500" level="0"><if_sid>5501</if_sid><field name="dstuser">backup-svc</field> <description>Suppress: backup service PAM session (ticket SEC-142)</description></rule> - Overwrite a stock rule’s level:
<rule id="5715" level="2" overwrite="yes">…same body…</rule> - Scope collection with agent groups (
/var/ossec/etc/shared/<group>/agent.conf). - Raise
<alerts><log_alert_level>only as a last resort; it hides everything below globally.
Rules: never edit /var/ossec/ruleset/; every suppression has a description with a reason and ticket; re-measure weekly; track “alerts per analyst per day” as the number you are managing.
The first three parts built detections. This one makes them survivable. A Wazuh deployment with fifty agents and the stock ruleset produces tens of thousands of alerts a day. Nobody reads tens of thousands of anything. Alerts that are not read are not detections; they are storage costs with a false sense of coverage attached.
Tuning is the discipline of deciding, rule by rule, what deserves a human, and encoding that decision so it survives upgrades and staff changes. It is not “turn down the level until the dashboard looks calm”. Done that way, you go blind and do not know it.
Measure before touching anything
The indexer already has the answer to “where is the volume coming from”. Ask it. Against the single-node lab, or a production indexer with read credentials:
curl -sk -u admin:"$INDEXER_PASS" "https://localhost:9200/wazuh-alerts-*/_search" \
-H 'Content-Type: application/json' -d '{
"size": 0,
"query": {"range": {"timestamp": {"gte": "now-7d"}}},
"aggs": {
"by_rule": {
"terms": {"field": "rule.id", "size": 20},
"aggs": {
"desc": {"terms": {"field": "rule.description", "size": 1}},
"level": {"max": {"field": "rule.level"}},
"agents": {"cardinality": {"field": "agent.id"}}
}
}
}
}' | jq -r '.aggregations.by_rule.buckets[] | "\(.doc_count)\t\(.key)\tL\(.level.value|floor)\t\(.agents.value) agents\t\(.desc.buckets[0].key)"' | column -t -s $'\t'
You get a table like this:
48213 5501 L3 47 agents PAM: Login session opened.
41877 5502 L3 47 agents PAM: Login session closed.
19406 2902 L7 12 agents New dpkg (Debian Package) installed.
12550 550 L7 9 agents Integrity checksum changed.
9931 31101 L5 3 agents Web server 400 error code.
4120 5715 L3 47 agents sshd: authentication success.
...
Two things are always true of this table. The top ten rules account for most of the volume. And most of them are level 3 to 7 “something normal happened” events that were useful once, on one host, for one investigation. Fix these ten and you have removed the majority of the noise without touching a single high-level detection.
Keep this query. Run it weekly. The number you are managing is not total alerts; it is alerts a human is expected to look at, per analyst, per day. Write that number down before you start so you can show the change.1
Tool one: the level-0 child
This is the most precise tool and should be the first thing you reach for. A stock rule fires for a case you have decided is fine. Instead of changing the stock rule, add a child that matches the specific case and sets level 0. The child wins because it is more specific, and the event is evaluated, matched, and discarded before it becomes an alert.
The backup service opens a PAM session on every host every night:
<!-- local_rules.xml -->
<group name="tuning,">
<rule id="100500" level="0">
<if_sid>5501, 5502</if_sid>
<field name="dstuser">^backup-svc$</field>
<description>Suppress: backup-svc PAM sessions are scheduled (SEC-142)</description>
</rule>
</group>
Rule 5501 still fires for every other user. Nothing about the stock ruleset changed. The suppression has an id, a description that names the reason and a ticket, and lives in a file that upgrades never touch. Six months from now someone can grep local_rules.xml for “backup-svc” and understand exactly why those events vanish.
Use the same shape for the web server’s 400 errors from the health checker, the package installs from the patching pipeline’s user, and the file integrity changes under a path that a deployment tool rewrites on schedule:
<rule id="100501" level="0">
<if_sid>550, 553, 554</if_sid>
<field name="file">^/opt/app/releases/</field>
<description>Suppress: /opt/app/releases is rewritten by deploys (SEC-150)</description>
</rule>
A level-0 child with only <if_sid> and no narrowing field silences the whole parent. That is a deletion with extra steps, and it will hide the one time the rule mattered. Every suppression narrows by user, path, address, process, or agent. If you cannot name what makes this case safe, you have not understood it yet.
Tool two: the overwrite
Sometimes the stock level is simply wrong for your environment, for every case. SSH authentication success at level 3 is a good default for an internet-facing box and pointless for a fleet where every login comes through a bastion you already log. For this you re-declare the rule with the same id and overwrite="yes":
<rule id="5715" level="2" overwrite="yes">
<if_sid>5700</if_sid>
<match>^Accepted|authenticated.$</match>
<description>sshd: authentication success.</description>
<group>authentication_success,pci_dss_10.2.5,gpg13_7.1,gpg13_7.2,gdpr_IV_32.2,hipaa_164.312.b,nist_800_53_AU.14,nist_800_53_AC.7,tsc_CC6.8,</group>
</rule>
Two cautions. The overwrite must carry the complete rule body, because it replaces the stock definition entirely; copy it from the stock file under /var/ossec/ruleset/rules/ and change only the level. And an overwrite at level 2 makes the event disappear from alerts entirely, because 2 is under the default alert_level of 3. That is usually what you want for “we log this elsewhere” events. For “keep it, but do not page” the answer is not a lower level; it is the next section.
Levels are a routing key, not a volume knob
Wazuh levels 0 to 15 are meant as a severity scale, and the tuning mistake is to treat them as a single dial. Treat them as routes instead. Decide once what each band means for your team and then tune rules into bands, not down to silence.
| Level | Meaning | Route |
|---|---|---|
| 0 to 2 | Evaluated, not stored | Nothing |
| 3 to 6 | Context. Useful during an investigation, never on its own. | Indexed, no notification |
| 7 to 9 | Worth a look today | Daily digest, ticket queue |
| 10 to 12 | Worth a look now | Chat channel, on-call during hours |
| 13 to 15 | Wake someone | Page |
The manager enforces the bottom band with alert_level and the top with email_alert_level in the <alerts> section of ossec.conf. The middle bands are enforced by whatever you integrate: the <integration> block can send levels 10 and up to Slack or a webhook, and the ticket digest is a scheduled indexer query.
<alerts>
<log_alert_level>3</log_alert_level>
<email_alert_level>13</email_alert_level>
</alerts>
<integration>
<name>slack</name>
<hook_url>https://hooks.slack.com/services/…</hook_url>
<level>10</level>
<alert_format>json</alert_format>
</integration>
Under this model, the brute-force rule from Part 2 at level 10 lands in the chat channel. The “login after brute force” rule at level 12 does too, with a different word in the description. If you want it to page, it goes to 13. The tuning conversation stops being “is this too noisy” and becomes “which band does this belong in”, which is a question a team can answer in a minute.
Setting log_alert_level to 7 makes the dashboard calm instantly and throws away every level 3 to 6 event on every host, including the PAM and sshd context you will want the next time you investigate a level 12. Prefer level-0 children for the noisy cases and leave the floor at 3.
Tool three: agent groups
Some noise is not a rule problem; it is a collection problem. The build servers generate thousands of package-install events a day because installing packages is their job. No rule tuning fixes that as cleanly as not collecting dpkg.log on that host class in the first place.
Agent groups push different configuration to different classes of host. Create a group, assign agents, and put a agent.conf in the group’s shared directory on the manager:
sudo /var/ossec/bin/agent_groups -a -g buildservers -q
sudo /var/ossec/bin/agent_groups -a -i 007 -g buildservers -q
<!-- /var/ossec/etc/shared/buildservers/agent.conf -->
<agent_config>
<!-- collect auth, not package churn -->
<localfile>
<log_format>syslog</log_format>
<location>/var/log/auth.log</location>
</localfile>
<!-- integrity monitoring only where it means something -->
<syscheck>
<directories check_all="yes" realtime="yes">/etc,/usr/bin,/usr/sbin</directories>
<ignore>/etc/mtab</ignore>
<ignore type="sregex">^/var/lib/docker/</ignore>
</syscheck>
</agent_config>
Agents pull the group configuration within a few minutes. The agent.conf merges with the agent’s local ossec.conf, so anything the agent already collects locally keeps flowing; to stop a noisy local localfile, remove it from the agent’s own config or centralise everything in groups from the start, which is the better habit.
Groups also let a rule fire for one class and not another without touching the rule. In local_rules.xml, <hostname> and <srcip> conditions can narrow a suppression to a naming pattern, and the alert carries agent.labels if you set labels per group, which rules can match on with <field name="agent.labels.role">.
Tool four: fix the source
The noisiest rule in most deployments is one an application is causing on purpose. A health checker that authenticates on every probe generates a login success every ten seconds. A cron job that runs as root produces a “user executed sudo” trail all night. The right fix is upstream: give the health checker an unauthenticated endpoint, run the cron as a dedicated user with a suppression scoped to that user, ship the application’s own structured log instead of scraping its access log.
Every hour spent on a source fix removes noise permanently and from every downstream tool, not only Wazuh. Every hour spent on a suppression removes it from one tool until the next log format change. Prefer the first when you have any influence over the application.
Reviewing the tuning itself
Tuning creates its own risk: a suppression written for one case widens over time as nobody remembers why it exists. Three habits keep that in check.
Every suppression is a rule with a reason. No bare level changes. The description names the case and the ticket, as in every example above.
Re-measure weekly. The same aggregation query, in a scheduled job, posted to the team channel. When a new rule enters the top ten, tune it that week, not at the next quarterly review.
Test suppressions like detections. Keep a fixture file of log lines that should still alert after tuning, and one of lines that should not, and run both through wazuh-logtest in CI before deploying ruleset changes. A level-0 child that accidentally swallows a real event fails the first fixture, and you find out before production does.
# ci/check-ruleset.sh: fail if any line in should-alert.log scores below 7
while IFS= read -r line; do
level=$(printf '%s\n' "$line" | sudo /var/ossec/bin/wazuh-logtest -q 2>/dev/null | awk -F"'" '/level:/ {print $2}')
[ "${level:-0}" -ge 7 ] || { echo "FAIL (level ${level:-0}): $line"; exit 1; }
done < fixtures/should-alert.log
Where this leaves you
The pipeline from Part 1 now collects only what each host class should send, decodes your own applications from Part 2, responds to brute force safely from Part 3, and routes each remaining alert to the band where someone will act on it. Alert volume is a number you measure weekly instead of a feeling you have about the dashboard.
What is left is the middle band: the level 10 to 12 alerts that need a human decision but not a page, arriving faster than one person can triage by hand. That is where an agent belongs, with the same allowlists, caps, and audit trails Part 3 built into a shell script. The next series on this site builds it.
- Alerts configuration reference log_alert_level and email_alert_level
- Custom rules: overwriting stock rules The overwrite attribute and the level-0 child pattern
- Centralized configuration (agent groups) agent.conf, group assignment, and how group config merges with local config
- Integration with external APIs The integration block for Slack, PagerDuty, and custom webhooks with level thresholds
- OpenSearch terms aggregation The aggregation used for the per-rule volume query
- Wazuh rules classification (levels) The intended meaning of each level band in the stock ruleset
Footnotes
-
A rough benchmark from SOC practice: an analyst can properly investigate somewhere between 20 and 40 alerts in a shift. If your level 10 and up volume per analyst exceeds that, everything above the line is being skimmed, whatever the dashboard says. ↩