← INDEX Detection engineering with Wazuh · 3 of 4

Wazuh Detection Engineering, Part 3: Active Response That Cannot Lock You Out

Attach automatic actions to Wazuh rules with stateful active response scripts, and build in the allowlists, timeouts, dry runs, and audit logs that keep automation from becoming the incident.

Copy for your AI agent. A condensed version of this entry written as a prompt, so Claude, Cursor, or Copilot can apply it to your own codebase.

View prompt · Safe Wazuh active response

Task: Add an automatic block for a brute-force rule without risking lockout.

Manager ossec.conf:

<command>
  <name>block-source</name>
  <executable>block-source.py</executable>
  <timeout_allowed>yes</timeout_allowed>
</command>
<active-response>
  <command>block-source</command>
  <location>local</location>
  <rules_id>100110</rules_id>
  <timeout>900</timeout>
</active-response>

Script contract (/var/ossec/active-response/bin/block-source.py, root:wazuh 750):

  • Read one JSON object from stdin: {"version":1,"origin":{...},"command":"add"|"delete","parameters":{"alert":{...},"extra_args":[]}}
  • On add: reply on stdout with {"version":1,"origin":{"name":"block-source","module":"active-response"},"command":"check_keys","parameters":{"keys":[<ip>]}}, read one line back: continue → do it, abort → exit 0.
  • On delete: undo the action for the same key.
  • Log every decision to /var/ossec/logs/active-responses.log.

Safety checklist:

  • Allowlist file of CIDRs that are never blocked (office, VPN, load balancers, monitoring, RFC1918 by default).
  • DRY_RUN=1 environment or file flag for the first week.
  • Cap on concurrent blocks (e.g. 50); refuse and alert above the cap.
  • Always set <timeout>; never block forever from automation.
  • Test with wazuh-logtest sequences, then with an agent in a lab, before production.

The brute-force rule from Part 2 fires at level 10 when one address fails to log in eight times in two minutes. Nobody wants to read that alert at three in the morning and then type a firewall command. This is what active response is for: analysisd matches the rule, and a script runs. The address is blocked before a human has finished waking up.

The same mechanism, configured carelessly, is how a security team locks itself out of production, blocks the load balancer that fronts every customer, or drops the monitoring probe and then wonders why the dashboards went dark. Every one of those has happened to a real team. This part builds the block, and then spends most of its length on making sure it cannot do any of that.

How active response actually works

Active response is not a separate product. It is three pieces.

A <command> in the manager’s ossec.conf names a script and says whether it accepts a timeout.

An <active-response> block binds a command to a trigger: which rule ids or levels fire it, and where it runs.

A script in /var/ossec/active-response/bin/ on whichever host location points at. The manager ships a few: firewall-drop adds an iptables rule, host-deny appends to /etc/hosts.deny, disable-account locks a user, restart-wazuh restarts the agent. On Windows there is netsh for the firewall.

When a bound rule matches, wazuh-execd on the target host runs the script with the alert as JSON on standard input. If the command allows a timeout and the binding sets one, the script is called again with "command":"delete" when the timeout expires. That second call is what turns a permanent block into a temporary one.

location is the decision that matters most:

locationRuns onUse when
localThe agent that generated the alertBlocking at the host that saw the attack
serverThe managerActions that need manager credentials or API access, and low-risk first deployments
defined-agentOne specific agent (<agent_id>)A dedicated gateway or firewall host
allEvery agentAlmost never. A bad allowlist here blocks the address everywhere at once.

The built-in block

The quickest working version uses the stock script. On the manager:

<!-- ossec.conf on the manager -->
<ossec_config>
  <command>
    <name>firewall-drop</name>
    <executable>firewall-drop</executable>
    <timeout_allowed>yes</timeout_allowed>
  </command>

  <active-response>
    <command>firewall-drop</command>
    <location>local</location>
    <rules_id>100110</rules_id>
    <timeout>900</timeout>
  </active-response>
</ossec_config>

Restart the manager, replay the eight failed logins from Part 2 against a lab agent, and check on that agent:

sudo iptables -L INPUT -n | grep 203.0.113.9
sudo tail -n 5 /var/ossec/logs/active-responses.log

You will see a DROP rule and a log line naming the script, the command, and the source address. Fifteen minutes later the delete call removes it.

This already works. It also has no idea that 203.0.113.9 might be your office. The stock script blocks whatever srcip the alert carries. A custom script is where the safety goes.

Do not put the safety in the rule

It is tempting to add <srcip negate="yes">10.0.0.0/8</srcip> to the rule so it never fires for internal addresses. Then the rule stops alerting on internal brute force too, and you lose detection to protect the response. Keep detection and response separate: the rule alerts on everything, the script decides what it is allowed to touch.

A custom script with the safety built in

Custom scripts speak the same stdin and stdout protocol as the stock ones. Here is one that blocks a source address with nftables, honours an allowlist, supports dry runs, caps how many blocks can be active, and logs every decision as JSON.

#!/var/ossec/framework/python/bin/python3
"""block-source: Wazuh active response that blocks an alert's source IP with nftables.

Protocol: read one JSON message on stdin. For "add", ask execd whether the key is
already handled (check_keys), then act on "continue". For "delete", undo the block.
"""
import ipaddress
import json
import os
import subprocess
import sys
import time

LOG = "/var/ossec/logs/active-responses.log"
ALLOWLIST = "/var/ossec/etc/ar-allowlist.txt"   # one CIDR per line
STATE = "/var/ossec/var/run/block-source.state"  # currently blocked addresses
MAX_ACTIVE = 50
DRY_RUN = os.path.exists("/var/ossec/etc/ar-dry-run")
NFT_SET = "inet filter blocked"                  # created once, see below

def log(**fields):
    fields.setdefault("ts", time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()))
    fields.setdefault("script", "block-source")
    with open(LOG, "a") as fh:
        fh.write(json.dumps(fields) + "\n")

def allowed_networks():
    nets = [ipaddress.ip_network(n) for n in ("10.0.0.0/8", "172.16.0.0/12", "192.168.0.0/16", "127.0.0.0/8")]
    if os.path.exists(ALLOWLIST):
        for line in open(ALLOWLIST):
            line = line.split("#", 1)[0].strip()
            if line:
                nets.append(ipaddress.ip_network(line, strict=False))
    return nets

def is_allowlisted(ip):
    addr = ipaddress.ip_address(ip)
    return any(addr in net for net in allowed_networks())

def active_blocks():
    return set(open(STATE).read().split()) if os.path.exists(STATE) else set()

def save_blocks(blocks):
    with open(STATE, "w") as fh:
        fh.write("\n".join(sorted(blocks)))

def nft(*args):
    cmd = ["nft", *args]
    if DRY_RUN:
        log(action="dry-run", cmd=" ".join(cmd))
        return
    subprocess.run(cmd, check=True, capture_output=True)

def main():
    msg = json.loads(sys.stdin.readline())
    command = msg.get("command")
    alert = msg.get("parameters", {}).get("alert", {})
    ip = alert.get("data", {}).get("src_ip") or alert.get("data", {}).get("srcip")
    rule = alert.get("rule", {}).get("id")

    if not ip:
        log(action="skip", reason="no source ip", rule=rule)
        return

    if command == "add":
        # Ask execd whether this key is already being handled (dedupes repeat alerts).
        print(json.dumps({
            "version": 1,
            "origin": {"name": "block-source", "module": "active-response"},
            "command": "check_keys",
            "parameters": {"keys": [ip]},
        }), flush=True)
        reply = json.loads(sys.stdin.readline())
        if reply.get("command") == "abort":
            log(action="skip", reason="already handled", ip=ip, rule=rule)
            return

        if is_allowlisted(ip):
            log(action="refuse", reason="allowlisted", ip=ip, rule=rule)
            return

        blocks = active_blocks()
        if len(blocks) >= MAX_ACTIVE:
            log(action="refuse", reason="cap reached", ip=ip, rule=rule, active=len(blocks))
            return

        nft("add", "element", *NFT_SET.split(), f"{{ {ip} }}")
        blocks.add(ip)
        save_blocks(blocks)
        log(action="block", ip=ip, rule=rule, dry_run=DRY_RUN)

    elif command == "delete":
        blocks = active_blocks()
        if ip in blocks:
            nft("delete", "element", *NFT_SET.split(), f"{{ {ip} }}")
            blocks.discard(ip)
            save_blocks(blocks)
        log(action="unblock", ip=ip, rule=rule, dry_run=DRY_RUN)

if __name__ == "__main__":
    try:
        main()
    except Exception as exc:  # never let a crash leave execd hanging
        log(action="error", error=str(exc))
        sys.exit(1)

The nftables set it adds to is created once on the host, outside of Wazuh:

sudo nft add table inet filter
sudo nft add set inet filter blocked '{ type ipv4_addr; flags timeout; }'
sudo nft add rule inet filter input ip saddr @blocked drop

Install the script and wire it up:

sudo install -o root -g wazuh -m 750 block-source.py /var/ossec/active-response/bin/block-source.py
sudo touch /var/ossec/etc/ar-dry-run           # dry run for the first week
printf '203.0.113.0/24  # office\n198.51.100.7    # load balancer\n' | sudo tee /var/ossec/etc/ar-allowlist.txt
<command>
  <name>block-source</name>
  <executable>block-source.py</executable>
  <timeout_allowed>yes</timeout_allowed>
</command>

<active-response>
  <command>block-source</command>
  <location>local</location>
  <rules_id>100110</rules_id>
  <timeout>900</timeout>
</active-response>

The check_keys exchange is the part most custom scripts get wrong. When a command allows timeouts, execd tracks which keys are already under an active response so that repeated alerts for the same address do not stack. Your script must send check_keys and wait for continue or abort before acting, or execd will never learn about the block and will never send the delete.1

Every safety mechanism, and why it is there

The allowlist is a file, not a constant. Operations can add the new VPN range without a code change and a restart of the manager. Private ranges are in the script’s default list because blocking them is never what you want from an internet-facing brute-force rule.

Dry run is a file flag. Run the script for a week with /var/ossec/etc/ar-dry-run present. Every decision is logged as if it were real, and nothing changes. Read the log. If anything with action: "block" is an address you recognise, the allowlist is incomplete. Remove the flag only when a week of dry runs contains no surprises.

The cap stops runaway automation. A misconfigured rule, a log format change, a scanner sweeping from a thousand addresses: any of these can trigger hundreds of responses in minutes. Fifty concurrent blocks is enough for a real attack. Above that, the script refuses and the refusal is logged, which is itself a signal that something is off.

The timeout is not optional. Automation never gets to make permanent decisions. Fifteen minutes stops a brute-force attempt. If the attacker returns, the rule fires again and the block renews. A human can make it permanent later.

Structured logging is the audit trail. Every add, delete, refuse, and skip is one JSON line with the rule id and address. When someone asks “why was the customer in Frankfurt blocked at 09:14”, the answer is grep, not archaeology. Ship active-responses.log to the manager too: add a localfile for it on the agent, and you can alert on action: "refuse" with a rule.

Why nftables and a set

Appending individual iptables rules, the way the stock firewall-drop does, gets slow past a few hundred entries and is easy to leave in a mess after a crash. A single nftables set with one drop rule is O(1) per lookup, and flags timeout means the kernel can expire entries itself as a second safety net if the delete call never arrives.

Choosing what to automate

Not every rule deserves a response. A useful test: would you be comfortable with this action being taken a hundred times in an hour with nobody watching? Blocking an external address for fifteen minutes passes. Disabling a user account does not, because the hundredth one might be the CEO during a password-manager migration. Quarantining a host does not, because the hundredth one might be the database.

For actions that fail the test, the right automation is a ticket or a page, not the action. Wazuh can call any script, so the script can open an incident in your tracker with the alert attached and a one-click approval to run the real action. Part 4’s tuning work makes those pages rare enough to be read; the next series on this site builds the agent that handles them.

Testing before production

Three levels, in order.

  1. Unit test the script by piping a handcrafted message into it with the dry-run flag set:
    echo '{"version":1,"origin":{"name":"test","module":"wazuh-execd"},"command":"add","parameters":{"extra_args":[],"alert":{"rule":{"id":"100110"},"data":{"src_ip":"198.51.100.7"}}}}' \
      | sudo -u wazuh /var/ossec/active-response/bin/block-source.py
    Then send the continue line by hand and check the log shows refuse with reason allowlisted.
  2. Integration test in the lab by replaying the Part 2 failures against a lab agent and confirming the block appears in nft list set inet filter blocked, then disappears after the timeout.
  3. Dry run in production for a week, reading the log daily.

Only then remove the dry-run flag, and only on location=local for the host class that actually faces the internet. Agent groups from Part 1 scope which hosts receive the binding.

// REFERENCES
  1. Active response overview How commands, active-response bindings, and locations fit together
  2. Custom active response scripts The stdin/stdout JSON protocol including check_keys, continue, and abort
  3. Default active response scripts firewall-drop, host-deny, disable-account, and the Windows equivalents
  4. Active response configuration reference Every element of the command and active-response blocks, including level, rules_group, and agent_id
  5. nftables wiki: Sets Named sets with timeouts, used by the script in this part
  6. Wazuh stock firewall-drop script source Read what the default script does before replacing it

Footnotes

  1. The full protocol, including the abort and continue messages and the extra_args field for passing static arguments from the command definition, is in the custom active response documentation linked below. The format changed in 4.2; scripts written for 3.x read positional arguments instead of JSON and will silently do nothing. ↩