Does your detection actually work? Test it against real attacks first Back to Blog
Product

Does your detection actually work? Test it against real attacks first

Tyler Oliver October 6, 2026 ~8 mins

Most detection rules ship without ever seeing the attack they're meant to catch. The Huntbase test lake lets you, and Scout, explore recorded real-world attacks in Explorer, test a rule straight from a watcher, see the exact clause that stopped a match, measure noise on real enterprise activity, and score hunts against labelled red-team ground truth.

TL;DR: Most detection rules go live without ever seeing the attack they're written to catch. Huntbase now gives every detection engineer, and Scout, a test lake: more than 1,250 recorded attacks covering 300+ MITRE ATT&CK techniques across all 14 tactics, plus hundreds of millions of events of real enterprise activity with labelled red-team ground truth. Browse any recording in Explorer. Test a rule against 18 recordings in about 10 seconds, straight from a watcher. See the exact clause that stopped a match, and know how noisy a rule is before it ever pages anyone.

The problem: detection rules are tested in production

Ask most security teams how they know a detection rule works and the honest answer is: it hasn't fired yet, so it's probably fine.

That's not negligence. Testing a rule properly is hard:

The result: silent gaps found during incidents, noisy rules that train analysts to ignore alerts, and hunts that conclude "clean" when they never could have found anything.

In Huntbase, a detection rule runs as a watcher: the watcher fires when it matches, a firing can create a detection, and the detection opens a hunt or starts another workflow. The test lake lets you check a watcher, or a rule that isn't one yet, before any of that happens for real.

By the numbers

A dataset page in the test lake: event counts, events over time, sample events and the Atomic Red Team tests in the recording
A dataset page in the test lake: event counts, events over time, sample events and the Atomic Red Team tests in the recording

What changes across the detection lifecycle

StageWithout a test lakeWith Huntbase
Research: what does this technique look like?Read write-ups, or spend days building a lab and running the attackReal rows from a recording in seconds: open it in Explorer, or ask Scout
Build and testOne recording, if you have one, tested by handEvery matching recording, 18 for LSASS dumping, in about 10 seconds
Diagnose missesGuessworkThe exact clause that blocked each match, the signals the attack left behind, and a drafted rule to catch them
Tune for noiseEnable it and wait for complaintsLinted for broad terms, then run against hundreds of millions of real enterprise events and your own recent telemetry before it's enabled
Review and approve"Looks right to me"A test report attached to every rule: detected, missed, can't run, false positives, verdict
HuntZero rows, and no way to tell a miss from a clean resultQueries checked against a recording where the attack is known to be, and scored against labelled red-team activity
MaintainRules drift untested as logs and tools changeRe-test any watcher in seconds whenever it, or your telemetry, changes

Explore a recording like it's your own data

Every recording in the test lake opens in Explorer, the same place you search your own telemetry. Pick a test dataset and Explorer switches to the dataset clock: time ranges follow the recording, not today's date, so "last hour" means the last hour of the attack. The volume chart, field facets, filters and click-to-zoom all work as they do on live data.

Large recordings stay fast. Explorer shows a sampled slice of a big window and says exactly what it sampled, then lets you narrow in. Filtering the 226-million-event LANL week down to a single host returns in about 4–7 seconds.

Every view is marked as test data. Reference rows are never mixed with your environment and can't be used as evidence in a customer hunt.

Explorer on the LANL cyber1 recording: the dataset clock, a host filter, the volume chart and a sampled view of about 15,600 events, marked as test data
Explorer on the LANL cyber1 recording: the dataset clock, a host filter, the volume chart and a sampled view of about 15,600 events, marked as test data

Test a rule before it goes live

Paste a Sigma rule, pick one from the watcher library, or open one of your watchers and choose Test on attack data. Huntbase runs it against every recording that matches its ATT&CK techniques, plus the baselines. You get a straight answer for each recording: detected, missed, likely false positives, or can't run (the recording lacks the fields the rule needs, so you know the gap is in the data, not the rule).

A real example: a widely used "LSASS dump via comsvcs.dll" rule detected exactly the 3 recordings that use comsvcs.dll, and none of the 15 that dump LSASS another way. In about 10 seconds you know precisely what the rule covers, and what other rules you need for full credential-theft coverage.

Testing a comsvcs.dll Sigma rule against 18 LSASS recordings: 3 detected, 15 missed
Testing a comsvcs.dll Sigma rule against 18 LSASS recordings: 3 detected, 15 missed

Every result has a reason. Testing a pass-the-hash rule by its ATT&CK tags gave 1 detected, 3 missed and 3 can't run. Each miss links to Why it missed, and each can't-run names what the recording lacks, for example "no authentication events (Windows logons, Kerberos, NTLM)".

A rule tested by ATT&CK tag: detected, missed and can't-run results per recording, each with its reason
A rule tested by ATT&CK tag: detected, missed and can't-run results per recording, each with its reason

Turn every miss into a better rule

For each recording a rule misses, Huntbase tells you why, down to the clause. Take a rule that matches Image: powershell.exe on the APT29 emulation. Miss analysis reports that the Image clause matched none of the 40,644 Process Activity events, because Image holds full paths and 24 of them end with \powershell.exe. It suggests Image|endswith: '\powershell.exe', and confirms that this clause alone blocks every candidate: the rule's other terms do occur.

Then it shows what the attack did leave behind, by comparing the attack window with the rest of the recording to surface the distinctive signals, and drafts a rule from them.

Tools such as Nanodump and Outflank Dumpert run as in-memory payloads, so a process-creation rule will never see them. Huntbase points straight to what does: process access to lsass.exe with telltale access rights.

Miss analysis: the blocking Image clause named with a fix, and "What the attack left" listing the distinctive events in the attack window
Miss analysis: the blocking Image clause named with a fix, and "What the attack left" listing the distinctive events in the attack window

Ship rules you already know are quiet

Every test includes a quality check:

You get one verdict, ready, needs work or noisy, with the reasons.

A baseline only counts if the rule could have matched in it. When a baseline lacks the fields the rule reads, the report says can't run or can't tell instead of reporting zero hits as quiet.

A real example: an NTLM network-logon rule on the LANL data detected the attack, then the quality check rated it Noisy: about 10.9 million estimated hits on normal activity, measured from a 3.7% sample. You see that before the rule pages anyone. If you still want it, Create watcher anyway is one click away.

A detection test report: detected 1 of 1, rated Noisy with about 10.9 million estimated baseline hits, and Create watcher anyway
A detection test report: detected 1 of 1, rated Noisy with about 10.9 million estimated baseline hits, and Create watcher anyway

Scout tests its own work

AI-written detection rules are only as good as their verification. Scout, our AI analyst, verifies everything it builds against the test lake before it reaches you, and cites the recordings behind every answer.

Scout querying a PurpleSharp Active Directory recording: 17 Kerberos service-ticket rows, marked as test data, with Open in Explorer
Scout querying a PurpleSharp Active Directory recording: 17 Kerberos service-ticket rows, marked as test data, with Open in Explorer

Hunt against ground truth

A clean result only means something if the hunt could have found the attack. The LANL data includes the red team's own record of what it did: 549 labelled red-team events inside the cyber1 days 8–14 recording, alongside 225,974,262 events of normal enterprise activity.

That turns a hunt into something you can score. Ask Scout to find the intruder and it works from the labelled activity: it names the computer the compromise came from, the hijacked accounts, and the hosts they logged on to, each one a chip that opens in Explorer. Run your own hunt query over the same week and you can see whether it would have found the same hosts.

Scout's answer on the LANL recording: the intruder's source computer and hijacked accounts, drawn from 549 labelled red-team events
Scout's answer on the LANL recording: the intruder's source computer and hijacked accounts, drawn from 549 labelled red-team events

Built on open security research

The test lake is built on the security community's best open datasets: OTRF Security-Datasets (MIT), Splunk Attack Data (Apache-2.0) and Los Alamos National Laboratory's cyber-security datasets (CC0). Every dataset page credits its source and gives the citation. It's reference material, kept entirely separate from your organisation's telemetry.

→ See it on your own rules: Watchers docs · Watcher health docs

#Detection #AI #Platform Update
More Articles