Back to Blog
Does your detection actually work? Test it against real attacks first
Most detection rules ship without ever seeing the attack they're meant to catch. The Huntbase test lake lets you, and Scout, explore recorded real-world attacks in Explorer, test a rule straight from a watcher, see the exact clause that stopped a match, measure noise on real enterprise activity, and score hunts against labelled red-team ground truth.
TL;DR: Most detection rules go live without ever seeing the attack they're written to catch. Huntbase now gives every detection engineer, and Scout, a test lake: more than 1,250 recorded attacks covering 300+ MITRE ATT&CK techniques across all 14 tactics, plus hundreds of millions of events of real enterprise activity with labelled red-team ground truth. Browse any recording in Explorer. Test a rule against 18 recordings in about 10 seconds, straight from a watcher. See the exact clause that stopped a match, and know how noisy a rule is before it ever pages anyone.
The problem: detection rules are tested in production
Ask most security teams how they know a detection rule works and the honest answer is: it hasn't fired yet, so it's probably fine.
That's not negligence. Testing a rule properly is hard:
- You need the attack in your logs. That means building a lab, running the technique, collecting the telemetry and replaying it. That's days of work per technique, and most teams don't have the lab.
- One recording isn't enough. There are a dozen ways to dump LSASS memory. A rule that catches ProcDump says nothing about comsvcs.dll, Nanodump or direct syscalls.
- A rule that fires isn't a good rule. It also has to stay quiet on normal activity. Without realistic background data, you learn how noisy a rule is after it's paged someone at 3 a.m.
- Rules arrive faster than anyone can test them. Community rule sets, vendor content and AI-written rules all land in the queue. The bottleneck isn't writing rules; it's knowing which ones to trust.
- Hunt queries fail quietly. A query with the wrong field name returns zero rows, and zero rows looks exactly like "nothing to see here".
The result: silent gaps found during incidents, noisy rules that train analysts to ignore alerts, and hunts that conclude "clean" when they never could have found anything.
In Huntbase, a detection rule runs as a watcher: the watcher fires when it matches, a firing can create a detection, and the detection opens a hunt or starts another workflow. The test lake lets you check a watcher, or a rule that isn't one yet, before any of that happens for real.
By the numbers
- 1,250+ recorded attacks from OTRF Security-Datasets and Splunk Attack Data, including Atomic Red Team runs and full campaign emulations such as APT29.
- 8.5 million attack events across 94 log types and 9 platforms: Windows, Linux, AWS, Azure, Microsoft 365, Okta, network and more.
- 300+ ATT&CK techniques and sub-techniques across all 14 tactics. LSASS credential dumping alone has 14 recordings, one for each major dumping method.
- Hundreds of millions of events of real enterprise activity from Los Alamos National Laboratory (authentication, process, network flow and DNS), with red-team compromise events labelled as ground truth. One week of it, LANL cyber1 days 8–14, is 225,974,262 events.
- About 10 seconds to test a rule against 18 recordings and the baselines.

What changes across the detection lifecycle
| Stage | Without a test lake | With Huntbase |
|---|---|---|
| Research: what does this technique look like? | Read write-ups, or spend days building a lab and running the attack | Real rows from a recording in seconds: open it in Explorer, or ask Scout |
| Build and test | One recording, if you have one, tested by hand | Every matching recording, 18 for LSASS dumping, in about 10 seconds |
| Diagnose misses | Guesswork | The exact clause that blocked each match, the signals the attack left behind, and a drafted rule to catch them |
| Tune for noise | Enable it and wait for complaints | Linted for broad terms, then run against hundreds of millions of real enterprise events and your own recent telemetry before it's enabled |
| Review and approve | "Looks right to me" | A test report attached to every rule: detected, missed, can't run, false positives, verdict |
| Hunt | Zero rows, and no way to tell a miss from a clean result | Queries checked against a recording where the attack is known to be, and scored against labelled red-team activity |
| Maintain | Rules drift untested as logs and tools change | Re-test any watcher in seconds whenever it, or your telemetry, changes |
Explore a recording like it's your own data
Every recording in the test lake opens in Explorer, the same place you search your own telemetry. Pick a test dataset and Explorer switches to the dataset clock: time ranges follow the recording, not today's date, so "last hour" means the last hour of the attack. The volume chart, field facets, filters and click-to-zoom all work as they do on live data.
Large recordings stay fast. Explorer shows a sampled slice of a big window and says exactly what it sampled, then lets you narrow in. Filtering the 226-million-event LANL week down to a single host returns in about 4–7 seconds.
Every view is marked as test data. Reference rows are never mixed with your environment and can't be used as evidence in a customer hunt.

Test a rule before it goes live
Paste a Sigma rule, pick one from the watcher library, or open one of your watchers and choose Test on attack data. Huntbase runs it against every recording that matches its ATT&CK techniques, plus the baselines. You get a straight answer for each recording: detected, missed, likely false positives, or can't run (the recording lacks the fields the rule needs, so you know the gap is in the data, not the rule).
A real example: a widely used "LSASS dump via comsvcs.dll" rule detected exactly the 3 recordings that use comsvcs.dll, and none of the 15 that dump LSASS another way. In about 10 seconds you know precisely what the rule covers, and what other rules you need for full credential-theft coverage.

Every result has a reason. Testing a pass-the-hash rule by its ATT&CK tags gave 1 detected, 3 missed and 3 can't run. Each miss links to Why it missed, and each can't-run names what the recording lacks, for example "no authentication events (Windows logons, Kerberos, NTLM)".

Turn every miss into a better rule
For each recording a rule misses, Huntbase tells you why, down to the clause. Take a rule that matches Image: powershell.exe on the APT29 emulation. Miss analysis reports that the Image clause matched none of the 40,644 Process Activity events, because Image holds full paths and 24 of them end with \powershell.exe. It suggests Image|endswith: '\powershell.exe', and confirms that this clause alone blocks every candidate: the rule's other terms do occur.
Then it shows what the attack did leave behind, by comparing the attack window with the rest of the recording to surface the distinctive signals, and drafts a rule from them.
Tools such as Nanodump and Outflank Dumpert run as in-memory payloads, so a process-creation rule will never see them. Huntbase points straight to what does: process access to lsass.exe with telltale access rights.

Ship rules you already know are quiet
Every test includes a quality check:
- Rule lint: catches overly broad match terms, unanchored image matches and missing log sources.
- Baselines: runs the rule against real enterprise activity to count false positives. Large baselines are measured on a sample and scaled up, and the report says so.
- Your own data: runs it against your recent telemetry to see what it would have fired on.
You get one verdict, ready, needs work or noisy, with the reasons.
A baseline only counts if the rule could have matched in it. When a baseline lacks the fields the rule reads, the report says can't run or can't tell instead of reporting zero hits as quiet.
A real example: an NTLM network-logon rule on the LANL data detected the attack, then the quality check rated it Noisy: about 10.9 million estimated hits on normal activity, measured from a 3.7% sample. You see that before the rule pages anyone. If you still want it, Create watcher anyway is one click away.

Scout tests its own work
AI-written detection rules are only as good as their verification. Scout, our AI analyst, verifies everything it builds against the test lake before it reaches you, and cites the recordings behind every answer.
- Show me how this looks. Scout answers with real rows from a recording, always labelled as reference data and never confused with your environment. Asked what Kerberoasting looks like, it found 7 matching datasets and returned 17 rows of 4769 service-ticket events in about 19 seconds, with a link to open them in Explorer.
- Would this rule catch it? Scout runs the rule and answers with dataset chips you can click. An encoded-PowerShell rule against the APT29 emulation hit once on day 1 and 7 times on day 2, so Scout says it would partially catch the campaign and names the stages it misses.
- Checked claims. Before Scout answers, it checks its claims against the rows it cites, and the recordings it relied on appear as chips you can open.
- Write me a detection rule. Scout drafts a rule and tests it across the matching recordings and the baselines. It explains its misses and revises the rule until it passes the quality check. Then it proposes the rule, with the test report attached, for you to approve.
- Check the hunt before it runs. Before Scout runs a hunt query on your data, it proves the query finds the attack in a matching recording, and fixes the fields or filters if it doesn't. Hunt cells show when a query has been checked.

Hunt against ground truth
A clean result only means something if the hunt could have found the attack. The LANL data includes the red team's own record of what it did: 549 labelled red-team events inside the cyber1 days 8–14 recording, alongside 225,974,262 events of normal enterprise activity.
That turns a hunt into something you can score. Ask Scout to find the intruder and it works from the labelled activity: it names the computer the compromise came from, the hijacked accounts, and the hosts they logged on to, each one a chip that opens in Explorer. Run your own hunt query over the same week and you can see whether it would have found the same hosts.

Built on open security research
The test lake is built on the security community's best open datasets: OTRF Security-Datasets (MIT), Splunk Attack Data (Apache-2.0) and Los Alamos National Laboratory's cyber-security datasets (CC0). Every dataset page credits its source and gives the citation. It's reference material, kept entirely separate from your organisation's telemetry.
→ See it on your own rules: Watchers docs · Watcher health docs