From Manual Testing to Repeatable Detection Validation with Atomic Red Team

From Manual Testing to Repeatable Detection Validation with Atomic Red Team

in

From Manual Testing to Repeatable Detection Validation with Atomic Red Team

How to Validate PowerShell Detections in Splunk with Sysmon, Atomic Red Team, and MITRE ATT&CK


Introduction

A detection rule can look perfectly reasonable on paper and still fail when it encounters real endpoint telemetry.

In this article I document a small detection-validation exercise I performed using PowerShell, Sysmon, Windows PowerShell logging, Splunk, and Atomic Red Team. The objective was not to build a sophisticated detection framework. Instead, I wanted to answer a practical SOC question:

If I manually generate a suspicious PowerShell behavior and then reproduce the same general behavior through an ATT&CK-mapped Atomic Red Team test, will my existing Splunk detections identify it?

The validation process followed this workflow:

Manual PowerShell test
        ↓
EncodedCommand
        ↓
Splunk detection
        ↓
Atomic Red Team T1027
        ↓
Equivalent behavior
        ↓
Does the existing detection identify it?

This exercise ultimately exposed two different aspects of detection engineering: whether the expected endpoint telemetry is generated, and whether the SIEM detection is actually searching that telemetry correctly.

Establishing a Baseline

Before creating a repeatable testing environment, I performed a simple baseline validation.

I manually invoked notepad.exe from PowerShell and then checked both Windows Event Viewer and Splunk to determine whether the expected events appeared.

An unexpected result occurred during this initial test. Launching Notepad interactively from Windows Explorer generated a Sysmon Event ID 1 process-creation event, while the attempted execution from PowerShell did not produce the expected process-creation result in the same way.

Rather than treating this individual application behavior as the basis for the entire validation methodology, I used it to identify a limitation of command-by-command manual testing.

The initial baseline was therefore useful, but limited.

01-01 notepad test.png

01-02 winevt ps logged notepad.png

01-03 splunk pw nloged notepad.png

Initial manual validation of PowerShell and process-creation telemetry. The baseline demonstrated that endpoint behavior and the resulting telemetry should be validated independently rather than assumed.

The first baseline produced an interesting result:

Telemetry Result
Sysmon Event ID 1 Failed to observe expected result
PowerShell Event ID 4104 Successful

The important lesson was that manual testing is difficult to repeat consistently and does not scale well when validating multiple ATT&CK techniques. I therefore moved to a repeatable testing methodology using Atomic Red Team.

For each test, I recorded three things:

  • Whether the expected endpoint telemetry was generated.
  • Whether that telemetry reached Splunk.
  • Whether the corresponding detection identified the activity.

This provided a simple validation chain:

Attack behavior
      ↓
Endpoint telemetry
      ↓
SIEM ingestion
      ↓
Detection logic
      ↓
Detection result

Installing Atomic Red Team

I installed Atomic Red Team and first verified the installed PowerShell version, since the official Invoke-AtomicRedTeam documentation lists PowerShell 5.0 as the minimum supported version.

03-PSversion.png

PowerShell version verification before installing and executing Atomic Red Team tests.

During installation, I encountered an important practical issue: Windows Defender blocked portions of the Atomic Red Team repository.

This is not particularly surprising for a security-testing framework. Atomic Red Team contains scripts and test artifacts that can resemble offensive tooling and may trigger endpoint security controls.

4- windows defender blocking all suspicious files from atomic red team.png

Windows Defender detecting suspicious contents after downloading Atomic Red Team

The official Invoke-AtomicRedTeam documentation also explains that files within the Atomics repository can trigger antivirus detections and supports working with individual tests rather than necessarily retrieving the entire collection.

Rather than disabling Windows Defender to install a large collection of security-testing payloads, I chose a more controlled approach: retrieve and work with only the Atomic test required for the validation.

This is also closer to how I would approach security testing in a controlled environment: minimize unnecessary changes to security controls and use only the test material required for the objective.

Selecting the Atomic Red Team Test

I selected MITRE ATT&CK T1027 — Obfuscated Files or Information, specifically Atomic Test 2, because it produces PowerShell encoded-command activity similar to the behavior I had previously generated manually.

The test definition was:

Attribute Value
Technique T1027 — Obfuscated Files or Information
Atomic Test Execute base64-encoded PowerShell
Atomic Test Number 2
Atomic Test GUID a50d5a97-2531-499e-a1de-5544c74432c6

The test creates a Base64-encoded representation of:

Write-Host "Hey, Atomic!"

and executes the resulting value using:

powershell.exe -EncodedCommand

Before executing it, I reviewed the Atomic definition to understand exactly what behavior it would generate.

This is an important part of the validation process. Atomic Red Team should not be treated as a black box. A SOC analyst should understand what a test is expected to do before deciding whether the resulting telemetry represents the expected behavior.

4 - t1027 installed and path confirmed.png

Local Atomic Red Team test for T1027.

4 - available  tests for 1027.png

Available Atomic Red Team tests for the T1027 technique.

Reviewing the Selected Test

The selected test was T1027-2 — Execute base64-encoded PowerShell.

The relevant test definition described the creation and execution of Base64-encoded PowerShell code and identified the expected output as:

Hey, Atomic!

5 - T1027 - 2 detail.png

Atomic Red Team T1027-2 definition reviewed before execution.

I then prepared the test using:

Invoke-AtomicTest T1027 `
  -TestNumbers 2 `
  -ShowDetails `
  -PathToAtomicsFolder "C:\AtomicRedTeam\atomics"

Reviewing the test before execution gave me a clear expectation of what I should later find in Windows and Splunk telemetry.

Running the Atomic Test

The reviewed Atomic test was subsequently executed.

6 - dates and test.png

Execution of the reviewed Atomic Red Team test.

At this point, the validation question was straightforward:

Does the existing detection identify the same general behavior that I had previously tested manually?

I evaluated the two existing detections independently.

Detection 1 — Suspicious PowerShell Execution Parameters

The first detection searched Sysmon Event ID 1 for PowerShell processes using potentially suspicious execution parameters:

index=* EventCode=1 Image="*powershell.exe"
(CommandLine="*-EncodedCommand*"
OR CommandLine="*-WindowStyle Hidden*")
| stats count by host User CommandLine ParentImage

6-1 splunk detection.png

Detection 1 identified PowerShell process creation containing a suspicious execution parameter.

The search returned two records with the same command line but different user fields.

Rather than immediately assuming that two records represented two separate executions, I inspected the underlying Sysmon events:

index=* EventCode=1 Image="*powershell.exe"
CommandLine="*-EncodedCommand*"
| table _time host User Image CommandLine ParentImage ParentCommandLine ProcessId ProcessGuid
| sort _time

7-1-2 check results in splunk - now 1 entry.png

Reviewing the underlying Sysmon fields to determine whether the apparent duplicate represented separate process executions.

This additional check was important because a SOC analyst should avoid interpreting duplicate-looking SIEM results as multiple executions without validating the underlying telemetry.

The Atomic test generated a Sysmon Event ID 1 containing the -EncodedCommand parameter, and the existing detection successfully identified the activity.

This provided the first positive validation result:

The manually developed Detection 1 successfully identified the equivalent behavior generated by Atomic Red Team.

6-2 splunk detection detail.png

Sysmon Event ID 1 recorded the PowerShell process creation and exposed the -EncodedCommand parameter in the process command line.

I also opened the underlying Windows event to verify that the detection result corresponded to the Atomic execution rather than an unrelated PowerShell process.

6-3 winevent detection.png

Windows Event Viewer confirmation of the Sysmon process-creation event.

Detection 1 therefore passed the validation.

Detection 2 — Suspicious PowerShell Script Blocks

The second detection used PowerShell Script Block Logging, Event ID 4104.

The original detection searched for several PowerShell commands commonly associated with execution, downloading, or .NET-based network activity:

index=* sourcetype="WinEventLog:Microsoft-Windows-PowerShell/Operational"
EventCode=4104
(
    ScriptBlock="*Invoke-WebRequest*"
    OR ScriptBlock="*Invoke-Expression*"
    OR ScriptBlock="*DownloadString*"
    OR ScriptBlock="*FromBase64String*"
    OR ScriptBlock="*Net.WebClient*"
)
| stats count by host User ScriptBlock

7-1 - 0 results for detection 2.png

The initial Detection 2 query returned no results for the Atomic Red Team execution.

Unlike Detection 1, this detection failed.

However, a failed detection does not immediately tell us why it failed.

I therefore investigated the generated 4104 events before changing the detection.

Investigating the 4104 Telemetry

The Atomic test had generated PowerShell Event ID 4104, so the next step was to inspect the event directly.

The relevant Script Block content was:

$OriginalCommand = 'Write-Host "Hey, Atomic!"'
$Bytes = [System.Text.Encoding]::Unicode.GetBytes($OriginalCommand)
$EncodedCommand =[Convert]::ToBase64String($Bytes)
$EncodedCommand
powershell.exe -EncodedCommand $EncodedCommand

The original detection searched for:

Invoke-WebRequest
Invoke-Expression
DownloadString
FromBase64String
Net.WebClient

None of those indicators appeared in the Atomic test. Instead, the test generated:

ToBase64String
-EncodedCommand

This immediately explained one part of the detection failure.

The 4104 telemetry was generated correctly, but the original detection criteria were too narrow for the behavior produced by this Atomic test.

However, there was another issue: when I inspected the fields available in Splunk, the ScriptBlock field was empty while the actual Script Block content was contained in the Message field.

I confirmed this by searching directly for the observed indicator:

index=* sourcetype="WinEventLog:Microsoft-Windows-PowerShell/Operational"
EventCode=4104
"-EncodedCommand"
| table _time host User ScriptBlock Message

The result showed the actual Script Block content in Message.

7-2 - 4104 event details.png

PowerShell Event ID 4104 contained the Atomic execution in the Message field, while the ScriptBlock field was not populated in this Splunk environment.

This was an important distinction: the problem was not that PowerShell failed to generate Script Block logs, but that Detection 2 was searching a field that did not contain the relevant content in this particular Splunk environment.

Additional Confirmation Through Event ID 4103

I also reviewed the corresponding PowerShell Module Logging event, Event ID 4103.

The event showed the resulting command invocation:

CommandInvocation(Write-Host): "Write-Host"
ParameterBinding(Write-Host): name="Object"; value="Hey, Atomic!"

It also showed the PowerShell host application executing with the -EncodedCommand parameter.

7-3 - 4103 event details.png

PowerShell Event ID 4103 provided additional evidence of the resulting command execution.

This provided useful corroborating telemetry.

The 4104 event showed the Script Block that constructed and executed the encoded command, while 4103 showed the resulting command invocation. Sysmon Event ID 1 independently showed the PowerShell process creation and command line.

Together, the events provided a consistent view of the same Atomic execution.

For this detection, however, I kept the detection logic focused on 4104 Script Block Logging rather than attempting to build a separate detection around every available telemetry source.

Refining Detection 2

The initial detection therefore required two changes:

  • Search the field that actually contained the Script Block content in this Splunk environment.
  • Add the -EncodedCommand indicator observed during the Atomic test.

The first refined query was:

index=* sourcetype="WinEventLog:Microsoft-Windows-PowerShell/Operational"
EventCode=4104
(
    Message="*Invoke-WebRequest*"
    OR Message="*Invoke-Expression*"
    OR Message="*DownloadString*"
    OR Message="*FromBase64String*"
    OR Message="*Net.WebClient*"
    OR Message="*-EncodedCommand*"
)
| stats count by host User Message

7-4 Detection 2 improved.png

The refined Detection 2 successfully identified the Atomic Red Team execution after correcting the field and expanding the detection criteria.

The detection now successfully identified the activity.

However, the result exposed another practical SOC consideration: the complete 4104 Message field can be very large and difficult to review directly in a detection results table.

The detailed event remains valuable for investigation, but it is not necessarily the best format for initial analyst triage.

Making the Detection More Practical for SOC Triage

Rather than displaying the complete Script Block in the initial search results, I refined the query once more.

The purpose was simple:

Use the detailed event content for detection, but present the analyst with the specific indicator that caused the match.

The final query was:

index=* sourcetype="WinEventLog:Microsoft-Windows-PowerShell/Operational"
EventCode=4104
(
    Message="*Invoke-WebRequest*"
    OR Message="*Invoke-Expression*"
    OR Message="*DownloadString*"
    OR Message="*FromBase64String*"
    OR Message="*Net.WebClient*"
    OR Message="*-EncodedCommand*"
)
| eval Indicator=case(
    like(Message,"%Invoke-WebRequest%"), "Invoke-WebRequest",
    like(Message,"%Invoke-Expression%"), "Invoke-Expression",
    like(Message,"%DownloadString%"), "DownloadString",
    like(Message,"%FromBase64String%"), "FromBase64String",
    like(Message,"%Net.WebClient%"), "Net.WebClient",
    like(Message,"%-EncodedCommand%"), "-EncodedCommand"
)
| table _time host User Indicator

7-5 Detection 3 refined -final.png

Final Detection 2 output. The query searches the detailed 4104 event but presents a concise indicator for initial analyst triage.

The resulting output was considerably easier to work with:

_time                  host        User             Indicator
2026-08-27 12:10:42    SOC-WIN11   NOT_TRANSLATED   Invoke-Expression

The analyst can immediately see:

  • When the activity occurred.
  • Which host generated it.
  • Which user was associated with the event.
  • Which indicator caused the detection.

If further investigation is required, the underlying 4104 event remains available.

This creates a useful balance between detection coverage and analyst usability.

What the Validation Demonstrated

The two detections produced different results against the same Atomic Red Team execution.

Detection Telemetry Result Finding
Detection 1 Sysmon Event ID 1 PASS Existing command-line detection identified -EncodedCommand
Detection 2 PowerShell Event ID 4104 FAIL → PASS Initial query missed the behavior; field targeting and detection criteria were refined

This distinction is important.

A detection test should not simply answer:

“Did my SPL return an event?”

It should help determine where the detection chain succeeded or failed.

In this case:

Atomic behavior
      ↓
Endpoint telemetry generated
      ↓
Telemetry ingested by Splunk
      ↓
Detection 1 → PASS
Detection 2 → FAIL
      ↓
Investigate 4104
      ↓
Identify field mismatch + missing indicator
      ↓
Refine detection
      ↓
Detection 2 → PASS

The failed detection was therefore useful. It exposed a weakness that would otherwise have remained hidden.

Limitations

This detection is indicator-based and should not be treated as proof of malicious activity.

Legitimate PowerShell scripts can use some of the same commands and parameters, including -EncodedCommand and web-request functionality. A detection match should therefore be treated as an investigation lead rather than confirmation of malicious activity.

The effectiveness of the query also depends on the available PowerShell logging and Splunk field extraction.

In this environment, the relevant 4104 Script Block content was available through the Message field rather than the ScriptBlock field. Other environments may use different field mappings or normalized data models.

Finally, the observed user field was displayed as NOT_TRANSLATED. This did not prevent the detection from functioning, but additional identity enrichment would improve the usefulness of the resulting alert in a production SOC.

Lessons Learned

1. Validate the telemetry before changing the detection

A failed detection does not necessarily mean that the endpoint failed to generate telemetry. In this case, Event ID 4104 was present and contained the expected activity. The problem was found by examining the actual event rather than immediately rewriting the detection.

2. Detection logic should reflect observed behavior

The Atomic test revealed -EncodedCommand, which was not included in the original Detection 2 criteria. Repeatable testing therefore helped identify a genuine coverage gap.

3. Field selection matters

The initial query searched the ScriptBlock field, but in this Splunk environment the relevant content was available in Message. A detection can therefore fail even when the required telemetry is present if the query is built around the wrong field.

4. Detection results should support analyst triage

Returning the entire 4104 Script Block made the result difficult to review. The final query retained the detailed event for investigation while presenting a concise Indicator field in the initial results. This is a practical approach for L1/L2 analysts: identify the reason for the alert first, then investigate the underlying event when necessary.

5. Atomic Red Team provides repeatable validation

Manual testing was useful for establishing the initial baseline, but Atomic Red Team provided a more consistent way to reproduce ATT&CK-mapped behavior. This makes it easier to validate whether detections actually cover the techniques they are intended to identify.

Final Takeaway

The most useful result of this exercise was not simply that Detection 2 was modified to detect -EncodedCommand. The more important lesson was the validation process itself.

A detection can fail because:

  • The expected endpoint telemetry was never generated.
  • The telemetry was not ingested.
  • The data was extracted into a different field than expected.
  • The detection does not contain an indicator for the observed behavior.
  • The detection produces results that are difficult for an analyst to interpret.

Atomic Red Team made it possible to separate these possibilities and investigate them systematically.

The final workflow was therefore:

Manual validation
      ↓
Identify detection objective
      ↓
Create repeatable ATT&CK-mapped test
      ↓
Execute Atomic Red Team
      ↓
Verify endpoint telemetry
      ↓
Validate existing detection
      ↓
Investigate detection failures
      ↓
Refine field/indicator selection
      ↓
Improve analyst-facing output
      ↓
Validate again

For a SOC analyst, this is the key takeaway:

A detection is not finished when the SPL query looks correct. It should be validated against observable endpoint behavior, investigated when it fails, and refined until the resulting alert provides useful information for the analyst who has to respond to it.