Client 0.2.0 · 82 tests · 97 ATT&CK techniques

Does your endpoint stack actually notice?

A standardised benchmark for endpoint security posture — and for whether your EDR, AV and SOC really do react when something hostile happens. Run it on a disposable machine, get a report, then compare yourself against everybody else.

What the data says today

No published reports yet — so nothing is claimed yet

The comparison switches on as soon as 3 runs have been shared publicly. Below that, an "average" would be one contributor's posture with a percent sign after it. Be one of the first: run the client, then set your report to public.

What it inspects

  • Installed EDR — SentinelOne, CrowdStrike, Tanium, Elastic Agent, Cortex, Carbon Black and more
  • AV and other protection software, with live status and definition age
  • Microsoft Defender state, including real-time and tamper protection
  • Group Policy, MDM profiles and GRC-relevant policy compliance
  • Active Directory / Entra membership and join type
  • Firewall, disk encryption and patch currency
  • Secure boot, TPM and firmware state
  • Logging: audit configuration, log agents, whether anything is forwarded off the host
  • Certificate trust store, proxy inspection, backup and cloud management
  • Remote-access paths, browsers and installed software
  • Local administrators and account hygiene
  • Platform, hardware details and attached USB devices
  • A set of CIS-aligned hardening controls, mapped onto CIS v8, ISO 27001, NIS2 and SOC 2

What it triggers

  • EICAR and other public, inert AV test signatures
  • Ransomware-style mass file renaming, inside a sandbox it created itself
  • Obfuscated PowerShell, living-off-the-land binaries and process injection
  • Credential-store access, persistence and discovery behaviour
  • Attempted creation of a privileged local user, immediately reverted
  • Download of offensive tooling such as mimikatz — fetched, never executed, deleted
  • Egress to known-bad networks, Tor, outbound SSH and a beaconing pattern
  • Exfiltration over DNS and HTTP, and synthetic canary documents to consumer cloud storage
  • Whether hacking services — scan databases, exploit archives, breach data, anonymisers — are reachable at all, by name and by request
  • Nearest-neighbour scanning of the local subnet

Optionally, real malware samples from MalwareBazaar with your own key — downloaded and unpacked only, to see whether that is noticed. Nothing executes a sample, and no code path exists to.

Every test declares the ATT&CK techniques it exercises, runs in its own child process so a kill is a measurement rather than a crash, and writes down how to undo itself before changing anything. The payload set is designed to be extended over time.

Run this on a disposable machine only

The client intentionally behaves like malware in order to provoke a reaction. Use a dedicated laptop you can wipe, or a virtual machine you can roll back.

Never run it on an employee's production machine, on a server, or on anything whose loss or lockdown would hurt. Expect quarantine, alerts, and possibly the machine being isolated by your own SOC — that is a successful test, not a malfunction.

Read the full safety notes

How it works

  1. Download the client One binary for Windows, macOS or Linux. No installer, no dependencies. Current version 0.2.0.
  2. Run it on a dedicated test machine Results stream to the console as each check and trigger completes. The main thread survives even when your EDR kills an individual test worker — that kill is a result.
  3. Get a JSON report Full posture inventory, the hardening control checks, a per-test detected / not-detected verdict with its ATT&CK mapping, how long the host's own logs took to notice, a posture score and a grade. The same run can also be written as a self-contained HTML report or an ATT&CK Navigator layer.
  4. Read it locally, or upload it The report is always written to disk first, and nothing leaves the machine unless you ask: either a curl to the API with your token, or --upload at the end of the run. Skip it and the client is still useful on its own.
  5. Choose what is shared Public feeds the comparative statistics. Organisation-only or fully private keeps it to yourself — you can change your mind on any report at any time.

Why compare?

A detection rate on its own means very little. Knowing that comparable organisations catch a particular technique far more often than you do is actionable.

The statistics break results down by technique, ATT&CK coverage, protection product, platform, posture score, hardening controls, what the network lets out, how fast anything noticed and policy compliance — so you can see which tooling earns its licence cost, and which techniques slip past almost everyone.

Only reports explicitly marked public are ever aggregated, so no single contributor can be identified from an average.

Any group of fewer than 3 reports is withheld entirely.