Skip to content

How detection works

ShadowClaw asks three different questions about the same process and only reports when the answers agree.

Plane A

Inference heartbeat

Is this process doing AI? Sustained compute, model-sized resident memory, and the local model runtimes themselves.

  • Source: ps CPU-time deltas
  • No privilege required
  • Read more

Plane B

Shadow egress

Where is it talking? Per-process sockets attributed to a provider, or to an inference-shaped endpoint in no catalog.

  • Source: lsof, DNS capture, PTR
  • Root widens socket visibility
  • Read more

Plane C

Agent actions

What did the agent then do? Credential reads, identities, privilege, persistence, exfiltration — gated on agent lineage.

  • Source: Endpoint Security
  • Needs root and Full Disk Access
  • Read more

The first two planes answer is this process doing AI? Plane C says nothing about provider reach and everything about what an agent then does with the machine it is running on.

Why correlation rather than any single signal

Every individual observation here has an innocent explanation.

  • A build server burns sustained CPU.
  • A developer machine holds TLS connections open to a hundred hosts.
  • An engineer runs sudo, reads a credential file, and writes a launch agent — on a normal Tuesday.

A detector that pages on one of those would be turned off within a week. So signals are additive, nothing fires on a single weak observation, and the host-plane weights are deliberately set so that no single tactic reaches high on its own.

Withheld from the public documentation

Exact weights, thresholds, and defaults are kept out of the published site: together they are enough to tune activity to sit below the reporting floor. They are in the repository, beside the code that applies them, for operators who need to reason about a score.

Confidence discounts weight

Attribution is not binary. An endpoint resolved from a captured DNS answer is exact; one inferred from forward-resolving the catalog against an anycast CDN is not. A signal whose supporting evidence has confidence below 0.9 contributes proportionally less than its nominal weight, so a probable match cannot score like a certain one.

The poll loop

flowchart LR
  A["ps<br/>process table"] --> C
  B["lsof<br/>sockets + listeners"] --> C
  D["tcpdump :53<br/>DNS answers"] --> R
  E["eslogger<br/>ESF events"] --> K
  F["config watcher<br/>agent config files"] --> K
  R["AddressResolver<br/>IP → provider"] --> C
  C["CorrelationEngine<br/>Plane A + B per process"] --> O
  K["AgentChainEngine<br/>lineage + Plane C sessions"] --> O
  O["Finding<br/>scored, deduplicated"] --> L["Ledger<br/>written first"]
  L --> X["findings.jsonl · console · OTLP"]

Each poll is independent. Nothing is held in memory that could not be rebuilt from the next two polls, except the state that is the detection: per-process CPU-time history, session windows, and the accumulating attribution index.

Withheld from the public documentation

Exact weights, thresholds, and defaults are kept out of the published site: together they are enough to tune activity to sit below the reporting floor. They are in the repository, beside the code that applies them, for operators who need to reason about a score.

What it deliberately does not do

ShadowClaw does not block, terminate, quarantine, or make exceptions. There is no allowlist, and a test fails the build if one is added. sanctioned_endpoints labels approved traffic rather than hiding it — the label is what makes gateway_bypass possible.

An AI-enabled IDE on the machine is, from here, indistinguishable from shadow AI. That is the correct answer. Deciding it is acceptable is a policy call, made downstream, with the finding in hand.

Next