Skip to content

Data flow

Each poll is independent. Nothing is held in memory that could not be rebuilt from the next two polls — except the state that is the detection.

Startup

flowchart TB
  A["CLI arguments"] --> B["resolve_config_path()"]
  B --> C{"named explicitly?"}
  C -->|yes, missing| D["exit — never fall back<br/>to built-in defaults"]
  C -->|yes, present| E["Settings.from_dict()"]
  C -->|no| F["search: repo config,<br/>then /etc, then defaults"]
  F --> E
  E --> G["preflight checks<br/>advisory only"]
  G --> H["banner: config provenance,<br/>coverage, warnings"]
  H --> I["Sensor.start()"]

A config named by --config or $SHADOWCLAW_CONFIG must exist. The sensor exits rather than quietly running on built-in defaults, because a sensor running on defaults while the operator believes their policy is loaded looks identical to one that is working. That is the whole failure mode.

The banner always names the file that won:

  config          : /Users/you/ShadowClaw/config/shadowclaw.json  (repository)

Sensor.start() then brings up the optional sources — the DNS sniffer, the Endpoint Security subscriptions, the agent-config watcher — each of which may be unavailable, and each of which reports that fact rather than failing silently.

One poll

sequenceDiagram
  participant S as sensor.run_once()
  participant P as ProcessSampler
  participant L as lsof
  participant R as AddressResolver
  participant C as CorrelationEngine
  participant A as AgentChainEngine
  participant D as Ledger
  participant X as Exporters

  S->>P: sample process table
  P-->>S: ProcessSnapshot (CPU-time deltas, RSS)
  S->>L: acquire sockets + listeners
  L-->>S: Connection[], Listener[]
  S->>R: attribute peer addresses
  R-->>S: provider, category, confidence, source
  S->>C: snapshot + endpoints
  C-->>S: Plane A + B findings per process
  S->>A: ESF events + config changes
  A-->>S: lineage, tactic sessions, chain scores
  S->>D: write finding (first, always)
  D-->>S: sequence number, chain hash
  S->>X: findings.jsonl, console, OTLP logs + metrics, KTP

Acquisition

Step Source Notes
Process table ps -Ao Every process, every user, unprivileged. CPU time is differenced against the previous poll.
Connection initiation PKTAP (tcpdump -k NP -Q out) Streamed TCP SYN and QUIC long-header evidence with process name and pid; catches connections shorter than one poll. macOS, root.
Sockets and listeners lsof -iTCP -iUDP Settled state and duration; includes shutdown states, not just ESTABLISHED. UDP covers QUIC/HTTP-3. Own user only unless root.
DNS answers tcpdump port 53 Optional, root. Feeds the resolver with exact name-to-address pairs.
Kernel events eslogger Optional, root. Two subscriptions, one of them prefiltered.
Agent configs stat + content hash Unprivileged. Stands down when Endpoint Security is watching the same files.

Attribution

AddressResolver turns a peer address into a provider, and records how it managed:

  1. A captured DNS answer — dns, exact.
  2. The forward-resolved catalog index — catalog, probabilistic.
  3. Reverse DNS — ptr, weak.

The index accumulates across refresh rounds rather than being replaced, and re-resolves on demand when unattributed egress appears, because anycast pools rotate and a single snapshot would miss most of them.

Correlation

Two engines run against the same poll:

CorrelationEngine
Accumulates Plane A and Plane B signals per process, applies confidence weighting, de-duplicates repeat sightings, caps the sum at 100, and bands the result.
AgentChainEngine
Rebuilds process lineage, attributes each host-plane observation to the agent responsible for it, groups observations into AgentSessions inside the chain window, and awards the chain bonus for three or more distinct tactics in forward order.

Withheld from the public documentation

Exact weights, thresholds, and defaults are kept out of the published site: together they are enough to tune activity to sit below the reporting floor. They are in the repository, beside the code that applies them, for operators who need to reason about a score.

Output ordering is deliberate

flowchart LR
  F["Finding"] --> L["ledger.db<br/>+ ledger.jsonl"]
  L --> A["findings.jsonl"]
  A --> B["console"]
  B --> C["OTLP logs"]
  C --> D["OTLP metrics"]
  D --> E["KTP risk factors<br/>+ envelope receipts"]

The ledger is written first, before any file, console, or network export.

Splunk may be unreachable. The Collector may be down. make clean wipes data/. None of that is allowed to cost you the record, so the local copy is the one thing that always exists. See The local ledger.

What gets exported

Shape Kind One record per
shadowclaw.finding.recorded log Scored finding.
shadowclaw.provider.reached log Distinct provider reached — so secondary providers do not disappear behind the finding's headline one.
shadowclaw.agent.activity log Newly observed tactic.
shadowclaw.ktp.risk_factors log Poll.
shadowclaw.ktp.envelope log Attributed agent action.
shadowclaw.* gauges and counters metric Poll.

Every resource carries product and authorship attributes, host identity, and a derived os.type. That last one was once the literal string "darwin" in otlp.py — which made every record from a non-macOS host a false statement about itself.

Full field lists in Event schemas and Metrics.

Redaction happens before anything is written

The detector reads command lines, which is where API keys live. Everything it emits passes through shadowclaw/redact.py first: provider key prefixes (sk-, sk-ant-, AIza, hf_, ghp_, AWS), KEY=value shapes, bearer headers, and JWTs.

A monitoring control that mirrors credentials into a log index is worse than no control. See Security notes.

Next