Data flow¶
Each poll is independent. Nothing is held in memory that could not be rebuilt from the next two polls — except the state that is the detection.
Startup¶
flowchart TB
A["CLI arguments"] --> B["resolve_config_path()"]
B --> C{"named explicitly?"}
C -->|yes, missing| D["exit — never fall back<br/>to built-in defaults"]
C -->|yes, present| E["Settings.from_dict()"]
C -->|no| F["search: repo config,<br/>then /etc, then defaults"]
F --> E
E --> G["preflight checks<br/>advisory only"]
G --> H["banner: config provenance,<br/>coverage, warnings"]
H --> I["Sensor.start()"]
A config named by --config or $SHADOWCLAW_CONFIG must exist. The sensor exits rather than quietly running on built-in defaults, because a sensor running on defaults while the operator believes their policy is loaded looks identical to one that is working. That is the whole failure mode.
The banner always names the file that won:
Sensor.start() then brings up the optional sources — the DNS sniffer, the Endpoint Security subscriptions, the agent-config watcher — each of which may be unavailable, and each of which reports that fact rather than failing silently.
One poll¶
sequenceDiagram
participant S as sensor.run_once()
participant P as ProcessSampler
participant L as lsof
participant R as AddressResolver
participant C as CorrelationEngine
participant A as AgentChainEngine
participant D as Ledger
participant X as Exporters
S->>P: sample process table
P-->>S: ProcessSnapshot (CPU-time deltas, RSS)
S->>L: acquire sockets + listeners
L-->>S: Connection[], Listener[]
S->>R: attribute peer addresses
R-->>S: provider, category, confidence, source
S->>C: snapshot + endpoints
C-->>S: Plane A + B findings per process
S->>A: ESF events + config changes
A-->>S: lineage, tactic sessions, chain scores
S->>D: write finding (first, always)
D-->>S: sequence number, chain hash
S->>X: findings.jsonl, console, OTLP logs + metrics, KTP
Acquisition¶
| Step | Source | Notes |
|---|---|---|
| Process table | ps -Ao |
Every process, every user, unprivileged. CPU time is differenced against the previous poll. |
| Connection initiation | PKTAP (tcpdump -k NP -Q out) |
Streamed TCP SYN and QUIC long-header evidence with process name and pid; catches connections shorter than one poll. macOS, root. |
| Sockets and listeners | lsof -iTCP -iUDP |
Settled state and duration; includes shutdown states, not just ESTABLISHED. UDP covers QUIC/HTTP-3. Own user only unless root. |
| DNS answers | tcpdump port 53 |
Optional, root. Feeds the resolver with exact name-to-address pairs. |
| Kernel events | eslogger |
Optional, root. Two subscriptions, one of them prefiltered. |
| Agent configs | stat + content hash |
Unprivileged. Stands down when Endpoint Security is watching the same files. |
Attribution¶
AddressResolver turns a peer address into a provider, and records how it managed:
- A captured DNS answer —
dns, exact. - The forward-resolved catalog index —
catalog, probabilistic. - Reverse DNS —
ptr, weak.
The index accumulates across refresh rounds rather than being replaced, and re-resolves on demand when unattributed egress appears, because anycast pools rotate and a single snapshot would miss most of them.
Correlation¶
Two engines run against the same poll:
CorrelationEngine- Accumulates Plane A and Plane B signals per process, applies confidence weighting, de-duplicates repeat sightings, caps the sum at 100, and bands the result.
AgentChainEngine- Rebuilds process lineage, attributes each host-plane observation to the agent responsible for it, groups observations into
AgentSessions inside the chain window, and awards the chain bonus for three or more distinct tactics in forward order.
Withheld from the public documentation
Exact weights, thresholds, and defaults are kept out of the published site: together they are enough to tune activity to sit below the reporting floor. They are in the repository, beside the code that applies them, for operators who need to reason about a score.
Output ordering is deliberate¶
flowchart LR
F["Finding"] --> L["ledger.db<br/>+ ledger.jsonl"]
L --> A["findings.jsonl"]
A --> B["console"]
B --> C["OTLP logs"]
C --> D["OTLP metrics"]
D --> E["KTP risk factors<br/>+ envelope receipts"]
The ledger is written first, before any file, console, or network export.
Splunk may be unreachable. The Collector may be down. make clean wipes data/. None of that is allowed to cost you the record, so the local copy is the one thing that always exists. See The local ledger.
What gets exported¶
| Shape | Kind | One record per |
|---|---|---|
shadowclaw.finding.recorded |
log | Scored finding. |
shadowclaw.provider.reached |
log | Distinct provider reached — so secondary providers do not disappear behind the finding's headline one. |
shadowclaw.agent.activity |
log | Newly observed tactic. |
shadowclaw.ktp.risk_factors |
log | Poll. |
shadowclaw.ktp.envelope |
log | Attributed agent action. |
shadowclaw.* gauges and counters |
metric | Poll. |
Every resource carries product and authorship attributes, host identity, and a derived os.type. That last one was once the literal string "darwin" in otlp.py — which made every record from a non-macOS host a false statement about itself.
Full field lists in Event schemas and Metrics.
Redaction happens before anything is written¶
The detector reads command lines, which is where API keys live. Everything it emits passes through shadowclaw/redact.py first: provider key prefixes (sk-, sk-ant-, AIza, hf_, ghp_, AWS), KEY=value shapes, bearer headers, and JWTs.
A monitoring control that mirrors credentials into a log index is worse than no control. See Security notes.