Skip to content

Metrics

Emitted every poll as OTLP gauges and monotonic counters. Enabled by emit_otlp_metrics, disabled by --no-otlp-metrics.

Loki does not accept these

Loki serves /otlp/v1/logs only and 404s the metrics path. Send metrics to a Collector or Prometheus. See Grafana and Loki.

Sensor health

Metric Kind Meaning
shadowclaw.sensor.up gauge Liveness.
shadowclaw.findings.active gauge Findings in the current poll.
shadowclaw.esf.running gauge Emitted on every poll, including when zero.
shadowclaw.esf.enabled gauge Whether the host plane was asked for.
shadowclaw.esf.events.received counter Endpoint Security events parsed.
shadowclaw.esf.events.dropped counter Events lost to the bounded queue.
shadowclaw.esf.queue.depth gauge Current queue occupancy, ceiling 20,000.
shadowclaw.configwatch.files gauge Agent-config files being polled.

esf.running is the one to alert on

It is exported on every poll and reads zero when the source is down. If it were emitted only while healthy, a subscription that died would leave no trace at all — and absence is the hardest thing to alert on.

An empty host-plane dashboard row is not a clean host. This metric is what tells the two apart, and splunk/shadowclaw-detection.spl ships an "Endpoint Security went dark" search for exactly this.

esf.events.dropped and queue.depth

A rising dropped counter means the reader is falling behind. eslogger open is a firehose — tens of thousands of events per second on an idle machine — which is why the credential stream runs behind a grep -F prefilter.

Measure before you deploy:

sudo python3 scripts/esf-volume-spike.py

If the credential stream is not viable on your hardware, set esf_credential_stream: false. Exec-argv credential detection continues to work without it.

Detection

Metric Kind Meaning
shadowclaw.risk.score gauge Per finding.
shadowclaw.egress.connection gauge Attributed egress connections.
shadowclaw.process.cpu.percent gauge Per implicated process, from CPU-time deltas.
shadowclaw.process.memory.rss gauge Per implicated process.
shadowclaw.candidate.cpu.percent gauge CPU for candidate runtimes not yet at threshold.

candidate.cpu.percent is the leading indicator for inference_heartbeat: a process climbing here is one that has not yet exceeded cpu_percent_threshold for consecutive_samples polls.

Host plane

Metric Kind Meaning
shadowclaw.agent.sessions.active gauge Agent sessions inside the chain window.
shadowclaw.agent.tactic.observed counter Tactic observations, by tactic.

The count connector in config/otel-collector.agentic.yaml derives fleet-wide tactic rates from these, so a fleet trend question is answerable from the metrics store rather than a log scan.

Kinetic Trust Protocol

Metric Kind Meaning
ktp.risk_factor.adversarial_pressure gauge Stress term, [0,1].
ktp.risk_factor.evidence_density gauge Stress term, [0,1].
ktp.risk_factor.trust_trend gauge Stress term, [0,1].
ktp.risk_factor.update_resistance gauge Stress term, [0,1].
ktp.aggregation.silent_processes gauge Tracked processes that stopped appearing.

Each factor carries these attributes:

Attribute Meaning
ktp.degraded Read this before reading the value.
ktp.degraded_count, ktp.degraded_inputs Which inputs were substituted.
ktp.feeds_active, ktp.feeds_total Coverage behind the value.
ktp.confidence Confidence in the reading.
ktp.risk_domain node
ktp.spec.version 2.0.0
ktp.source_id Emitting sensor identity.

1.0 means hostile or unobserved

Anything the detector could not observe reports 1.0, never 0. On an unprivileged sensor adversarial_pressure never drops below 0.714.

# Wrong. Fires permanently on any unprivileged host.
ktp_risk_factor_adversarial_pressure > 0.7

# Right.
ktp_risk_factor_adversarial_pressure{ktp_degraded="false"} > 0.7

# Also worth alerting, at lower urgency.
ktp_risk_factor_adversarial_pressure{ktp_degraded="true"}

Full explanation in Risk Factors.

Resource attributes

Every metric carries the same resource as every log record — product, authorship, the attribution.intact official-build provenance indicator, host.name, and a derived os.type. See Event schemas. The indicator is product metadata, not a license-compliance signal.

Host metrics from the Collector

config/otel-collector.yaml also runs the Collector's own hostmetrics receiver in two pipelines:

Pipeline Scraper
metrics/ai-processes Process CPU, RSS, and thread counts.
metrics/host Host-wide network and load.

Those are host-wide — they tell you bytes moved, never which pid moved them. That is precisely why the sensor exists. See Why a sensor as well as the Collector.

Alert on findings rather than on raw host metrics: correlation happens at the endpoint where per-process attribution exists, and redoing it downstream is strictly weaker.

Next