From b2f89d417c86f7c6a578b9f672ac650896747420 Mon Sep 17 00:00:00 2001 From: Jeff Larson Date: Sat, 4 Jul 2026 18:11:49 -0700 Subject: [PATCH] docs(adr): truth-up 0014 + define the Retire-Falco corroboration-parity bar (JEF-305) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Un-defer per-objective corroboration in ADR-0014 (it shipped in engine/src/engine/reason/proof/corroborate.rs, wired at reason/proof/mod.rs) and record the four F1 decisions the Retire-Falco epic builds on: 1. per-objective corroboration is landed, not deferred (Status block corrected); 2. the parity bar is measured decision-path corroboration coverage (F6/F7 gate), NOT Falco YAML rule-replication (that would re-create the removed coupling); 3. "retire Falco" = retire the Falco adapter + external deploy, KEEP the tool-agnostic Behavior::Alert + /behavior port (ADR-0003 — any sensor); 4. "alarming-now -> blanket corroboration" is engine-side classifier policy (JEF-113 pattern; wire type stays pure data, classification in observe::exec_class). Light amendment to ADR-0009 (its "live Falco signal" is now the tool-agnostic, per-objective corroboration of 0014), README index statuses updated, and the stale "only an alert behavior corroborates" comment in charts/protector/values.yaml fixed. Docs + one chart comment only — no code/behavior change. Co-Authored-By: Claude Opus 4.8 (1M context) Claude-Session: https://claude.ai/code/session_01VtjoJttCvBY4dzCoE4f9vP --- charts/protector/values.yaml | 5 +- docs/adr/0009-asymmetric-action-bar.md | 9 +++ docs/adr/0014-behavioral-telemetry-ebpf.md | 84 ++++++++++++++++++---- docs/adr/README.md | 4 +- 4 files changed, 87 insertions(+), 15 deletions(-) diff --git a/charts/protector/values.yaml b/charts/protector/values.yaml index 9511f27..5ff98af 100644 --- a/charts/protector/values.yaml +++ b/charts/protector/values.yaml @@ -354,7 +354,10 @@ engine: # # OFF by default: it needs the protector-agent image and the eBPF probes load-tested on # your kernel (see agent/README.md). When enabled it runs in SHADOW — the signals enrich -# the model + the engine's output state; only an `alert` behavior corroborates the action bar. +# the model + the engine's output state. Corroboration is per-objective (ADR-0014): an +# alerting signal corroborates any chain, and the agent's own behaviors each corroborate +# the objective class whose ATT&CK tactic they evidence (internet egress → exfil, secret +# read → credential-access, vuln-library load / notable exec → foothold). agent: enabled: false image: diff --git a/docs/adr/0009-asymmetric-action-bar.md b/docs/adr/0009-asymmetric-action-bar.md index b71df31..a90dd40 100644 --- a/docs/adr/0009-asymmetric-action-bar.md +++ b/docs/adr/0009-asymmetric-action-bar.md @@ -3,6 +3,15 @@ - Status: Accepted - Date: 2026-06-12 +> **Amendment (JEF-305, 2026-07-04):** this ADR describes the live-corroboration signal +> as "a live Falco signal" because Falco was the only sensor when it was written. That +> `corroborated-now` predicate is now **tool-agnostic and per-objective** +> ([ADR-0014](0014-behavioral-telemetry-ebpf.md)): any sensor (Falco, Tetragon, the +> first-party eBPF agent) feeds a normalized behavioral port, an *alerting* signal +> corroborates any chain, and the agent's own behaviors corroborate the objective class +> whose ATT&CK tactic they evidence. Read every "Falco" below as "a live corroboration +> signal from whichever sensor is present." The asymmetry decided here is unchanged. + ## Context [ADR-0001](0001-async-mitigation-engine.md) set the bar for automated response as diff --git a/docs/adr/0014-behavioral-telemetry-ebpf.md b/docs/adr/0014-behavioral-telemetry-ebpf.md index 8acda97..4425c54 100644 --- a/docs/adr/0014-behavioral-telemetry-ebpf.md +++ b/docs/adr/0014-behavioral-telemetry-ebpf.md @@ -85,16 +85,24 @@ telemetry without requiring any third-party sensor. *promotes* toward action behind the existing reversible, self-reverting cut — never a new kind of action. - *Status (phase 1, shadow):* per-objective corroboration — the - `corroborates(behavior, objective)` relation described above (egress→exfil, - lib-load→foothold) — is **deferred**. What ships today is the flat, - `is_alert()`-gated corroboration: only an *alerting* signal (a Falco critical) - corroborates, and it does so for any chain on the workload. Mundane behaviors - (connections, secret reads, library loads) flow to the model as evidence but do **not** - yet promote the action bar — widening the flat predicate to admit them would corroborate - everything (every workload connects). The per-objective relation belongs at - objective/action-bar matching, not in that predicate; see the `entry_corroborated` NB in - `engine/src/engine/proof.rs`. It lands only after the shadow bake (rollout step 3 below). + *Status (landed, shadow):* per-objective corroboration — the + `corroborates(behavior, objective)` relation described above — **has shipped**. It lives + in [`engine/src/engine/reason/proof/corroborate.rs`](../../engine/src/engine/reason/proof/corroborate.rs) + (`corroborates` / `corroborated_for`) and is wired into chain proof at + [`reason/proof/mod.rs`](../../engine/src/engine/reason/proof/mod.rs). Each behavior + corroborates only the objective class whose ATT&CK *tactic* it evidences: internet + egress → EXFILTRATION (T1041), secret read → CREDENTIAL_ACCESS (T1552), vuln-library + load → the INITIAL_ACCESS / EXPLOIT_PUBLIC_FACING foothold (T1190, matched against the + entry's foothold tactic per JEF-77 as well as the objective's). An *alerting* signal + (`Behavior::Alert`) still corroborates **any** chain — "an attack is happening now" + regardless of which objective — and a *notable* exec (interactive shell / package + manager, JEF-55/JEF-117) corroborates broadly the same way, as the agent-side + equivalent of Falco's shell/pkg-mgr criticals. A *bare* `ProcessExec` and mundane + in-cluster connections remain model-evidence only, so the predicate never becomes the + "everything corroborates everything" blanket. This is entirely **shadow-gated**: the + arms only set `corroborated`; actuation stays behind `mode: enforce` / `enforceScope` + (ADR-0021). The earlier "flat, `is_alert()`-gated, deferred until a shadow bake" + framing is obsolete — it is superseded by the per-objective relation now in the tree. Behavioral signals are **never graph structure**: they don't mint edges or nodes, so the proof layer's "reach is proven, not guessed" invariant is untouched. They are TTL'd @@ -156,5 +164,57 @@ Shadow-first, mirroring the engine's posture ([ADR-0001](0001-async-mitigation-e 2. **The agent.** Build `protector-agent` (aya) — network connections first, then secret reads and library loads — emitting normalized signals. Graceful degradation per hook. 3. **Deploy.** Add the DaemonSet to the chart with the scoped capabilities, **in shadow** - — signals only enrich the engine's output state and the model's prompt. Only after a bake does any - behavioral signal feed the action-bar corroboration that can promote a cut. + — signals enrich the engine's output state and the model's prompt, and feed the + per-objective action-bar corroboration relation (now landed; see *Status* above). + Every corroboration path is shadow-gated: it can only ever *promote a cut* under + `mode: enforce` within `enforceScope` (ADR-0021), never by the mere presence of a + signal. + +## Addendum — retiring Falco: the corroboration-parity bar (JEF-305, 2026-07-04) + +Falco 0.44.1 crash-loops on the cluster's `7.0.0-1014-raspi` arm64 kernel (a libsinsp +ABI mismatch against the syscall tracepoints it parses), leaving live corroboration down +on half the nodes; the first-party agent runs healthy on the same kernel because it +attaches to LSM fentry + kprobes, not those tracepoints. The "Retire Falco" epic +therefore moves live corroboration fully to the agent. This addendum records the four +decisions the rest of that epic builds on. **It is a decision/documentation change only — +no behavior changes with it.** + +1. **Per-objective corroboration is landed, not deferred.** The original *Consumers → + action bar* status note called the `corroborates(behavior, objective)` relation a + deferred phase-1 item to fill after a shadow bake. That is stale: the relation shipped + and lives in + [`corroborate.rs`](../../engine/src/engine/reason/proof/corroborate.rs), wired at + [`reason/proof/mod.rs`](../../engine/src/engine/reason/proof/mod.rs). The *Status* + block above has been corrected to match the tree. + +2. **The parity bar is measured decision-path coverage, not Falco rule-replication.** + Retiring Falco does **not** mean porting its YAML ruleset — doing so would re-create + exactly the upstream coupling ADR-0003 removes. The agent must cover what protector's + *decision path* actually consumes: the corroboration signal that reaches + `corroborates` (today `Behavior::Alert` / the `is_alert`-style broad signal, plus the + notable-exec equivalent of Falco's shell/pkg-mgr criticals). Parity is that + decision-path corroboration coverage, **measured** — the later F6/F7 gate in this epic + proves the agent reproduces the corroboration the deployed Falco was contributing + before the adapter and external workload retire. Rule-for-rule fidelity is a non-goal. + +3. **"Retire Falco" = retire the Falco *adapter* + external deploy — keep the + tool-agnostic port.** What retires is the Falco-specific ingest adapter (the legacy + `/` alert shape) and the external Falco / falcosidekick workload. What **stays** is the + normalized `Behavior::Alert` variant and the `/behavior` ingest port: they are + tool-agnostic (ADR-0003), so any sensor — Tetragon, a future collector, the + first-party agent — can still POST an alerting corroboration signal. Retiring one + sensor's adapter must never remove the contract other sensors depend on. + +4. **"Alarming-now → blanket corroboration" is an engine-side classifier policy.** The + decision that an *alerting* signal corroborates any chain (and that a notable exec does + the same) is **classification policy that lives engine-side**, following the JEF-113 + pattern: the wire behavior type stays pure data, and the "is this alarming now?" + judgement is made in the engine (as `observe::exec_class` already does for + shell/pkg-mgr execs). A new sensor does not encode the blanket-corroboration policy on + the wire; it emits data, and the engine classifies. This keeps the port honest and the + policy in one place we can audit. + +None of these four touch the honesty, zero-egress, or shadow-by-default framing: the +agent stays observe-only, the graph and evidence stay in-cluster, and corroboration only +ever promotes a cut behind the existing reversible, self-reverting, `enforce`-gated bar. diff --git a/docs/adr/README.md b/docs/adr/README.md index 8da5a29..bafefcf 100644 --- a/docs/adr/README.md +++ b/docs/adr/README.md @@ -16,12 +16,12 @@ Copy [`0000-template.md`](0000-template.md) to start one. | [0005](0005-attack-objectives.md) | Objectives are ATT&CK outcomes, not just secrets | Accepted | | [0006](0006-build-vs-adopt.md) | Build the substrate; treat KubeHound/IceKube as catalogue and optional provider | Accepted | | [0007](0007-live-cuts-via-adminnetworkpolicy.md) | Live network cuts are additive AdminNetworkPolicy Deny rules | Accepted | -| [0009](0009-asymmetric-action-bar.md) | Asymmetric action bar: live evidence acts, latent exposure proposes | Accepted (amended by 0011, 0013, 0016, 0017, 0022) | +| [0009](0009-asymmetric-action-bar.md) | Asymmetric action bar: live evidence acts, latent exposure proposes | Accepted (amended by 0011, 0013, 0016, 0017, 0022; corroboration made tool-agnostic + per-objective by 0014/JEF-305) | | [0010](0010-flannel-actuator-workload-isolation.md) | Flannel actuator: quarantine the source with a default-deny NetworkPolicy | Accepted (amended by 0022) | | [0011](0011-positive-judgement.md) | The model corroborates positively; operator access is out of scope, defended in depth | Superseded in part by 0013 | | [0012](0012-exposure-observed-or-declared.md) | Exposure is observed where possible, declared (annotation) where it can't be — tunnels | Accepted | | [0013](0013-proof-winnows-model-decides.md) | Proof winnows the search space; the model makes the exploitability call (positive gate + breach-relevance) | Accepted (amended by 0016) | -| [0014](0014-behavioral-telemetry-ebpf.md) | First-party behavioral telemetry via eBPF, behind a tool-agnostic port (potential vs actual) | Accepted | +| [0014](0014-behavioral-telemetry-ebpf.md) | First-party behavioral telemetry via eBPF, behind a tool-agnostic port (potential vs actual) | Accepted (amended by JEF-305: per-objective corroboration landed; the Retire-Falco parity bar = measured decision-path coverage, retire the adapter not the port) | | [0015](0015-advisory-evidence-egress.md) | Advisory evidence is mounted-snapshot-only (zero egress); structurally extracted + capped for injection safety | Accepted (advisory feed retired per JEF-242; Rekor egress carve-out amended by 0020) | | [0016](0016-severity-vs-urgency.md) | The breach model: prove chains, enrich them, the model decides and isolates until clear | Accepted (amended by 0017) | | [0017](0017-isolation-persists-on-the-breach-condition.md) | Isolation persists on the breach condition: chain ∧ enrichment fingerprint (revert keys on `entry_fingerprint`) | Accepted |