Kyvern Threat Model¶
Last updated: 2026-10-07 · Status: pre-1.0
Scope¶
Kyvern is accountability infrastructure for autonomous systems. It provides cryptographically signed, append-only decision audit chains and AI advisory fail-safes. This document defines what Kyvern does — and does not — defend against.
Knowing what a system does not defend against is at least as important as knowing what it does. Overclaiming is the failure mode; this document is written to avoid it.
Primary Threats Kyvern Defends Against¶
Threat 1: Insider/operator post-hoc tampering of decision history¶
Scenario. After an incident, an operator, employee, or other insider attempts to alter, delete, or fabricate decision log entries to mislead investigators, regulators, or courts. The motive is plausible deniability or liability avoidance.
Defense. Every decision is Ed25519-signed and SHA-256 chain-linked to its
predecessor. The signing and linking happen at write time inside
sign_decision() in services/decision/audit_chain.py. Any subsequent
mutation — of payload, action, timestamp, or any other field — breaks chain
verification. The break is detectable at kyvern-verify time:
kyvern-verify chain.jsonl --policy config/policies/default.yaml \
--pubkey signing.pub
# → ✗ Chain broken at index 3 (if tampered)
The tamper is detectable; the original state is unrecoverable to the attacker without the signing private key.
Deleting entries from the end of the chain is different: the shorter
chain is still a valid chain, and signatures alone cannot show that anything
is missing. Anchoring (see Recording Guarantees) detects it for anchored entries:
the receipt for an entry that is gone fails verification. Run
kyvern-verify --require-anchors so that a chain without a valid receipt
fails instead of passing.
Result. The attack surface shifts from the log file — low-protection,
on-disk — to key management, which is HSM/KMS territory and a much harder
problem for an insider. An attacker who cannot forge signatures cannot silently
rewrite history. See docs/architecture.md §4 for the
full audit chain design.
Threat 2: AI-induced unsafe escalation¶
Scenario. An LLM advisor — potentially manipulated via prompt injection, adversarial input, or simply faulty reasoning — recommends a dangerous action that the rule engine alone would not produce. The system acts on the recommendation.
Defense. Three independent layers prevent this:
(a) LLM advisory ceiling. The LLM advisor is structurally prohibited from
recommending ENGAGE — the highest-severity action. Its maximum proposal is
HANDOFF. This is enforced in llm_graph.py (_reconcile_action): the
advisor can escalate LOG → ALERT → HANDOFF but nothing beyond, and cannot
lower the rule engine's action. An LLM answer of ENGAGE (or of an unknown
action) is ignored and the rule engine's action stands
(tests/decision/test_llm_graph.py).
(b) Guardrail downgrade-only invariant. Guardrails run after LLM
reconciliation and can only reduce severity, never increase it. The
_SEVERITY ordering in guardrails.py enforces this mathematically: a
guardrail that returns downgrade_to=ENGAGE when the current action is ALERT
is silently discarded. The worst-case outcome of a guardrail bug is an
unnecessary downgrade to LOG — the safest possible state. See
docs/architecture.md §6 for the full downgrade-only
invariant.
(c) Prompt injection quarantine. Track data never reaches the LLM prompt
unfiltered. sanitize.py (sanitize_track_for_llm) applies allowlist
filtering, control-character stripping, and injection-pattern detection before
any track field enters the prompt. This limits the blast radius of adversarial
sensor inputs.
Result. AI is structurally an advisor; the rule engine is the decider. A compromised, manipulated, or simply wrong LLM cannot cause the system to take an action more aggressive than the rule engine would have taken without it.
Recording Guarantees¶
- No unrecorded decision.
run_graph()returns a decision only after it has been appended to the chain and fsynced. If that fails — disk error, lock timeout, corrupt chain tail, missing signing key for an existing chain — it raisesAuditWriteErrorand returns nothing; the caller decides how to fail safe. - One chain per host. Decisions and RuntimeEvents from several processes on the same host append through one inter-process lock, so the chain does not fork. Several hosts writing one chain is not supported.
- Every entry names its key.
key_idis part of the signed payload. A verifier given the wrong key reports "signed by unknown key" instead of a generic failure, and the signing key is never silently re-created next to a non-empty chain. - Anchoring the chain head. Whoever holds the signing key can rewrite
the chain from any point and re-sign it forward, and nothing inside the
chain proves that did not happen.
kyvern-anchorcloses this for anchored entries: it sends the hash of the current chain head to an RFC 3161 Time Stamping Authority (IdenTrust by default) and stores the signed timestamp inchain.anchors.jsonl.kyvern-verifychecks every receipt against the chain: an anchored entry that was rewritten — even re-signed with the right key — no longer matches its receipt and verification fails. Limits: - entries after the latest anchor can still be rewritten or deleted until
the next
kyvern-anchorrun, so run it on a schedule (for example hourly); - without receipts nothing detects entries deleted from the end of the
chain;
kyvern-verify --require-anchorsfails a chain that has no valid receipt; - a receipt proves when an entry existed, not the time the entry claims,
so someone holding the signing key could write a backdated entry and
anchor it later;
kyvern-verify --max-lag DURATIONfails an anchored entry dated more than DURATION before (or after) the first receipt that covers it, which bounds the backdating to about one anchoring interval; - the receipts file is the evidence: copy
chain.anchors.jsonloff the host (log shipping, object storage), otherwise an attacker with host access can delete it; - the TSA receives only a hash, never chain contents.
Rotating the signing key¶
- Stop every process that appends to the chain.
- Note the current key id:
python -c "from pathlib import Path; from cryptography.hazmat.primitives import serialization; from services.decision.audit_chain import key_id; print(key_id(serialization.load_pem_public_key((Path.home()/'.kyvern'/'keys'/'signing.pub').read_bytes())))" - Copy
~/.kyvern/keys/signing.pubto~/.kyvern/keys/signing-<old key id>.puband movesigning.keyto offline storage. - Create the new pair:
from pathlib import Path
from cryptography.hazmat.primitives import serialization
from cryptography.hazmat.primitives.asymmetric import ed25519
keys = Path.home() / ".kyvern" / "keys"
key = ed25519.Ed25519PrivateKey.generate()
(keys / "signing.key").write_bytes(key.private_bytes(
serialization.Encoding.PEM, serialization.PrivateFormat.PKCS8,
serialization.NoEncryption()))
(keys / "signing.pub").write_bytes(key.public_key().public_bytes(
serialization.Encoding.PEM, serialization.PublicFormat.SubjectPublicKeyInfo))
- Restart the writers, then verify with both keys:
kyvern-verify chain.jsonl --policy <policy> --pubkey ~/.kyvern/keys/signing-<old key id>.pub --pubkey ~/.kyvern/keys/signing.pub
Threats Kyvern Explicitly Does NOT Defend Against¶
These are deployment-operator responsibilities. We list them to avoid overclaiming in due diligence, interviews, or compliance assessments.
Signing key compromise¶
If an attacker obtains the signing private key, they can produce arbitrary
"valid" chains that pass kyvern-verify. Kyvern does not perform key
management. Operators must use an HSM, KMS, or equivalent hardware-backed key
store, with the private key never materialized on the host running Kyvern.
Wholesale chain replacement¶
If an attacker replaces the entire chain file with one signed by a key they
control, and the public-key trust anchor itself is compromised or unverified,
kyvern-verify will report the replacement chain as valid. Trust anchor
verification — how a verifier obtains and trusts the public key — is the
operator's responsibility. Kyvern provides the verification mechanism; it
cannot solve the key-distribution problem.
Sensor-level truth¶
Kyvern trusts its inputs. If a sensor lies, or upstream data is poisoned
before reaching FusionService, the resulting decision will be logged
faithfully — including the full provenance of why that decision was made.
The logged record is cryptographically correct; the underlying premises are
not. Sensor integrity, calibration auditing, and input-source authentication
are out of scope.
Runtime intrusion and process isolation¶
Container escapes, privilege escalation, or direct compromise of the host process running Kyvern are not Kyvern's responsibility. Standard host hardening, OS-level security, and container isolation policies apply. If the process itself is compromised, the signing key in memory is compromised too.
Network-level attacks¶
The MCP server (kyvern-mcp) is stdio-only in v1 — there is no network
listener and therefore no network attack surface. This is a deliberate design
choice that limits exposure, not an active network defense mechanism. Future
SSE/HTTP transport (Phase 2) will require TLS, authentication, and
rate-limiting; those are deployment concerns, not Kyvern guarantees. See
docs/integrations/mcp.md for the MCP threat model
note.
Honest One-Line Summary¶
Kyvern defends against insider post-hoc tampering of decision history and AI-induced unsafe escalation. It does not defend against signing key compromise, sensor-level deception, runtime intrusion, or network attacks — those are operator responsibilities.
What This Means for Compliance Claims¶
Kyvern produces tamper-evident decision audit trails suitable for the
Article 12(2) general logging requirements under Regulation (EU) 2024/1689
(risk identification per Art.79(1), post-market monitoring per Art.72,
operation monitoring per Art.26(5)). See
docs/compliance/eu_ai_act.md.
Kyvern is not a substitute for:
- Key management policy (HSM, KMS, key rotation procedures)
- Sensor calibration auditing or input-source authentication
- Host hardening and container isolation
- Network security (TLS, authentication, rate-limiting)
- Third-party conformity assessment under Article 43 of the EU AI Act