This page is a copy of spec/IEB-SPEC.md in the Lapilli repository, made when the site was last deployed. The repository is the normative source. A verifier must not fetch this page: schema_version is compared as a string, verification is offline, and a verifier that asked a website what the rules are would have given that property away.

Incident Evidence Bundle (IEB) — Reference Bundle Layout

Status: lapilli.dev/ieb/v1, frozen from the first release (v0.1.0) under ../docs/COMPATIBILITY.md. The verification contract below is normative. This is a reference bundle layout, not (yet) a “standard” — that word is earned only when an independent producer or consumer adopts it (see ../docs/design-review-round2.md, Path C).

The IEB is a portable representation of a single Kubernetes incident. Lapilli is its reference producer; the layout is intended to be readable by other tools (e.g. HolmesGPT, k8sgpt, homegrown scripts).

Container

Layout (ieb/v1)

manifest.json        # schema version, bound incident identity, hash tree, coverage, image digest
timeline.json        # normalized time-sorted events across sources        [collector: events]
events.json          # raw Kubernetes events for the target               [collector: events]
resources/           # point-in-time JSON of the pod + owner chain         [collector: resources]
  pod.json           #   Pod → ReplicaSet → Deployment
  replicaset.json
  deployment.json
  statefulset.json   #   (optional) when a StatefulSet owns the pod directly
  daemonset.json     #   (optional) when a DaemonSet owns the pod directly
logs/                # log tails, bounded by lines AND bytes               [collector: logs]
  index.json         #   which instance each file came from + gaps (see below)
  <container>-current.log
  <container>-previous.log   # last-terminated instance (the timing-sensitive win)
changes.json         # change indicators (generation, managedFields incl. subresource,
                     #   revision)                                        [collector: changes]
diffs/               # before/after pod-template diffs                     [collector: changes]
  index.json         #   expected objects, one entry per revision pair, status, timing, actor
  <ns>/<Kind>/<name>/<n>.json   # field changes of one pair (pending.json: unrolled edits)
metrics/             # PromQL range snapshots (optional)                  [collector: metrics]
  index.json         #   queried range, step, rendered queries, per-query status
  <name>.json        #   raw Prometheus query_range response, verbatim
redaction.json       # redaction policy version, mode, per-file counts, dropped fields
signature/           # (optional) detached signature over manifest.json — present only if signing enabled

diffs/

Before/after diffs of the pod template, read at capture time from the revision history Kubernetes already keeps: ReplicaSets of a Deployment, ControllerRevisions of a StatefulSet or DaemonSet. No watch, no stored state. The full rules and their rationale are in ../docs/design-change-diff.md.

diffs/index.json:

{ "normalization": "v1",
  "expected": [ { "namespace": "lapilli-demo", "kind": "Deployment", "name": "checkout" } ],
  "entries": [ {
    "namespace": "lapilli-demo", "kind": "Deployment", "name": "checkout",
    "status": "ok", "source": "replicaset-history",
    "before": { "revision": "1", "object": "ReplicaSet/checkout-5c8f4588f5" },
    "after":  { "revision": "2", "object": "ReplicaSet/checkout-7f8b6b66f8" },
    "changed_at": "…", "changed_at_source": "creationTimestamp",
    "seconds_relative_to_firing": -2, "after_firing": false, "in_range": true,
    "actor": "demo-deployer", "actor_kind": "fieldManager (client-asserted)",
    "pod_revision_is_current": true, "kind_of_change": "spec", "warnings": [],
    "summary": ["containers[name=app].env[name=CACHE_WARMUP].value: lazy → eager"],
    "file": "diffs/lapilli-demo/Deployment/checkout/0.json" } ] }

A change file is a list of {op, display, path_before?, path_after?, before?, after?, changed}. display is the normative identity: list elements are addressed by merge key (containers[name=app], ports[containerPort=8080,protocol=TCP], with \ ] = , escaped), so reordering is not a change. path_before/path_after are RFC 6901 pointers into each side, absent where the element does not exist. It is not an RFC 6902 patch. Values are redacted with the same policy as resources/; changed stays true when both sides are "<redacted>".

Redaction and redaction.json

Captured objects pass through redaction before they are written, so every file only ever holds redacted values. Policy v1 is best-effort (see docs/design-change-diff.md); strict mode exists for deployments that need a guarantee.

Candidate (pod spec / pod template / object metadata) Name rule Value rule
env[].value (name = env[].name) ✅ ✅ per token
command[], args[], lifecycle and probe exec.command[] ✅ on --name=v, -Dname=v, NAME=v, -name v, -u user:pass ✅ per token
lifecycle/probe httpGet.httpHeaders[].value (name = header) ✅ ✅
metadata.annotations, spec.template.metadata.annotations (except *.kubernetes.io/*, *.k8s.io/*) ✅ ✅; JSON values per inner key
event message (events.json, timeline.json) ✅ on name=v tokens ✅ per token
kubectl.kubernetes.io/last-applied-configuration dropped —
everything else: not redacted — see the list below not redacted —

“Everything else” is named, not hand-waved, because an independent implementer reads this table as the contract and a consumer has to classify what it receives. No mode, strict included, touches any of these:

This is a statement about redaction, not about verification: none of it affects a verdict. What it means for an operator who has to classify a bundle before installing the reference producer is in ../docs/data-handling.md.

redaction.json:

{ "policy_version": "v1", "mode": "default", "plaintext_names": [],
  "redacted_values": { "resources/pod.json": 2, "resources/replicaset.json": 2 },
  "dropped_fields": [],
  "not_redacted": ["logs/", "metrics/"],
  "not_redacted_fields": ["metadata.labels", "metadata.managedFields", "spec.nodeName",
                          "spec.serviceAccountName", "spec.containers[].image", "…"] }

The two lists answer different questions and a reader must not merge them. not_redacted holds whole trees the policy never visits — everything under those prefixes is as the cluster gave it. not_redacted_fields holds fields that survive inside the files the policy does visit, so a consumer cannot conclude from “this file was redacted” that any particular value in it was. Both are advisory: §7 fixes only mode, a verifier does not check either list, and a bundle that omits them is still valid. A consumer reporting them MUST report them as what the capture recorded rather than as a property it verified.

lapilli verify prints a warning for a bundle captured with mode: off. A bundle without redaction.json is FAILED — exit 1, problem code manifest, plus integrity when the hash tree still lists the file — not merely noted: §7 makes the file required, and the frozen fixture fail-no-redaction.ieb has pinned that verdict, with both codes, since v0.1.0.

An earlier version of this sentence said a missing redaction.json was noted. It contradicted §7 of this same document, and the fixture had been enforcing §7 all along.

metrics/index.json

{ "prometheus_url": "http://prometheus.monitoring:9090",
  "start": "2026-09-19T03:46:03Z", "end": "2026-09-19T03:51:03Z",
  "end_capped_at_capture": true, "step_seconds": 5,
  "queries": [
    { "name": "memory_working_set_bytes", "query": "container_memory_working_set_bytes{namespace=\"lapilli-demo\",pod=\"checkout-…\",…}",
      "status": "ok", "series": 1, "file": "metrics/memory_working_set_bytes.json" },
    { "name": "memory_limit_bytes", "query": "…", "status": "ok", "series": 0, "file": "…" },
    { "name": "custom", "query": "…", "status": "error", "error": "HTTP 400: bad_data parse error …" }
  ] }

logs/index.json

Log files alone are ambiguous, because of two kubelet behaviors observed on real clusters:

So each container gets an entry listing its instances:

{ "containers": [ { "container": "app", "instances": [
  { "which": "current",  "file": "logs/app-current.log", "state": "terminated",
    "container_id": "containerd://be73…", "reason": "OOMKilled", "exit_code": 137,
    "finished_at": "2026-09-18T14:17:00+00:00" },
  { "which": "previous", "file": null, "state": "terminated",
    "container_id": "containerd://3201…", "reason": "OOMKilled", "exit_code": 137,
    "finished_at": "…", "unavailable": "unable to retrieve container logs for containerd://3201…" }
] } ] }

file is null when nothing was captured, with the reason in unavailable. A consumer looking for “the crash’s last words” takes the terminated instance with the latest finished_at that has a file. A previous entry appears only if the container has restarted at least once.

A tail is bounded by lines and by bytes, and a bound that was hit is recorded. A line is workload-controlled — containerd and CRI-O split a log entry at 16 KiB and each fragment comes back as its own line — so a line count alone bounds nothing a producer can put in a memory budget. A producer therefore SHOULD also bound the tail in bytes (the reference producer reads the last 2000 lines with a 4 MiB per-instance limit). A producer that truncates for its own bound MUST say so on the instance entry, because a log file cannot say it about itself and a consumer would otherwise read a cut tail as a whole one:

{ "which": "current", "file": "logs/app-current.log", "state": "terminated",
  "truncated": { "limit_bytes": 4194304, "bytes": 4194304, "cut": "newest" } }

cut says which end is missing. It is "newest" for the reference producer because the kubelet spends the byte budget forward from the start of the tail window, so the lines nearest the crash are the ones dropped — a consumer hunting last words MUST NOT present the last line of a cut: "newest" file as the crash’s last words. truncated is informational and does not change the verdict: the bound is a producer setting that was honoured, not an error in the sense of rule 6, so logs stays in collectors_run and the bundle is not PARTIAL — the same reasoning that keeps a 2000-line tail out of PARTIAL and that makes coverage.deferred a notice. truncated is absent when the tail fit, and it is optional and additive within ieb/v1 (rule 6: “status values in index files are informational”), so a verifier that does not know it ignores it.

Two kinds of unavailable are not the same, and rule 6’s producer obligation separates them. A kubelet in-band report — HTTP 200 with a one-line error, the instance was garbage-collected — is a fact about the workload, so the collector did its job and logs stays in collectors_run. Any other failure to read a log is an error in the sense of rule 6 — the log API refusing (403 because pods/log was not granted), the API server erroring, a timeout — and the producer MUST then leave logs out of collectors_run, making the bundle PARTIAL. Otherwise a bundle whose every entry is "…403 Forbidden…" verifies OK at 100% coverage, which is what the reference producer did before ../docs/design-review-round24.md §3.

lapilli verify accepts either a .ieb file or an already-unpacked directory. A .ieb is streamed and never extracted: entries are hashed as they pass, in one pass, with no seeking and no temp directory. That is why the reader rules below are written about entries rather than about files on disk, and why the traversal defence is a name check on each entry name (hashtree::check_path) rather than symlink hardening on an extracted tree. lapilli unpack extracts; verification does not.

Packing happens after sealing, so the tar byte-stream is never what the hash tree covers.

An earlier version of this paragraph said a .ieb was “unpacked to a temp dir, with path-traversal/symlink hardening”. That was never how the verifier worked — verify.rs’s own module note reads “a .ieb file is not extracted” — and it mattered, because it pointed an independent implementer at a different verification model with a different threat surface.

The ieb/v1 verification contract (normative)

Frozen for lapilli.dev/ieb/v1. A conforming verifier implements exactly these rules; a conforming producer writes bundles that pass them. The words MUST/MUST NOT are normative.

1. Container

2. Paths

Every file path other than manifest.json and the files under signature/ MUST:

3. Hash tree

4. Outside the tree

5. Manifest

manifest.json is a JSON object. A duplicate member name anywhere in it makes the bundle FAILED (otherwise one reader could use the first value and another the last). Readers MUST ignore fields they don’t know. Fields:

Field Meaning
schema_version "lapilli.dev/ieb/v1". Checked first (rule 9).
incident {id, cluster_id, trigger: {rule, firing_ts}, window: {start, end}, target?: {namespace, pod}}; bound context, compared when the caller asserts it. target is optional and additive (a bundle sealed before it existed has none): the pod the capture was about, so a lookup by what an alert carries never unpacks a bundle. It is not compared by verify; resources/pod.json is the evidence, this is the index.
producer {version, image_digest}; self-reported
signing null (unsigned) or {alg: "ecdsa-p256-sha256", key_id} where key_id = lowercase_hex(SHA-256(DER of the SubjectPublicKeyInfo with the EC point **uncompressed**)), i.e. the 91-byte DER for P-256, whatever form the key file uses
hash_tree rule 3
coverage {collectors_run: [..], collectors_intended: [..], deferred?: [..]} (rule 6)
timing {capture_started, sealed_at, capture_to_seal_ms}; self-asserted by the producer clock

incident and its members are untrusted text. incident.id, cluster_id, trigger.rule, trigger.firing_ts, window.start, window.end and target.namespace / target.pod are strings chosen by whoever sent the alert or named the workload, not by the operator, and firing_ts in particular is not required to parse as a timestamp. A consumer that renders them into a document, a message, a table or an LLM’s context MUST escape them for that sink first — otherwise a value that can end a table row can write its own heading, and a bundle becomes an injection vector into whatever reads it. This is a reader obligation about rendering and does not affect the verdict: a manifest whose incident strings are hostile is still a well-formed manifest.

6. Coverage and required files

7. Redaction record

redaction.json MUST be present (it is in the tree like any file). Its mode is one of default, strict, off; a reader MUST treat any other value as off, and an absent mode as off too — the safe reading is the one that claims least.

mode is the only field this section fixes. not_redacted and not_redacted_fields (§”Redaction and redaction.json”) are advisory, so a producer MAY omit them and a verifier MUST NOT fail a bundle that does. No mode means the bundle makes no redaction claim at all; it does not mean the policy ran.

8. Signature

9. Version dispatch and verdicts

10. Limits

A producer MUST NOT write more than 50,000 files or 1 GiB (sum of file sizes), nor a manifest.json, redaction.json or signature/ file larger than 16 MiB. A verifier’s limits MUST NOT be lower; a bundle over them (in either form) is cannot evaluate.

Why these rules (non-normative)

Contents are hashed individually and the signed payload is the literal manifest bytes, so sealer and verifier agree on bytes without JSON canonicalization or reproducible tar: this is exactly cosign’s blob model. Cross-capture reproducibility is not required. Test vectors for every rule are in test/fixtures/ieb/; test/spec/build_from_spec.py builds a bundle from this section alone (it shares no code with Lapilli) and CI verifies it.

Verification in practice

lapilli verify <bundle> [--cluster <id>] [--incident <id>] [--key <trusted.pub>]

--key given? bundle signed? result
yes yes, by that key signed:trusted-key: integrity + producer authenticity
yes by another key, or invalid FAILED
yes no FAILED
no yes signed:unpinned: self-consistent only, does not affect the verdict
no no unsigned

Authenticity comes only from a key the verifier obtained out of band. The signature/cosign.pub a producer embeds is a convenience for tools like openssl and is never trusted by lapilli verify: anyone able to rewrite a bundle can re-seal it with their own key and swap that file, or re-seal it unsigned with signing: null. Without --key, a bundle proves nothing against anyone who could write to it.

lapilli keygen writes a key pair in the formats this expects: lapilli.key (PKCS#8 PEM, for the controller’s Secret) and lapilli.pub (SPKI PEM, for --key).

cosign interop is pinned. When signing is enabled, the signature is cosign-compatible ECDSA-P256 over manifest.json, verifiable with cosign v2.x:

cosign verify-blob --key <pub> --signature <sig> --insecure-ignore-tlog manifest.json

Signature encoding (critical, get it right the first time): cosign expects base64 of the ASN.1 DER ECDSA signature (P-256, SHA-256) — not the fixed-width 64-byte IEEE-P1363 r‖s form. With RustCrypto ecdsa/p256, Signature::to_bytes() yields P1363 (cosign rejects it); use signature.to_der() then base64. Producers should emit the canonical low-S form (s ≤ n/2; RustCrypto’s p256 does not normalize by itself, and a KMS doesn’t promise it, so normalize explicitly). Verifiers must accept both forms, as openssl, Go and cosign do: (r, n − s) is the same signature. Private key = PKCS#8 PEM; publish the public key as SPKI PEM via VerifyingKey::to_public_key_pem(LineEnding::LF) (that’s what cosign --key cosign.pub consumes).

Conformance anchor = openssl (not cosign), decided after testing against real binaries. cosign v3 (verified against 3.1.3) has removed the detached --signature flag — verify-blob now requires a --bundle (.sigstore.json protobuf) even with --key. So a pinned-cosign-CLI gate is a moving target. Instead the v0.1 conformance gate verifies the signature with openssl — a neutral, version-stable check that the bytes are a correct ECDSA-P256-SHA256 DER signature over manifest.json:

base64 -D -i signature/manifest.sig -o sig.der           # (Linux: base64 -d)
openssl dgst -sha256 -verify signature/cosign.pub -signature sig.der manifest.json
# -> "Verified OK"

This is run by scripts/verify-conformance.sh and gated in CI. Emitting a cosign v3 .sigstore.json bundle for direct cosign verify-blob --bundle --key interop is a v0.2 item, tied to sigstore-rs maturing. Our signature is already standard-correct (openssl confirms), so v3 interop is packaging, not cryptography.

Signing configs (see ../DESIGN.md §5 for the full capability matrix)

Config Integrity after sealing Authenticity Independent time Air-gap Availability
unsigned (default) ⚠️ accidental change only (hash tree) ❌ ❌ ✅ v0.1
static-key ECDSA ✅ ✅ ❌ ✅ v0.1 (opt-in)
KMS ECDSA ✅ ✅ ❌ ✅ v0.1 (opt-in)
+ RFC 3161 TSA ✅ ✅ ✅ ✅ v0.3
keyless + Rekor ✅ ✅ ✅ ❌ v0.3 (spike)

The first column is integrity after sealing, against whoever can reach the bundle — not “the hash tree is well-formed” (it is, in every row). Unsigned, the tree catches a corrupted byte and catches nothing at all against anyone able to rewrite the file, because they recompute the tree and rewrite the manifest: see “Verification in practice” above, which is the same statement in prose. That row read ✅, and KMS availability read v0.2, in an earlier version of this table, contradicting both ../DESIGN.md §5 and the spec’s own prose. The table was the error; §5 remains the full capability matrix and this is the same five rows.