Kode-1
Playbook

AI evidence and audit readiness

For risk and compliance leaders, and the platform owners who hold the evidence

15 July 2026

When scrutiny arrives — auditor, regulator, board, or journalist — the questions about an AI system are predictable. Why did it decide what it decided? Who approved it? What data shaped it? What changed since, and who approved that? How do you know it still works? Organisations rarely fail these questions for lack of governance. They fail them because the evidence was never captured, or was captured somewhere nobody can find under pressure. This playbook is the architecture for capturing it. 1. The questions, written down Start by writing the actual questions your obligations imply, in plain language, for each material system. The prudential lens supplies most of them: operational risk wants to know the system's failure modes and their controls; information security wants to know its access and its data paths; the board wants to know who is accountable. A question you have written down is a question you can rehearse. A question you have not is the one that finds the gap. 2. Evidence at design Every material system gets a decision record at inception: its purpose, the decision it touches, the data it consumes and where that data comes from, the evaluation criteria it must meet before production, and the named approvals. This is one document, produced once, kept current. It is the anchor every later question hangs from — and it is nearly impossible to reconstruct honestly after the fact. 3. Evidence at deploy Version everything that shapes behaviour: model, prompts, retrieval configuration, guardrails. Record the evaluation results that gated the release, and who accepted them. A deployment you cannot reproduce is a deployment you cannot explain; an explanation that starts with "we believe it was configured to…" has already failed. 4. Evidence at run Production evidence accumulates on its own if the platform is built to keep it: monitoring against the evaluation baseline, drift measurements, incidents and their resolutions, and the outcomes of human review — including the overrides, which are the richest signal you hold about where the system's judgement and yours diverge. Retention follows your record-keeping obligations, decided once, applied automatically. Capture is designed once, at the platform layer, so every new system inherits it — evidence per system should get cheaper as the estate grows, not more expensive. 5. Attestation posture AI reaches most boards through the standards that already exist — operational risk and resilience, information security, privacy — not through a parallel AI regime. Architect the evidence to serve those attestations. When the operational-resilience attestation comes due, the AI systems inside critical operations should appear in it with the same rigour as any other system: mapped, evidenced, tested. An AI evidence base that cannot feed the existing attestation calendar is a second filing cabinet, not a capability. 6. The dry run Rehearse before it counts. Pick one production system, convene the people who would face a real review, and ask the written questions cold. Time how long each answer takes and note where the answer is a person's memory rather than a record. An answer that takes three weeks to assemble is a finding — you have simply chosen to discover it privately. The gaps become the backlog, in priority order. Repeat the rehearsal twice a year: the second one is where the improvement shows, and the habit is the capability. 7. The first 90 days Write the questions for your two most material systems. Create their design decision records, backfilled honestly. Turn on deploy-time versioning and gate-keeping for one of them end to end. Schedule the first dry run for day 60, and let what it surfaces set the next quarter's work. Evidence produced as a by-product of operating always beats evidence assembled the week before the audit — in cost, in credibility, and in what it does to the organisation's nerves when the request actually lands.