Skip to content
NOMARK
← Thinking

13 September 2026

Decision Record

TL;DR: Every AI rule written for financial services since 2024 converges on one demand. A record. Firms understand the obligation. They fail at the schema, the write path, and the retention decision. This piece is the build specification nobody publishes. * Show a decision your system made, prove what it was, why it happened, which control was in force, and who owned it. * APRA says it in principle-based language. The EU says it in Article 12. The SEC said it in 2022 without mentioning AI at

TL;DR: Every AI rule written for financial services since 2024 converges on one demand. A record. Firms understand the obligation. They fail at the schema, the write path, and the retention decision. This piece is the build specification nobody publishes.

  • Show a decision your system made, prove what it was, why it happened, which control was in force, and who owned it.
  • APRA says it in principle-based language. The EU says it in Article 12. The SEC said it in 2022 without mentioning AI at all.
  • All three are asking for the same artefact. Almost nobody has built it.
  • The failures are consistent: wrong schema, broken write path, poor retention decisions.
  • The result twelve months later: a data lake full of telemetry that cannot answer a supervisor's question.

This is a build specification for that artefact. It assumes you accept the obligation and moves to what nobody publishes: what goes in the record, where it sits in the stack, how long you keep it, and how you know it works.

What the rules actually say

Start with the Australian stack, because it is the most specific about operations and the least specific about AI.

CPS 230 commenced on 1 July 2025. It requires a register of critical operations, with tolerance levels covering maximum disruption duration, maximum acceptable data loss, and minimum service levels under alternative arrangements. Paragraph 27 is the one that matters here: entities must identify and document the processes and resources needed to deliver critical operations, including people, technology, information, facilities, and service providers, the interdependencies across them, and the associated risks, obligations, key data, and controls.

Read that clause as a data model and it is already most of a decision record. Obligations, key data, controls, interdependencies, owners. The standard does not tell you to build one. It tells you that you must be able to describe one.

Paragraph 32 requires incidents and near misses to be identified, escalated, recorded, and addressed in a timely manner. Paragraph 33 sets 72 hours for notification of incidents likely to have a material financial impact. Paragraph 42 sets 24 hours where a critical operation is disrupted beyond tolerance. Those clocks require a record to exist before the incident. A record assembled after the fact is not a record.

CPS 234 adds the testing obligation. Paragraph 27 requires systematic testing of information security controls, with frequency commensurate with the rate at which threats change, the criticality of the asset, the consequences of an incident, exposure to environments the entity cannot control, and the materiality and frequency of change to the asset itself. Paragraph 31 requires the sufficiency of that programme to be reviewed at least annually.

Two things follow. An AI system that changes weekly demands testing at that same frequency. The evidence of that testing has to be durable, because paragraph 36 gives you ten business days to notify APRA of a material control weakness you cannot remediate in time.

The accountability layer came next. The Financial Accountability Regime (FAR) commenced for banks on 15 March 2024 and for insurers and superannuation trustees on 15 March 2025. Administered jointly by APRA and ASIC. Accountable persons sign statements that must reflect actual accountability as it operates in practice, and an accountability map sets out reporting lines across the group.

The consequences are personal. At least 40 per cent of an accountable person's variable remuneration is deferred for a minimum of four years, with reduction obligations on breach. The regulators can disqualify an individual. Hold that four-year number: it sets a floor on how long your evidence has to survive.

Key point: CPS 230 read as a data model already describes most of a decision record. The standard does not tell you to build one. It tells you that you must be able to produce one.

What APRA actually wrote about AI

On 30 April 2026 APRA wrote to all regulated entities. The letter is signed by Therese McCarthy Hockey and it reports on a targeted engagement with selected large banks, insurers and superannuation trustees conducted in late 2025.

Two corrections to the commentary that followed it, because both are load-bearing.

The letter cites no prudential standard by number. Not CPS 230, not CPS 234, not FAR. It says instead that APRA's principle-based prudential framework is technology and vendor agnostic, and that APRA requires regulated entities to ensure appropriate risk management of AI, including setting risk appetite, managing AI-related exposures, and ensuring appropriate oversight and accountability. Published analysis asserts the letter placed AI governance under FAR. It did not. The letter deliberately declines to name an instrument. That is a choice, not an oversight.

The second correction matters more. Much of the commentary counted the word "gaps" and treated the tally as the finding. The tally is not the finding. This sentence is:

Gaps include weak controls over post deployment monitoring, weak model behaviour monitoring, change management, and decommissioning of AI capabilities.

Four control domains, named: post-deployment monitoring, model behaviour monitoring, change management, and decommissioning. Every one is a record-keeping problem before it is a control problem. You cannot monitor behaviour you did not capture. You cannot evidence a decommissioning you did not log.

The single most useful line in the letter is about assurance:

APRA also observed reliance on point in time and sample based assurance methods, despite these methods being ill suited to probabilistic models that learn, adapt and degrade over time.

That sentence ends the quarterly control test for AI systems. If the model drifts between samples, a sample tells you about a system that no longer exists. Continuous assurance of a system that changes continuously requires a record of every member of the population. Not a sample. The population.

APRA also noted a tendency to treat AI risk as "just another technology," and that many boards are still developing the technical literacy needed to provide effective challenge. The letter closes by stating that where entities fail to adequately identify, manage, or control AI risks proportionate to their size, scale, and complexity, APRA will take stronger supervisory action and, where appropriate, pursue enforcement.

In late August 2026 APRA and ASIC followed with an information paper drawn from nine industry roundtables held across June and July, attended by more than 600 people. The board-facing ask: settle risk appetite, escalation authority, recovery priorities, and communication strategies before a crisis, because frontier AI compresses incident response timeframes. Settled in advance means written down in advance. That is a record too.

Key point: APRA's four named control failures all share a root cause. They are record-keeping failures first. Control failures second.

The international picture, and one hole in it

The EU AI Act is the only instrument that specifies logging as a technical obligation. Article 12(1) requires that high-risk AI systems shall technically allow for the automatic recording of events over the lifetime of the system. Article 19 requires providers to keep those logs for a period appropriate to the purpose, at a minimum six months. Article 26 places the equivalent duty on deployers, which is the bank or insurer rather than the model vendor.

Article 19(2) is the clause almost nobody quotes. It says that providers which are financial institutions subject to internal governance requirements under Union financial services law shall maintain the logs as part of the documentation kept under that law. The EU expects AI logs to live inside your existing regulatory record-keeping, not in a parallel AI compliance silo. That is a design instruction, and it is the right one.

Scope is narrower than it appears. Annex III makes creditworthiness assessment and credit scoring of natural persons high-risk, with an explicit carve-out for systems used to detect financial fraud. It makes risk assessment and pricing high-risk for life and health insurance. General insurance is not named. Corporate credit is not named.

Timing has moved. The Digital Omnibus amending regulation entered into force in late July 2026 and deferred the high-risk obligations for standalone Annex III systems from 2 August 2026 to 2 December 2027. The logging duty is coming. It is not yet biting. Firms treating December 2027 as distant should note: the record you need in December 2027 describes decisions made well before it. You cannot retrofit a decision record onto decisions already taken.

Then the hole. On 17 April 2026 the Federal Reserve, the OCC and the FDIC jointly issued revised model risk management guidance that supersedes SR 11-7, the 2011 letter that has governed model risk for fifteen years. The new guidance retains effective challenge, narrows the definition of a model to exclude simple arithmetic and deterministic rules, and introduces a 30 billion dollar asset threshold for where it is most relevant. It also says this:

Generative AI and agentic AI models are novel and rapidly evolving. As such, they are not within the scope of this guidance.

Read it twice. The incumbent US model risk framework has been rewritten, and generative and agentic AI have been written out of it. The agencies signalled a request for information to follow. As of September 2026 there is no evidence it has been issued.

A US bank running an agentic system in 2026 is not covered by SR 11-7, because SR 11-7 is withdrawn. It is not covered by the replacement, because the replacement excludes it. The obligation has not disappeared. It has moved to general safety and soundness, consumer protection, and whatever the firm told its board it was doing. Firms that governed AI by mapping it to SR 11-7 have lost their map.

IOSCO published a supervisory toolkit for AI in capital markets in May 2026, naming recordkeeping and reporting as one of four supervisory focus areas alongside governance and risk management, third-party and outsourcing risk, and disclosure. The Financial Stability Board consulted in June 2026 on twelve sound practices for responsible AI adoption, explicitly stating they are not intended to establish an international standard. The ECB has asked significant institutions for an action plan on AI-driven cyberattack defence by 31 October 2026. In the UK, SS1/23 remains the model risk statement and the FCA has said it will not introduce additional AI regulation, relying on existing frameworks instead.

Key point: The US regulatory void on agentic AI is the most important structural fact in the current landscape. Firms that governed AI by mapping to SR 11-7 have lost their map. There is no replacement instrument yet.

What nobody has specified

I looked for a published reference architecture. A regulator or standards body describing, in technical terms, what an AI evidence layer contains and how it is assembled.

There is not one.

MAS comes closest through the Veritas initiative, which produced assessment methodologies and open-source code for credit scoring and customer marketing fairness. ISO/IEC 42001 certifies a management system rather than a schema, and adoption remains thin. The best available industry estimate puts certified organisations in the low hundreds by early 2026, concentrated among firms selling AI rather than firms running it inside regulated operations. The NIST AI Risk Management Framework is the most detailed public source. Its MEASURE 2.8 subcategory says to instrument the system for measurement and tracking by maintaining histories, audit logs, and other information, including human oversight statistics and override rates.

That is guidance, not a specification. Every regulator has specified a capability and left the artefact design to you. Nobody has said what the record looks like.

What follows is what gets built, and why each part is there.

Key point: No regulator has published a reference architecture for an AI decision record. The obligation is specified. The artefact is not. That is the problem this piece addresses.

The unit is the decision, not the log line

The first design error is logging at the wrong grain.

Application logs are ordered by time and scoped to a process. A decision record is ordered by decision and scoped to an obligation. One decision may span six services, two model calls, a retrieval step, a control evaluation, and a human override. It has to be retrievable as a single object by someone who knows nothing about your service topology.

The trade capture analogy is exact. A trade is not a sequence of messages across a bus. It is a booked object with an identifier, and every message is an attribute of that object. Firms that log AI the way they log infrastructure end up doing forensic joins across a timeline under deadline. A 72-hour notification clock does not permit that.

Define the boundary before the schema. A decision is the smallest unit for which someone can be held accountable. For a credit adjudication, that is one application. For an agentic settlement process, it might be one exception or one instruction, depending on where authority actually sits. Get this wrong and you produce either a record nobody can act on or a volume nobody can afford.

Key point: The first design error in AI evidence programmes is logging at the wrong grain. Log by decision, not by process. The record has to be retrievable as a single object.

Agents move the boundary

Everything above assumes you can point at a decision. Agentic systems make that harder, and the difficulty is structural rather than technical.

A single-shot model call has an obvious boundary. An agent that plans, calls three tools, revises, and acts has produced one outcome through a dozen intermediate steps, several of which changed state. Record only the outcome and you cannot explain it. Record every step as a decision and you will drown, and you will also be wrong, because most of those steps are not decisions anyone can be held accountable for.

The rule: a step is a decision when it changes state outside the system, or when it commits the firm to something. A retrieval is not a decision. A plan revision is not a decision. Sending an instruction, releasing a payment, adjusting a limit, closing a break: all of those are. Record the committing steps as decisions, and attach the intermediate trace as a referenced artefact rather than promoting it to a record of its own.

This matters most in the United States right now. The revised model risk guidance issued in April 2026 excludes generative and agentic AI from its scope. The guidance it replaced no longer applies. A firm running agentic settlement or agentic adjudication in that market has no instrument telling it what to document. It has the general duty, its board-approved statements, and whatever it can produce when asked.

The Bank of England has said publicly that existing frameworks were not built to contemplate autonomous agents. The FCA has said legislation will never keep up. Both are true. Neither helps in an examination. In the absence of an instrument, the accountable person's exposure is defined by what the firm said it would do and whether it can show it did.

Build the record now rather than waiting for the rule. When the rule arrives, it will ask about decisions made before it existed.

Key point: For agentic systems, record only the committing steps. Those are the steps where state changes outside the system or the firm is committed to something. Attach the full trace as a referenced artefact, not as a record of its own.

The record, field by field

Nine groups. Each earns its place against a named obligation.

Identity. A decision identifier unique across the estate, a timestamp with an explicit clock source, and the legal entity. Without a clock source you cannot defend ordering, and ordering is the first thing challenged.

Subject. What the decision is about, referenced rather than embedded. A customer reference, an account, an instruction, a position. Reference rather than copy, because the subject record has its own retention and erasure rules and you do not want to inherit them.

System. Model identifier, model version, provider, and a hash of the deployed artefact where you control it. Deployment configuration version. Prompt or policy template version. Sampling parameters. This group answers "what was running," and it is the group most often missing, because teams record the model name and not the version that produced the output.

Inputs. A hash of the full input, the identifiers and hashes of any retrieved context, and the feature set reference. Hashes rather than payloads by default. The hash proves what was seen without duplicating the customer data into a second store with a different retention clock.

Reasoning. The output, any structured rationale the system produced, confidence or score, and the alternatives considered where the architecture exposes them. This is the group people over-promise. Capture what the system actually emitted. A reconstructed post-hoc explanation is not evidence of reasoning, and presenting it as such is worse than capturing nothing.

Controls. Which controls were in force at the moment of decision, which evaluated, which fired, the thresholds applied, and the pass or fail result. This is the group that converts a log into a control record, and it is the group that answers CPS 234 paragraph 27 when a supervisor asks how you know a control was effective across a population rather than a sample.

Human oversight. Whether a person reviewed, who, what they were shown, what they changed, and the stated reason for an override. NIST names override rates explicitly as a measurement target. You cannot compute an override rate you did not record, and override rate is the single most diagnostic number in any human-in-the-loop design.

Outcome. The action taken and a reference to its downstream effect. A decision record that stops at the recommendation cannot tell you whether the recommendation was followed, and the gap between recommendation and action is where most operational loss actually occurs.

Accountability and integrity. The accountable person or role, resolved through the accountability map rather than typed in free text. Then the integrity fields: the hash of the previous record in the chain, and a signature.

That last group is what makes it evidence rather than data. It is also the group that gets cut first when the schema meets a sprint deadline.

Key point: Nine field groups. Each earns its place against a named obligation. The accountability and integrity group is the one that gets cut first. It is the one that cannot be cut.

Four properties that turn a log into evidence

Logging is for you. Evidence is for someone who does not trust you. That distinction drives every remaining design decision. Engineering teams are rarely asked to hold both in mind at the same time.

Four properties follow.

Attributable. Every record ties to an actor, human or machine, and to a specific version of a specific system. The ISO 42001 audit test, as practitioners report it: can you reconstruct what happened, who or what performed the activity, which resource was affected, and what action followed, from the records alone. From the records alone. Not from the records plus an engineer who remembers.

Complete where it counts. You cannot sample the thing you are required to prove. The cost argument against this is a misreading, and it is addressed in the capture section below.

Tamper-evident. Hash chaining gives you ordering, where each entry carries the hash of its predecessor so any alteration breaks the chain forward. Merkle trees give you efficient proof, so that demonstrating a single record belongs to the set costs logarithmic rather than linear work. Neither is exotic. Both are decades old and well understood.

Reconstructable. Given a record, can you re-derive how the decision was reached and explain why it came out that way? This is weaker than replay. Replay means running the same input through the same model and getting the same output, which for a probabilistic system under a changed model version you will often not achieve. Reconstruction means explaining the decision as it happened, with evidence of what was in force at the time. Design for reconstruction. Promise replay only where you have pinned versions and fixed seeds, and say so explicitly rather than letting a board assume it.

Key point: Four properties separate a log from evidence: attributable, complete where it counts, tamper-evident, and reconstructable. The last one is the hardest and the most misunderstood.

The WORM assumption is out of date

Firms still specify write-once media because someone remembers that regulation demanded it. That changed in 2022.

The SEC's amendments to Rule 17a-4 ended the exclusive WORM mandate for broker-dealers and introduced an audit-trail alternative. The alternative requires a complete time-stamped audit trail that includes all modifications to and deletions of a record or any part thereof, the date and time of actions that create, modify, or delete the record, and where applicable the identity of the individual doing so. Effective January 2023, compliance from May 2023.

This is the most useful precedent available, and it is not an AI rule. The canonical financial-services regulator, on the canonical records question, moved from "the medium must prevent change" to "the system must evidence change." Those are different engineering problems. The second is the one a hash-chained append-only store solves natively.

Design to the audit-trail standard rather than the media standard. It is a better fit for AI, where the record has to be queryable for continuous assurance rather than merely preserved.

Key point: The SEC's 2022 Rule 17a-4 amendment is the most useful precedent for AI record design. The move from write-once media to audit-trail evidence is the right model. Build for queryability, not just preservation.

The schema problem, stated honestly

There is one candidate for a vendor-neutral schema, and you need to know its actual status before you build on it.

The OpenTelemetry semantic conventions for generative AI define attributes for provider, request and response model, token usage, conversation identifier, input and output messages, and dedicated spans for agent invocation, chat and tool execution. Adoption across observability vendors is real.

Every one of those attributes carries Development status. None is stable. In June 2026 the conventions were moved out of the main semantic-conventions repository into a dedicated repository with no tagged release. Breaking changes have already happened: the earlier prompt and completion attributes were removed outright, and the system attribute was superseded by provider name. Different SDKs currently emit more than one generation of attribute names at once, so downstream consumers have to query both.

Two conclusions. Adopt the names, because inventing your own vocabulary buys nothing. Do not adopt the contract, because you cannot tell a supervisor your evidence conforms to a specification with no stable release that changed shape twice this year.

Put a normalisation layer between collection and storage. Your decision record schema is yours, versioned by you, frozen at a number you control, with a documented mapping to whatever the conventions say this quarter. The cost is one translation layer. The alternative is a schema migration in the middle of an examination.

Key point: The OpenTelemetry semantic conventions for generative AI are real, widely adopted, and entirely unstable. Adopt the names. Do not adopt the contract. Your schema needs to be versioned and controlled by you.

The record above the decision

One record per decision is not the whole obligation. There is a second artefact, at system level, and the two answer different questions.

The decision record answers what happened on this occasion. The system record answers what this thing is, what it was built for, what it was tested against, and where it should not be used. A supervisor asking about a single adverse outcome wants the first. A supervisor asking whether the system was fit for the purpose it was put to wants the second.

The tradition here is older than AI regulation. Model cards, as defined by Mitchell and colleagues in 2019, carry nine sections: model details, intended use, factors, metrics, evaluation data, training data, quantitative analyses, ethical considerations, and caveats and recommendations. Intended use is where out-of-scope uses belong. It is the field most often left blank and most often needed. Datasheets for datasets, from Gebru and colleagues in 2018, cover motivation, composition, collection process, preprocessing, uses, distribution, and maintenance.

NIST's framework references both directly, and its GOVERN 1.4 subcategory lists a documentation set that reads as a regulatory-flavoured superset. Among its thirteen items: business justification, scope and usage, expected risks, assumptions and limitations, training data description, algorithmic methodology, alternatives evaluated, testing and validation results, dependencies, and deployment and monitoring plans.

Keep the system record versioned, and have the decision record reference the version in force at the time. That link is what lets you answer the hardest question in any review: not whether the system was appropriate now, but whether it was appropriate then, under the configuration that actually ran.

Most firms have some of this scattered across model validation reports, architecture decision records, and vendor documentation. Scattered is not the same as retrievable.

Key point: Two artefacts are required: one at decision level, one at system level. The decision record and the system record answer different supervisory questions. Both need to be versioned and linked.

What you capture, and what you must not

Here is where cost and privacy collide, and where most implementations either overspend or quietly fail.

The standard practitioner advice for AI observability is tiered sampling: keep every error, keep anything unusually slow, keep anything a user rated badly, and keep a small random slice of successful traffic. The reasoning is that traces carrying full prompt and response text run to tens of kilobytes each, and full capture at volume is expensive.

That advice is correct for observability and wrong for evidence, and conflating the two is the expensive mistake.

For in-scope decisions, invert it. Every decision inside a critical operation or a high-risk use case gets a complete record, every time, with no sampling. Sample the diagnostic telemetry around it as heavily as you like. The volume argument collapses, because the decision record is small. Identifiers, hashes, versions, control results, and a rationale field run to a few kilobytes, not tens, once you stop copying payloads into it.

The schema above references and hashes inputs rather than embedding them. That is a cost decision and a privacy decision at the same time, and they point the same direction.

Redact at capture, in process, before anything durable is written. Masking by default, irreversible. Tokenisation only where you genuinely need to re-link a value later, and then vault-backed. Regular expressions handle fixed-shape identifiers cheaply and deterministically. Names, addresses, and employers have no fixed shape and need a model, which costs more, so scope it to the fields that require it.

For genuinely sensitive flows, record metadata only and reserve full content capture for a small, access-controlled, consented path. A decision record proving which control fired and who owned it does not require the customer's free text duplicated into a second store.

Key point: The cost and privacy problems in AI evidence resolve in the same direction. Stop copying payloads. Reference and hash. Complete coverage of in-scope decisions becomes affordable once the record is built from identifiers and hashes rather than full content.

The vendor problem

You cannot produce evidence your provider does not emit. This breaks more evidence programmes than any engineering decision. It is usually discovered late.

If the model runs behind someone else's API, the fields in the System group are only as good as what comes back. Providers change model versions behind stable endpoint names. They deprecate on their own schedule. Some return a version identifier, some return a family name, some return nothing useful at all. A decision record that names a product with no version cannot support a drift argument twelve months later. You cannot establish what changed or when.

CPS 230 gives you the lever and most firms do not pull it. Paragraph 49 requires a register of material service providers. Paragraph 51 requires that register to be submitted to APRA annually. Paragraph 54 requires formal, legally binding agreements covering service levels, data ownership, audit access, sub-contracting and termination.

An AI provider serving a critical operation is a material service provider. So put the evidence requirements in the contract, where they belong, rather than trying to reconstruct them from telemetry afterwards. Four clauses do most of the work.

A stable version identifier returned with every response, and notification before any change to the model serving a named endpoint. Without this the rest is decoration.

A defined deprecation notice period, long enough to revalidate rather than merely to migrate. Migration is an engineering problem. Revalidation is a control problem and takes longer.

Log export in a form you can hold, on your retention schedule, not theirs. Their retention is set by their business, and it will be shorter than your liability.

Audit access that survives sub-contracting, because the model you are buying may be served by an infrastructure provider you have never assessed, and concentration in that layer is precisely what the FSB has flagged as a systemic vulnerability.

IOSCO's supervisory toolkit makes the underlying point plainly: accountability stays with the firm regardless of who operates the model. The EU makes the same move by placing log-retention duties on deployers as well as providers. You cannot contract out of the obligation, so contract in the capability instead.

Where the provider will not supply it, that is a finding to put in front of the board with the exposure named, not something to absorb quietly into the architecture. The firms that will struggle most in 2027 are the ones running critical decisions on opaque endpoints under contracts written for ordinary software.

Key point: CPS 230 paragraph 54 gives you the lever. An AI provider serving a critical operation is a material service provider. The evidence requirements belong in the contract. Four clauses do most of the work: stable version identifiers, deprecation notice, log export on your retention schedule, and audit access that survives sub-contracting.

How long you keep it

There is a persistent belief that CPS 230 mandates a seven-year retention period for this material. I have read the standard looking for it. It is not there.

CPS 230 requires registers, documented processes, and recorded incidents. It does not state a numeric retention duration. Neither does CPS 234. Neither does CPG 235, which is a practice guide and creates no enforceable requirement. If your retention policy cites a CPS paragraph for a seven-year figure, the citation is wrong. A supervisor who checks will find it.

The actual numbers come from elsewhere, and they do not agree with each other. The EU AI Act sets a floor of six months for logs. FINRA's recordkeeping rule defaults to six years for records without a specified period. Your own jurisdiction's corporations and privacy legislation will set others.

So set retention by the liability, not by the log. Three questions decide it.

How long can someone be held personally accountable for this decision? Under FAR, variable remuneration is deferred at least four years, which means an accountable person can face a remuneration consequence for a decision four years after it was made. Evidence that expires at two years cannot defend them.

How long is the limitation period for the harm this decision could cause? A credit decision that damages a consumer has a longer tail than an internal routing decision.

How long does the model version live? You cannot assess drift against a baseline you deleted.

Take the longest of the three, then tier the storage rather than shortening the period. Full records hot for the operational window, compressed and chained in cold storage after, with the integrity chain intact across the tier boundary. The chain makes cold storage defensible, because it proves the archived records are the ones you wrote.

Key point: Set retention by the liability, not by the log. Three questions decide it: personal accountability duration, limitation period for harm, and model version lifespan. Take the longest answer.

Where it sits, and the decision people resist

This is the part of the design that gets argued about. The argument is always the same.

Write the decision record inline, synchronously, before the action commits. If the record cannot be written, the decision does not execute.

Teams resist this hard. It couples an operational path to a compliance store, adds latency, and introduces a new failure mode where the evidence layer can stop the business. Every one of those objections is true.

They are also the objections that were raised, and settled, about trade capture thirty years ago. You do not execute and then write the booking record on a best-efforts queue, because a queue that drops under load drops precisely when volume is highest, which is precisely when you will be asked about it. The failure correlates with the event you need evidence of.

An asynchronous evidence pipeline fails silently and selectively. It will be complete for ordinary traffic and incomplete for the incident. The incident is the only time anyone reads it. A firm discovered this during a post-incident review, with a queue that had shed load for four minutes, and a four-minute hole exactly where the decisions in question sat.

The architecture that works: a fail-closed write to a fast local append-only log, acknowledged before the action commits, with asynchronous shipping from there to the durable store. Local write is sub-millisecond. The chain is established at the point of decision. Shipping can retry without the business path depending on it. You get the guarantee where it matters and pay for it once.

Scope it to in-scope decisions only. This is not a rule for every inference in the building. It is a rule for decisions inside critical operations, and CPS 230 already made you enumerate those.

Key point: Write synchronously, before the action commits. The trade capture analogy is exact. An asynchronous evidence pipeline fails selectively. It will be incomplete for the incident, and the incident is the only time it matters.

Assurance changes shape

Return to APRA's line about point-in-time and sample-based assurance being ill suited to models that learn, adapt and degrade. If you accept it, your control testing has to change, and the decision record is what makes the change possible.

Sample-based testing asks whether twenty-five cases passed. Population testing asks what proportion of every decision in the quarter had the control in force, what proportion fired, and how that moved week by week. The second question is only answerable if the control result is a field in a complete record.

Three measures fall out immediately, and all three are computed from the record store rather than from a testing exercise.

Control coverage: the proportion of in-scope decisions where the control was present and evaluated. Anything below 100 per cent is a finding, and the informative part is always what failed coverage, not the aggregate number.

Control fire rate over time: a rate that moves without a corresponding change in traffic means either the population shifted or the model did. Both need explaining. This is drift detection using the controls rather than the model internals, which matters because you own the controls and often do not own the model.

Override rate by reviewer and by decision type: a rising rate signals model degradation or reviewer loss of confidence. A falling rate toward zero is the more dangerous reading, because it usually means review has become rubber-stamping. A human oversight control that never disagrees is not a control.

Report those continuously rather than quarterly. That is what continuous assurance means in practice. It is buildable the moment the record exists.

Key point: Three measures come directly from a complete record: control coverage, control fire rate over time, and override rate. None requires a testing exercise. All require a complete record.

Build it in ninety days

Sequence matters more than completeness. This order works because each step produces something defensible on its own.

Weeks one and two: enumerate. List the decisions inside your critical operations that involve a model, and define the decision boundary for each. This is an operating-model exercise, not an engineering one. Do it with people who know where accountability actually sits. Expect the list to be shorter than feared and the boundaries to be more contested than expected.

Weeks three and four: freeze schema version one. Nine field groups, documented, versioned, with a mapping to the OpenTelemetry names where they exist. Resist every request to add a field that no obligation requires. A schema that tries to anticipate everything ships late and is migrated anyway.

Weeks five to eight: instrument one critical operation end to end, with the fail-closed local write. One operation. The value of a single complete operation is substantial compared with partial coverage of six, because it proves the write path under real load and produces a population you can actually test against.

Weeks nine and ten: add the integrity chain and set the retention policy against the three questions above. Write the retention decision down with its reasoning, because the reasoning is what a supervisor asks about and nobody records it.

Weeks eleven and twelve: run the drill.

The drill that tells you it works

Pick a date in the past inside your instrumented operation. Pick one hundred decisions at random from it. Reconstruct them.

For each one, produce: the model and version that ran, the input as hashed and its verification against the source, the controls in force and their results, the human involvement if any, the action taken, and the accountable person at that moment. Then verify the integrity chain across the whole set.

Time it. Do it with the people who would actually do it, not the team that built the system.

The failures are predictable, and they are the point of the exercise: model version missing because only the name was captured; control result absent because the control ran in a service that was not instrumented; the accountable person resolving to a team rather than a name because the accountability map was never wired in; a retrieval context that cannot be verified because the source document was updated and nobody hashed it.

Every one of those is cheap to fix in week twelve. Each is extremely expensive to fix during a 72-hour notification window.

When one hundred records reconstruct cleanly in an afternoon, you have an evidence layer. Until then you have logging, however much of it you have.

That is the whole standard, and it is deliberately a low bar. One hundred records. One afternoon. Most firms cannot clear it today. The ones that can did not get there by writing a policy.


Frequently asked questions

What is an AI decision record in financial services?

An AI decision record is a structured artefact that captures what an AI system decided, why, which controls were in force, and who owned it. It is distinct from application logs or observability telemetry. It is scoped to obligation, not to process, and must be retrievable as a single object.

Which regulations require AI decision records?

No regulation specifies the artefact by that name. APRA's CPS 230 and CPS 234, the EU AI Act Articles 12, 19, and 26, FAR, NIST MEASURE 2.8, and IOSCO's supervisory toolkit all specify a capability: reconstruct the decision, evidence the control, name the accountable person. The artefact design is left to the firm.

Does APRA's April 2026 letter place AI governance under FAR?

No. The letter deliberately cites no prudential standard by number. It states that APRA's principle-based framework is technology and vendor agnostic. Published analysis that maps the letter to FAR is reading an implication that is not in the text.

What is the difference between a decision record and a log?

A log is ordered by time and scoped to a process. A decision record is ordered by decision and scoped to an obligation. One decision may span multiple services, model calls, and a human override. A log cannot answer a supervisor's question about that decision. A record can.

What happened to SR 11-7 for generative and agentic AI?

The Federal Reserve, OCC, and FDIC issued revised model risk guidance on 17 April 2026 that supersedes SR 11-7. The replacement explicitly excludes generative and agentic AI from its scope. As of September 2026, no successor instrument for those systems has been issued. US firms running agentic systems are governed by general safety and soundness obligations and their own board-approved statements.

How long should AI decision records be retained?

CPS 230 does not specify a numeric duration. Set retention by the liability. Take the longest of: the personal accountability period under FAR (at least four years), the limitation period for harm the decision could cause, and the lifespan of the model version. Tier storage across that period. Do not shorten the period to reduce cost.

What is the WORM requirement for AI records?

The SEC's 2022 amendments to Rule 17a-4 ended the exclusive write-once, read-many (WORM) mandate for broker-dealers and introduced an audit-trail alternative. The alternative requires a complete time-stamped trail of all modifications, deletions, and the identities responsible. A hash-chained append-only store satisfies this requirement and is better suited to AI systems that need queryable continuous assurance.

Why is synchronous evidence writing required?

An asynchronous pipeline fails selectively. It drops load precisely when volume is highest, which correlates with the events most likely to be examined. The trade capture analogy is exact: you do not book a trade on a best-efforts queue. A fail-closed write to a local append-only log, acknowledged before the action commits, with asynchronous shipping to the durable store, gives the guarantee at sub-millisecond cost.

What is the ninety-day build sequence for an AI evidence layer?

Enumerate decisions in critical operations (weeks one and two), freeze schema version one (weeks three and four), instrument one operation end to end with fail-closed write (weeks five to eight), add integrity chain and set retention policy (weeks nine and ten), run the reconstruction drill (weeks eleven and twelve).


Key takeaways

  • Every major AI regulation specifies the same capability: reconstruct the decision, evidence the control, name the accountable person. None specifies the artefact. The design is yours.
  • The three most common failures are schema, write path, and retention. Each produces the same outcome: telemetry that cannot answer a supervisor's question.
  • The unit of record is the decision, not the log line. Define the boundary by accountability, not by process or service.
  • Four properties convert a log to evidence: attributable, complete where it counts, tamper-evident, and reconstructable. Reconstructable is not the same as replayable.
  • The US regulatory void on agentic AI is the most significant structural fact in the current landscape. SR 11-7 is withdrawn. Its replacement excludes generative and agentic AI. No successor instrument has been issued.
  • Write synchronously, before the action commits. Asynchronous evidence pipelines fail precisely during the events that will be examined.
  • The drill is the test: one hundred records, one afternoon, reconstructed by the people who would actually do it. When it clears, you have an evidence layer. Until then, you have logging.

Get these as they’re published.

Long-form on AI governance in regulated firms — what the control gap actually is, what regulators are asking for, and what the evidence has to look like. Roughly monthly. No pitches.

Your address is used to send these pieces and nothing else. Unsubscribe from any email.