September 2026EU AI Act Article 12 is the record-keeping obligation for high-risk AI systems, and it is three paragraphs long. The problem is not that it is hard to read. The problem is that the pages explaining it have drifted apart into two camps, and both leave a build team stranded.
Law firm explainers restate the paragraph accurately and stop before any decision an engineer could act on. Developer guides go the other way and supply the missing specificity by inventing it, presenting mandatory field schemas and tamper-proofing clauses that are not in the Regulation. A team that builds to the second kind of page is building to law that does not exist, which is a strange way to spend a compliance budget.
So here is the text, then the three requirements that get attributed to it and are not there, then where your log design actually comes from.
First the scope gate, because it does most of the work. Article 12 sits in Chapter III, Section 2, among the requirements for high-risk AI systems. It is not a general rule that every LLM integration must keep the same logs. If your system is not high-risk in its intended use, Article 12 does not reach it, and the first question is always classification rather than architecture.
For systems that are in scope, paragraph 1 is the whole obligation:
High-risk AI systems shall technically allow for the automatic recording of events (logs) over the lifetime of the system.
Paragraph 2 sets the standard as traceability appropriate to the intended purpose, then names the uses that capability has to serve: identifying situations where the system may present a risk within the meaning of Article 79(1), identifying situations that may lead to a substantial modification, feeding the post-market monitoring of Article 72, and supporting the operational monitoring deployers carry out under Article 26(5).
Paragraph 3 is the only place in the article that names data fields, and it names them for one category: remote biometric identification systems under Annex III point 1(a). Those must at minimum record the period of each use, the reference database the input was checked against, the input data that produced a match, and the persons involved in verifying the result.
That is the article. There is no paragraph 4.
If you want the paragraph read as a build specification, with the deadline that moved in July 2026 and what a record-keeping failure costs under the fine schedule, that is the statutory walkthrough. This page is about what happens after you have read it.
The highest-ranking developer guide on this query presents a table of "six mandatory log fields" said to derive from Article 12(1), covering system ID, operator ID, input reference, output reference, timestamp and event type, followed by a list of generic minimum events attributed to Article 12(2).
Those are reasonable fields. They are not the Regulation. Article 12(2) names purposes, not events, and the Act deliberately gives no specific guidance on what concrete data should be recorded, leaving that to the intended purpose and the risk assessment. The text sets an outcome and a context-sensitive standard; it does not prescribe one universal event schema, storage product, gateway pattern, signing method, or failure mode.
The likely origin of the error is a merge. Annex IV Section 5 governs what goes in your technical documentation, Article 12 governs what the system records, and the two get collapsed into a single checklist.
The same guide states that Article 12(3) "requires that logging be designed to detect and prevent tampering", that logs must be append-only with no modification after write, and calls this "the statutory basis for immutable audit trails in EU AI systems".
Article 12(3) is the biometric minimum-field list quoted above. The claim describes a different paragraph than the one that exists. And the Regulation's vocabulary across the whole article is record-keeping, automatic recording of events and traceability: it never says audit trail and it never says tamper-evident.
A softer version of the same mistake is easy to make in good faith. A widely shared compliance checklist puts append-only storage, cryptographic hash chaining and independent verification in a phase that sits in the same list, in the same register, as the genuinely statutory scope and retention items. Nothing there is false as engineering advice. It just reads as law.
The developer guide cites an "Art.12(4) Market Surveillance Retention". Retention lives in Article 19 for providers and Article 26(6) for deployers, which matters because those two articles turn on the words under their control and the invented paragraph does not.
Why be pedantic about this? Because proportionality is the only lever Article 12 gives you. The standard is traceability appropriate to the intended purpose, which means a team is entitled to log less for a lower-risk high-risk system and must be able to justify the line it drew. A team that believes six fields and append-only storage are hard statutory minimums has lost the ability to reason about that trade-off at all. It will over-build in one place, under-build in another, and be unable to explain either to an auditor.
The derivation runs in five steps: define the intended purpose, conditions of use and pre-planned changes; identify the situations that could lead to risk or to substantial modification; identify the validation conditions; define what to log for each identified risk; then implement and report. Log scope is an output of your Article 9 risk work, not an input to it. This is also why the derivation has to be written down. The document justifying why you log what you log is the artefact an authority will ask for, and no field list can produce it for you.
The useful design question is whether your team can reconstruct the path from an input to the output and to any consequential action the application took, using the record alone. In practice that means stable request or case identifiers, the deployed model and relevant configuration, input, retrieval and policy references in proportionate form, the processing steps, the output, and the downstream decision. Then the second half of the test, which is easy to skip: whether those records can be linked without retaining more personal data than is necessary.
Two common designs fail the first half. Logging only errors fails, because a risk you did not anticipate does not arrive labelled as an error. Sampling fails, because the one inference that later gets challenged is the one that was not sampled.
One common assumption fails the boundary. A gateway or platform record makes routed model interactions observable, but it cannot by itself document application decisions taken outside that routed boundary, and those application-level facts are what the article cares about. Your load balancer does not know which policy fired.
Most of this exists in your stack under different names. Inference logs cover identifying the inputs that cause unwanted behaviour and the drift detection that feeds post-market monitoring; experiment tracking covers backtracking to training data; data versioning identifies training and reference datasets by unique identifier; orchestration run IDs let you reconstruct a full model generation process. That mapping onto Article 12(2)(a) to (c) is the cheapest path to a defensible design, because it starts from systems that already operate rather than from a greenfield compliance platform.
The concrete implementation guidance is coming, but it is not here yet. The draft standard prEN ISO/IEC 24970 on AI system logging is intended to carry it, designed to be used alongside a risk management system of the kind Article 9 requires. Until it is in force, the derivation is the deliverable.
Where Article 12(3) does not bind you, treat its four fields as a template rather than a carve-out. When, against what, on what evidence, and on whose authority are the four questions that arrive at incident time for any high-risk system. Biometrics is simply where the legislator wrote the answers down.
Providers keep the automatically generated logs under their control for a period appropriate to the intended purpose and in any case at least six months, under Article 19. Deployers carry the corresponding duty for logs under their control under Article 26(6). Both are qualified by other applicable Union and national law, in particular data protection law.
Six months moves in both directions.
Upward, sector rules bite hard. Financial regulation such as MiFID II and DORA can demand five to seven years, and healthcare can run longer; logs under investigation are preserved until that investigation concludes, regardless of your standard period.
Downward, GDPR Article 5(1)(e) storage limitation applies to the record itself. Logs about people are personal data, so keeping them in identifiable form beyond what the purpose requires is its own violation. Pseudonymisation reduces the exposure but does not remove it, since pseudonymised data can still be personal data.
The practical consequence: retention is a decision per record category, documented with its reason, not a single TTL applied to a log bucket. Two related pieces of hygiene are worth naming because they are expensive to retrofit. Scrub personal data before write rather than after, since raw data that hits disk and is redacted later still existed as a processing event. And pseudonymise consistently, so one person maps to one pseudonym across the record and the log stays investigable.
Article 12 creates the capability. Articles 19 and 26(6) create the custody, and both hang on the word control.
A provider running a hosted service holds most of the record. A provider shipping self-hosted software holds almost none of it and owes something different instead: instructions that let the deployer interpret the logs. A format only your engineers can read is not a compliant deliverable.
If you only deploy someone else's system, Article 26 still reaches you. Deployers carry duties even though they did not build the system, a contract can allocate the work but not move the statutory duty, and ambiguous classification should be treated as high-risk, because under-scoping is the more expensive mistake. Sequencing matters here too, since which parts of the AI Act are already enforceable differs from what arrives with the high-risk chapter.
Having spent a section insisting that integrity is not mandated, here is the other half of the honest answer: build it anyway.
A conventional pipeline can capture every event Article 12 asks for and still be mutable. Anyone with write access can alter or truncate it, and nothing inside the record demonstrates that it is complete. That record can be vouched for. It cannot be verified. The distinction is invisible until the single moment it matters, when a market surveillance authority, a notified body or a counterparty asks whether the log is the whole log, and the answer is your word.
Hashes, signing and fail-closed handling are sensible engineering controls that Article 12 simply does not compel. Independent verification is the property worth aiming at: a regulator should be able to confirm the chain's integrity without trusting your assertion. Build it as good evidence design, chosen deliberately and documented as your own decision. Do not build it because a page told you the statute demanded it, and do not let a vendor sell it to you on that basis.
Every regulation that compels a capability without specifying the hard part does the same thing to a market. It defines a technical problem precisely, attaches a date to it, and applies it across the entire Union at once. That is an unusually clean brief for an inventor, which is why regulatory text belongs in a filing strategy and not only in a compliance folder.
EX-IX read Article 12 that way. The first family listed on IX Markets, Edge Assist, is a method for a tamper-proof black box recorder for AI systems, aimed squarely at the automatic logging and traceability requirement this article creates. Application P00202606645 is filed in Indonesia and pending examination, with dual US and EU prosecution scoped. Pending means pending: it has not been examined, grant is not assured, and nothing here should be read as a claim that it has issued.
The claim we do make is about method. Before any family reaches a buyer it passes a digital twin validation gate: physics-grade reconstruction of the claims, detectability analysis and validity probability, benchmarked at 0.76 percent mean absolute percentage error against real-world outcomes. That gate is what separates a filing aimed at a real regulatory gap from one aimed at a press release, and it is the same evidence that determines whether a filed but unexamined application is worth anything to a third party. If you are holding a filing of your own against a compliance deadline, how a patent is actually valued is the question to answer before the deadline arrives rather than after.
Does Article 12 require tamper-proof or append-only logs?
No. The Regulation's vocabulary is automatic recording of events, record-keeping and traceability, and it never uses the words audit trail or tamper-evident. Append-only storage and hash chaining are engineering recommendations that several widely read guides present as statutory. They are worth building. They are not compelled, and a vendor telling you otherwise is describing its product rather than the law.
Which specific fields does Article 12 make me log?
For almost every high-risk system, none by name. The only enumerated list is Article 12(3), which binds Annex III point 1(a) remote biometric identification systems. Everything else is derived from the intended purpose and the risk assessment, and the derivation is what you document.
Is there a harmonised standard for AI system logging yet?
Not yet in force. prEN ISO/IEC 24970 on AI system logging is in draft and is intended to carry the concrete implementation guidance, designed to be used alongside an Article 9 risk management system. Until it lands, no page can hand you a compliant schema, including this one.
Can platform and infrastructure logs satisfy Article 12 on their own?
Rarely. Gateway and platform records make routed interactions observable but cannot document application-level decisions taken outside that boundary: which model version produced the output, which policy fired, whether a human reviewed and overrode it. Those facts are not in your infrastructure telemetry.
Does Article 12 apply to us if we only deploy someone else's high-risk system?
The logging capability is the provider's duty under Article 12. Operational monitoring under Article 26(5) and retention of the logs under your control under Article 26(6) are yours. Neither obligation absorbs the other, and a contract can allocate the work without moving the duty.
We use essential cookies to operate this website and, with your consent, optional cookies to understand site usage. You can accept or decline non-essential cookies at any time. See our Impressum for our contact and legal details.