Ask an automated pipeline how it produced a particular figure and you tend to get one of two answers. The first is "trust us". The second is forty thousand lines of logs, which is the same answer with more effort.
Neither is an audit trail. The difference is not volume. It is who the record is written for.
Who the record is for
A log is written for the engineer who is debugging tonight. They know the system, they know roughly what went wrong, and they want the most recent events in the order they happened.
An audit trail is written for someone who was not there. An auditor, a client, a colleague who inherited the project, occasionally a court. They arrive months later, they do not know how the system works, and they have one question about one output: where did this come from, and who or what decided it?
Everything worth saying about a provenance record follows from taking that reader seriously.
What it has to contain
Five layers, each answering a different question.
| Layer | What is recorded | The question it answers |
|---|---|---|
| Inputs | Which dataset, which version (a checksum, not just a filename), when it arrived | What was this built from? |
| Transformations | Each step, in order, with its parameters | What was done to it? |
| Automated judgements | The rule or model, the version of its instructions, what it saw, what it decided, how confident it was, and how independent runs voted | What did the machine decide, and on what basis? |
| Human decisions | Who, when, what they were shown, what they changed, and why | Who decided? |
| Outputs | Each figure or quote linked back to the rows or passages behind it | Where does this number come from? |
The middle two rows are the ones most often missing. Logs of an automated step usually say that it ran. They rarely say what it decided for each item, and almost never record that a person looked at the result and changed it.
Six design rules
1. Never overwrite the source. A correction creates a new version, and the record notes the difference. When approved codes are merged back into a dataset, the merge produces a new dataset and leaves the original exactly as it arrived. If the original can be changed, nothing downstream can be trusted.
2. Record as a by-product of doing the work. A record reconstructed afterwards is a story about what happened. A record written by the process itself, at the moment each thing happens, is evidence. The first is cheap to produce and the second is the only one that survives a challenge.
3. Record the version of the instructions, not only the model. A prompt reworded in March and re-run in June is a different automation, even if the model name did not change. The instructions are part of the method and belong in the record.
4. Treat human decisions as first-class events. "Reviewed" is not a record. "Reviewed 200 items, changed 12, and here are the 12" is. Approvals that changed nothing matter as much as corrections, because they show the review happened.
5. Keep the disagreement. When several independent runs judge the same item and split, that split is information. Store how each run voted, not only the answer that won. It shows which items were hard, and it is the reason a person was asked.
6. Make it readable without the system. If the only way to interpret the record is to run the software that produced it, the record dies with the software. Export something a person can read cold.
A test you can run on your own pipeline
Pick a number from a deliverable at random. Time how long it takes someone who did not build the system to reach the underlying rows and see every automated and human step in between.
If the honest answer is "a day, and the original developer", you have logs. If it is "a few minutes, from the report itself", you have an audit trail.
Why this is becoming an expectation
Rules on marking and disclosing AI-generated content are moving provenance from good practice towards something organisations will be asked to show. The EU AI Act's transparency obligations under Article 50 apply from August 2026, and the technical means of marking output are still settling. Buyers in regulated and public settings were already asking how a result was produced. The difference now is that "we keep good records" is going to need to be something you can demonstrate.
The question that stays open
Who should hold the record? If the supplier holds it, the client is trusting the supplier's account of its own work. If the client holds it, they need to be able to read it, which returns us to rule six. And if the supplier changes, does the record travel with the work?
I do not have a settled answer, but I lean towards the client holding a copy that they can read without the supplier's help, because a record that depends on the party being audited is weaker than one that does not. It is worth deciding before the first output ships, not after the first challenge.