Approach
Automation that can be checked
A method for building automation that can be explained, tested and traced, and for saying honestly what it can and cannot do.
Automation earns trust when a sceptical person can ask three questions and get a specific answer. What is it allowed to do? How do you know it is good enough? What happened in this case? Everything we build is designed around those three.
01
Protocols
What is it allowed to do?
Automate the effort, constrain the AI, escalate the interpretation to a person.
02
Proof
How do you know it is good enough?
Judged against expert golden sets with a proper statistical framework, and every figure reported with a confidence interval.
03
Provenance
What happened in this case?
An audit trail for everything, from raw data to finished output.
Protocols: what it is allowed to do
The principle is simple: automate the effort, constrain the AI's judgement, escalate the interpretation to a person. Every task an automation touches is one of three kinds, and each kind is handled differently.
- Deterministic steps: automate the effort. Cleaning, matching, calculating, formatting: work where there is a right answer and rules can find it. These are automated completely and tested like any other software. If a rule can do it, a model is not used.
- Routine judgement: constrain the AI. Categorising, extracting or summarising against a framework a person has already set. This is where AI earns its place, but on a short lead, as described below.
- Decisions about meaning: escalate to a person. What a theme is, what a finding implies, whether an exception applies, what to recommend. These go to a person, by design. The system's job is to put the decision in front of the right person, with the evidence they need.
How the constraint is done technically
- A closed set of permitted outputs. The model answers in a fixed structure against a fixed framework. It cannot invent a category, and it has no authority to change the framework.
- A way out. There is always a defined way to say this does not fit. An item that fits nothing is left unresolved and shown to a person, never forced into the nearest category.
- Voting. Where a judgement matters, the same item is judged several times, independently, and the agreement between the runs is the signal. Unanimous answers pass. Split answers are not averaged away: they are routed to a person, because disagreement between runs is the cheapest reliable sign that an item is hard.
- Adversarial testing. Before anything is relied on, we try to break it: cases built to be hard or ambiguous, inputs that try to redirect the model, and a separate checking process whose only job is to find what the first one got wrong. What survives goes into the test set permanently.
- Escalate by need. Human review starts with the items most likely to need judgement, such as split votes, low confidence, unusual or high consequence, so limited attention goes where it matters most.
Proof: how you know it is good enough
Where a tool involves judgement, we test it against work produced by experts in that field, using a proper statistical framework, and we do not call it ready until it matches that standard.
- Build a golden set first. The experts produce it by hand, before the tool sees the material. It is the standard the automation is judged against, and it is kept separate from anything used to tune the tool.
- Measure agreement properly. Chance-corrected statistics, per category as well as overall, not raw percentage agreement. Our note on validating an LLM that codes open ends shows why raw agreement flatters a model when categories are unbalanced.
- Report every figure with a confidence interval. No performance number appears without one. A result of 90% from fifty items and 90% from two thousand are very different claims, and the interval is what tells them apart. We plan the size of the golden set in advance so that the interval will be narrow enough to act on.
- Compare with the human baseline. Experts do not agree with each other perfectly either. The target is at least as good as the manual process, and the manual process is measured the same way.
- Say what was and was not tested. If a tool has not been validated yet, we say so and describe the checks that run on each piece of work instead. Results are shared openly, including the ones that are unflattering.
- Test again when something changes. A new model version, a reworded instruction or a different kind of data is a new thing to test.
Provenance: an audit trail for everything
Every output comes with a full audit trail from raw data to finished result, supplied as standard. It holds the inputs and their version, every transformation in order, every automated judgement (which rule or model, which version of the instructions, what it decided, how confident it was, how the votes fell), and every human decision (who, when, what they changed). Outputs link back so that any figure or quote can be traced to the rows or passages behind it.
An audit trail is written for someone who was not there, months later, with one question about one output. That is a different document from a system log, and it has to be designed. We wrote up what it needs to contain.
Governance runs through all three
Before work starts we agree, and write down:
- where data is processed and stored;
- who can access it, and how that access is limited;
- how long it is kept, and how it is deleted;
- whether any AI service is allowed to retain or learn from it (we choose services and settings so that it is not, and say which);
- whether personal details are removed before processing where they are not needed;
- which decisions need a named person's approval.
What this is not
Why a standard, not a one-off
When each automation gets its own bespoke controls, standards vary, each one has to be explained separately, and lessons learned on one are not applied to the next. Asking the same three questions of everything has three effects. Leaders get one story they can tell in a sentence. Checks are improved once and everything benefits. And there is one place to see what is running, what has been tested and what is not ready yet.
The trust sheet
Every automation we deliver gets a one-page trust sheet. It is written for the person who has to decide whether to rely on the thing, not the person who built it. The rule is simple: if something has not been validated yet, the sheet says so.
Template
Trust sheet: [automation name]
A worked example, for a tool we run ourselves: the trust sheet for AI Signal.
How we say what is ready
Each capability carries one of three statuses, and we use the same words in proposals and on trust sheets.
| Status | Meaning |
|---|---|
| Green | Tested against expert work, results available, safe to rely on within its stated limits. |
| Amber | Working and being validated. Used only with full human review or on low-consequence work, and we do not quote performance figures we have not measured. |
| Red | Not ready, or not something we would offer. We say so rather than stretch. |
How we talk about it
Claims should match the evidence and state their limits, so some phrases are off the table.
| We say | We avoid |
|---|---|
| Tested against expert work, with the result and what it does not cover | "Accurate", "100% reliable" |
| Every step is recorded and traceable | "Black box", "magic" |
| People make the decisions that matter | "AI does the analysis" |
| Faster on the routine steps | "Instant", "effortless" |
| "Not yet validated", when that is true | Saying nothing |