Your AI programme has a vocabulary problem
Four frameworks, four different jobs. Most organisations pick one, use it for everything, and discover the gap in a conversation they cannot reschedule.
By Hector Gonzalez-Stahl · · 9 min
Nothing here is legal advice
That sentence is not a disclaimer bolted on at the end. It is the frame for the whole piece. This article will not tell you whether a regime applies to you or what compliance requires, and anyone who tells you that in an article is selling something. Read the originals, and involve counsel where the answer has consequences.
What this will do is something more basic and, in our experience, more often missing: explain what each of these four things is for, so you can tell which conversation you are in.
The failure we see is not organisations breaking rules. It is organisations picking whichever framework they heard about first, applying it to every question, and being surprised when it does not answer the one in front of them. A model risk framework will not help you tell a customer you are trustworthy. A management standard will not tell you whether your use case is prohibited.
Model risk management, and SR 11-7
Referenced in US banking through supervisory guidance issued in 2011, model risk management is the oldest discipline of the four and the one with real institutional teeth. Its core requirements are unremarkable and demanding: models must be documented, independently validated by someone who did not build them, and monitored in production.
Its audience is a prudential regulator. Its question is whether this institution can demonstrate that it understands and controls the models it relies on.
Generative AI strains it in a specific way worth understanding. The discipline assumes a model is a definable artifact with inputs, a method, and outputs you can test. With a language application, the thing that determines behaviour is not just the weights. It is the prompt, the retrieval corpus, the tool definitions, and the configuration, any of which can change on a Tuesday without anyone calling it a model change. If your change control covers the model and not the prompt, you have a documented model and an undocumented system.
The practical implication: the unit you validate has to be the whole configured system, and you need a record of when any part of it changed.
The NIST AI Risk Management Framework
Voluntary, published by the US National Institute of Standards and Technology, and organised around four functions: Govern, Map, Measure, Manage. No enforcement mechanism, no certification, no regulator behind it.
That makes it sound weak, and its actual value is real but different from the others. It is shared vocabulary. When your risk team, your engineers, and your business sponsor disagree about an AI system, they usually disagree because they are answering different questions without noticing. This framework gives them the same four boxes to disagree inside, which converts an unproductive argument into a locatable one.
Its audience is your own organisation. Its question is whether the people responsible for this system are reasoning about it in a structured, comparable way.
Use it to organise an internal debate. Do not present it to a regulator as evidence of compliance, because it is not addressed to them.
The EU AI Act
The only one of the four that is law. It classifies uses by risk, from prohibited through high-risk to limited and minimal, with obligations that scale by tier. It reaches organisations outside the EU based on where a system is placed on the market or where its output is used, so it matters to plenty of firms that do not consider themselves European.
Its audience is a regulator with enforcement powers. Its question is what tier this specific use sits in, and whether the obligations for that tier have been met.
The structural point that catches people: the tier attaches to the use, not to the technology. The same underlying model can be minimal risk in one application and high risk in another, which means the classification exercise has to be repeated per use case rather than performed once for the platform. Organisations that assess their model and consider the question closed have assessed the wrong unit.
It also phases in over several years, which produces a particular hazard: teams check the date, conclude an obligation is not live yet, and build something that will need rework when it is.
ISO/IEC 42001
A management system standard for AI, published in late 2023, and the certifiable one. It follows the familiar pattern of ISO management standards: define a system, run it, audit it, improve it, and get a third party to attest that you do.
Its audience is a customer, a partner, or a procurement process. Its question is whether an independent party will vouch that you run AI governance as a managed system rather than as a set of good intentions.
This is the one to reach for when the question arriving from outside is show me you are trustworthy, rather than prove you complied. If your sales cycle keeps stalling on a security and governance questionnaire, this is usually the conversation you are actually in, and no amount of internal framework adoption will substitute, because the point is precisely that someone external attested.
How they fit together
They overlap heavily in content and they are not substitutes, because they differ in audience and in what they produce.
A large regulated organisation typically needs the model risk discipline for its regulator, the NIST vocabulary to make its internal debate productive, the EU tiering to know which obligations attach to which use case, and the ISO standard when a customer asks for proof. Those are four different deliverables for four different readers, and satisfying one does not satisfy the others.
The good news is that the underlying work is largely shared. Documentation, independent validation, monitoring, incident response, and change control appear in all four wearing different names. If you build those capabilities once and keep the evidence, you are mostly re-presenting the same substance to different audiences rather than doing the work four times.
The bad news is that the evidence has to be a byproduct of operating the system rather than a document assembled for each audience in turn. A control whose artifact is produced only when someone asks is a control that will eventually be asked about at a moment when nobody has time to assemble it.
Where to start if this is new
Start with an inventory, not a framework. List the AI uses in the organisation, at the level of use case rather than technology, and for each one record what it decides or influences, whose data it touches, and who would be harmed if it were wrong. That list is the input every one of these frameworks needs, and most organisations cannot produce it.
Then classify by consequence. A drafting assistant for internal documents and a system that influences a credit decision are not the same governance problem and should not receive the same treatment. The most common failure is uniform governance: heavy process applied everywhere, which is expensive where it is unnecessary and, because it is resented, quietly bypassed where it matters.
Only then pick the framework, and pick it by the audience you actually have to satisfy first. If that is a regulator, start with model risk. If it is your own arguing teams, start with NIST. If it is a customer, start with ISO. If you place systems in the EU, the tiering exercise is not optional and the answer determines the rest.
The vocabulary matters because these conversations move faster than the practice does. Knowing which of the four you are in is most of knowing what a good answer sounds like.
- The four frameworks answer different questions for different audiences: model risk for a regulator, NIST for your own internal debate, the EU AI Act for legal obligation, ISO/IEC 42001 for a customer asking for proof.
- In the EU AI Act the risk tier attaches to the use, not to the technology. The same model can be minimal risk in one application and high risk in another, so classification repeats per use case.
- Generative AI strains model risk management because the thing that determines behaviour is the prompt, retrieval corpus, and configuration as well as the weights. Change control that covers only the model leaves the system undocumented.
- NIST AI RMF has no enforcement and real value: it gives disagreeing teams the same four boxes to disagree inside. Do not present it to a regulator as evidence of compliance.
- ISO/IEC 42001 is the one that answers show me you are trustworthy rather than prove you complied, because a third party attests.
- The underlying work is largely shared across all four. Build documentation, validation, monitoring, and change control once and keep the evidence as a byproduct of operating.
- Start with an inventory of use cases and their consequences, not with a framework. Uniform governance is expensive where it is unnecessary and bypassed where it matters.
Sources
- This article describes the purpose and audience of four publicly published frameworks. It is not legal advice, no regime is interpreted, and no compliance determination is offered.
- Framework descriptions reflect their publicly stated scope and structure. Readers should consult the primary documents and appropriate counsel for any question with consequences.
- No client system, employer system, or third-party engagement is described or alluded to.
Let's find out what your operation is actually running on.
Bring us the process you're trying to fix. We'll tell you honestly whether it's ready for automation or still needs to be standardized first.