Someone on the team rewrites the system prompt, ships it, and feels like governance is done. I get the temptation. Prompts are visible. They feel like control. Then the vendor ships a new model name, behaviour shifts, and the carefully worded instructions stop doing what you thought they did.
Prompts are instructions to a model. Governance is how the organization decides who may use what, on which data, with what review, and how you prove it later. Those are different jobs.
Model upgrades can change tone, refusal patterns, tool-calling habits, and how strictly the model follows long instructions. Your “always cite the policy PDF” line might get softer. Your length limits might be ignored more often. Edge-case refusals might flip.
If the only control is the prompt text, you discover drift when a customer already saw the bad output. That is late.
Treat prompts as configuration, not policy. Configuration gets versioned and retested. Policy states outcomes and boundaries that stay true even if you swap models next quarter.
Write rules in the language of work, not model internals:
Keep the list short enough that a busy manager can explain it in five minutes. Long policy PDFs that nobody opens do not govern anything.
You do not need to store every token forever. You do need enough history to answer: what was asked, what context was attached, what came back, who approved it, and which version of the app and model was in play.
Useful minimums for many mid-market setups:
Without logs, “the AI did something weird” stays a story. With logs, it becomes a fixable defect. You also get a clearer picture of which features people actually use, which is useful when someone asks for more budget without evidence.
Define approval by risk of the outcome, not by which logo is on the API:
When a new model lands, re-run the same acceptance checks. Do not rewrite governance from scratch because marketing renamed the model. If you evaluate outputs with BA-style criteria, the checks travel better than “it felt smarter this week.”
A pattern that holds up:
When the vendor pushes a new default model, you retest layer three against layer one. You may tweak layer two. You should not invent a new philosophy of control every release.
Company knowledge retrieval needs the same discipline. A RAG setup can still surface stale or wrong documents if nobody owns content quality. Governance includes who updates the knowledge base, not only who wrote the system message.
If your control story is “we have a strong prompt,” you have a starting point, not a control framework. Add usage rules people can recite, logging you can query, and approval standards tied to business risk. Re-test when models change. Keep data ownership and access rules explicit so tools do not become a quiet second copy of the enterprise.
Hype cycles will keep renaming capabilities. Policies that describe work, risk, and accountability tend to survive the rename. Prompts alone usually do not.
I help teams put lightweight AI governance in place that still works after model upgrades. That can include:
Reach out for a quick chat on how I can help at Suganth@AruviConsultancyServices.com