Why most agency AI reviews fail before they start
Most agencies that try to govern AI adoption end up with one of two failure modes. Either every tool gets waved through because nobody owns the decision, or every tool gets stuck in a compliance queue because the review process has no scoring logic, just a checklist of yes/no questions that don't add up to a decision. Neither approach holds up when an inspector general, a state auditor, or a citizen complaint asks why a specific tool was approved.
The fix is not more policy language. It's a scoring method that produces a number, a rationale, and a paper trail, applied consistently whether the tool is a chatbot for constituent services, a resume screener for HR, or a fraud-detection model for benefits eligibility. NIST's AI Risk Management Framework (AI RMF 1.0) gives agencies the vocabulary, Govern, Map, Measure, Manage, but it deliberately does not hand you a scorecard. Building that scorecard is the work this article walks through.
The five-axis scoring model
Score every proposed AI use case on five axes, each on a 1-5 scale, before it touches production data. Keep the scoring sheet short enough that a program manager fills it out in under an hour, because a framework nobody uses is worse than no framework.
The goal isn't a single composite number that hides the details. It's five numbers plus a written rationale for each, reviewed by someone other than the person requesting the tool.
- Feasibility: Does the tool actually do what's claimed, on your data, at your scale? Score low if the vendor's demo used clean sample data unlike your production records, or if the use case requires accuracy the model can't realistically hit.
- Mission impact: What happens if it's wrong, and how often does 'wrong' occur? A tool that misroutes a help-desk ticket carries different stakes than one that flags a benefits application for fraud.
- Security posture: Does the tool meet your existing control baseline (NIST 800-53, 800-171 for CUI, FedRAMP for cloud services)? Where does data go, who can see it, and is there a model card or SSP-equivalent documentation?
- Legal and privacy exposure: Does the use case touch PII, HUD-protected fair-housing data, CCPA-covered personal information, or FOIA-discoverable records? Is there a documented lawful basis and a retention limit?
- Operational risk: What's the cost of reversing course, staff retraining, contract termination, reputational fallout, if the tool underperforms or gets pulled after deployment?
Scoring third-party and vendor AI specifically
Most AI risk in government doesn't come from models built in-house, it comes from vendor products layered with AI features nobody explicitly vetted. A records-management system that quietly adds an AI summarization feature in a routine update is a third-party AI risk event, even though no new contract was signed. Your framework needs a trigger for re-scoring when an existing vendor changes what their product does.
For any new AI vendor, add three questions to the standard five-axis score: Where does training data come from, and could your agency's inputs be used to improve the vendor's model for other customers? What does the contract say about liability if the tool causes a wrongful denial, a data breach, or a discriminatory outcome? And does the vendor provide audit logs and explainability sufficient to defend a decision in an appeal or a court challenge?
Contract risk-transfer matters here as much as technical review. Indemnification clauses, data ownership terms, and SLA-backed uptime and accuracy commitments should shift financial exposure toward the vendor, not just the agency's legal boilerplate. If a vendor won't commit contractually to data handling terms consistent with your security baseline, that's a finding, not a negotiating footnote to work around later.
Turning scores into a decision
Set thresholds in advance, before you score anything, so the process can't be bent to fit a preferred outcome. A common structure: any axis scoring 4 or 5 (high risk) triggers mandatory mitigation before deployment, regardless of how well the other axes scored. Two or more axes at 4-5 means the use case goes to a governance board, not a program-manager sign-off.
Document the rationale for every score in plain language, not just the number. When a state auditor or an IG asks why a facial-recognition pilot was approved for one agency but rejected for another, 'the scoring sheet said so' isn't a defense, the written reasoning behind each score is. This is also where fail-closed design earns its keep: if a tool's legal or security score can't be verified, the default answer is no, not a conditional yes pending paperwork that never arrives.
Re-score on a schedule, not just at initial approval. A tool that scored acceptably at launch can drift as vendors update models, as your data volume grows, or as new regulatory guidance (state AI laws, updated OMB memos, agency-specific policy) changes what's permissible. Build re-scoring into contract renewal cycles so it isn't optional.
Building this into how your agency actually works
A framework only holds up if it's used consistently across procurement, IT security, legal, and program offices, which usually means someone has to own the scoring sheet, the thresholds, and the re-scoring calendar, and enforce it even when a director wants to move fast. That ownership question is where most agency AI governance efforts stall, not on the scoring logic itself.
This is the same discipline VAERESOURCE applies when we build or evaluate AI systems for government clients: every tool gets scored against feasibility, impact, security, legal, and vendor risk before it goes near production data, every AI-assisted decision keeps a human in the loop with the authority to override it, and every system is designed to fail closed, denying by default when a check can't be verified, rather than quietly proceeding. Auditability isn't an add-on feature we bolt on for compliance; it's how the system is built from the first line of code, so that when someone asks why a decision was made, there's a real answer on file, not a reconstruction after the fact.
Building AI or data systems your agency can trust?
VAERESOURCE is an SBA-certified SDVOSB/VOSB/WOSB data-engineering and trusted-AI firm for federal, state, and local missions. See our services.
Start a conversation →