VAERESOURCEData Engineering & Trusted AI
Trusted AI · Insights

Is Your Agency Actually Ready for AI? Check the Data First

Leadership wants an AI strategy, but nobody has checked whether the underlying data can support one. Here's how to run an honest AI readiness assessment before you spend a dollar on a model.

VAERESOURCE Insights·September 9, 2026·8 min read

Why AI Projects Really Fail in Government

When a government AI pilot stalls or gets shelved, the postmortem almost always points at the model: it was biased, it hallucinated, it didn't perform in production the way the demo suggested. That diagnosis is usually wrong, or at least incomplete. In our experience running data and AI work across federal, state, and local agencies, the model is rarely the root cause. The data is.

Agencies tend to buy AI readiness the way they buy software: pick a vendor, pick a use case, pick a launch date. But AI systems are only as good as the data pipeline feeding them, and most agency data was never built for this purpose. It was built for case management, for compliance reporting, for grant tracking — systems designed around forms and workflows, not around statistical reliability. Layering a model on top of that without an honest assessment is how you end up with a system that looks impressive in a sandbox and falls apart against real records.

An AI readiness assessment isn't a formality before the fun part starts. It's the work. Done properly, it will tell you whether you should be building a model at all right now, or whether the higher-value project is fixing your data foundation first.

If you can't explain where your data came from, who touched it, and what's wrong with it, you're not ready for AI — you're ready for a very expensive lesson.

Data Availability and Quality: The First Honest Check

Start with a blunt inventory question: do you actually have the data this use case needs, in a usable form, with known provenance? Not a data dictionary someone wrote three reorganizations ago — the actual current state of the actual tables and files.

This is where most readiness efforts get uncomfortable, because the answer is often no. Records live in three systems that don't talk to each other. Field definitions changed in 2019 and nobody updated the metadata. A fifth of records have missing values in the field the model needs most. None of this is unusual, and none of it is disqualifying by itself — but it has to be found and documented before a model touches the data, not after.

Governance and Security Before Anything Ships

Data quality tells you whether the AI can work. Governance tells you whether you're allowed to let it. These are separate questions and agencies routinely skip the second one until a privacy officer or IG asks about it after the fact.

NIST's AI Risk Management Framework (AI RMF 1.0) is a useful structure here because it forces you to name who owns risk decisions at each stage — Govern, Map, Measure, Manage — instead of assuming risk management happens somewhere in the background. If your agency handles federal controlled unclassified information, NIST 800-171 controls around access management, audit logging, and data protection apply to the AI pipeline just as much as they apply to any other system touching that data. If you're in a state or local context with HUD program data, CCPA-covered personal information, or similar regimes, those obligations don't pause because a system is labeled 'AI' instead of 'database.'

Concretely: Who approves what data can be used to train or fine-tune a model? Who can see model outputs before a decision gets made on them? What's the retention and deletion policy for training data and for logs of model use? If these questions don't have named owners and written answers, the project isn't ready, regardless of how good the data looks.

Use-Case Fit: Not Everything Should Be a Model

A readiness assessment has to include the option of a no. Some of the best outcomes we've delivered involved telling a client that the use case they wanted didn't need AI at all — a rules engine, a better dashboard, or a data-cleanup project solved the actual problem faster and with far less risk.

The use cases where AI genuinely earns its keep in government tend to share traits: high volume, repetitive pattern recognition, a tolerance for probabilistic output, and — critically — a human decision-maker downstream who has time and authority to review the output before it affects someone's benefits, eligibility, or liberty. Use cases that fail this test are usually ones where the stakes are high, the data is thin, and someone wants the system to make the final call unsupervised. That combination is where agency AI projects get agencies sued or investigated.

Designing Human Oversight Before You Build, Not After

Human-in-the-loop is one of those phrases that gets written into every AI strategy document and then quietly ignored during implementation because it's inconvenient. Real oversight design means specifying, before the model ships, exactly what a human reviewer sees, how much time they realistically have to review it, what happens when they disagree with the model, and how those disagreements get tracked and fed back into evaluation.

It also means designing for fail-closed behavior: when the model is uncertain, or when input data falls outside the range it was trained on, the system should default to routing to a human, not to producing a confident-looking answer anyway. This is a design decision, not an afterthought, and it has to be built into the workflow and the interface, not left as an unenforced policy in a PDF.

This is the part of the readiness assessment that separates agencies with a real AI strategy from agencies with a slide deck. The technical work of building oversight — audit trails, override logging, decision documentation tied to specific model versions — is exactly the kind of unglamorous data engineering that determines whether a system survives an IG audit or a FOIA request.

What an Honest Readiness Score Looks Like

Put together, a real AI readiness assessment produces a plain answer across four dimensions — data quality, governance, use-case fit, and oversight design — and it's fine, even expected, for an agency to score well on some and poorly on others. The point isn't a pass/fail grade; it's a prioritized list of what has to get fixed before deployment, and what can proceed now.

At VAERESOURCE, this is how we scope every AI engagement: we assess the data foundation first, map it against NIST AI RMF categories and applicable security and privacy controls, and only then talk about model selection. We build systems that are auditable by design — every data source traceable, every model decision logged, every override captured — and fail-closed by default, so uncertainty routes to a person instead of getting papered over. That approach takes longer than jumping straight to a pilot. It's also the reason the systems we build are still running, and still trusted, well after the demo is over.

Filed under: Trusted AI · Data Governance · AI Readiness · Public Sector IT · NIST AI RMF

Building AI or data systems your agency can trust?

VAERESOURCE is an SBA-certified SDVOSB/VOSB/WOSB data-engineering and trusted-AI firm for federal, state, and local missions. See our services.

Start a conversation →