The U.S. government faces a critical gap in its ability to oversee artificial intelligence development, as national security experts warn that federal agencies lack the necessary tools to verify the performance claims made by frontier AI labs. A new survey conducted by the Institute for Security and Technology (IST) reveals that chronic underinvestment in independent testing and evaluation has left policymakers reliant on industry-provided data rather than objective, government-led benchmarks. This dependency raises significant concerns about the integration of AI into military and critical infrastructure systems, where the inability to validate model behavior before deployment could lead to unforeseen operational risks.
This assessment of the federal government's oversight capacity arrives as the Cybersecurity and Infrastructure Security Agency (CISA) faces its own internal challenges, including a significant reduction in its workforce and the recent termination of several free cybersecurity assessment services for critical infrastructure operators. The convergence of these developments highlights a broader struggle within the federal government to maintain a robust, independent security posture in an era of rapid technological advancement. As AI capabilities evolve, the lack of a standardized, government-verified testing framework threatens to leave the nation’s most sensitive systems vulnerable to both technical failures and adversarial exploitation.
The Verification Gap in Frontier AI
The IST survey, which gathered insights from 111 national security practitioners between late April and mid-July 2026, identifies the absence of independent evaluation processes as a primary impediment to secure AI adoption. According to the report, the government currently lacks reliable methods to benchmark frontier AI capabilities, forcing leaders to accept capability claims on faith. This reliance on industry self-reporting creates a vulnerability to regulatory capture, where the government’s understanding of AI risks is shaped primarily by the companies it is tasked with overseeing.
Experts emphasized that the risks associated with AI are not limited to the technology’s unique capabilities but are exacerbated by fundamental cybersecurity weaknesses. The survey noted that many practitioners view hallucinations as an intrinsic feature of current models, making them highly problematic for independent decision-making. As one respondent observed, AI can make a good analysis better, but it can also make a bad analysis significantly worse. Without rigorous, independent testing, the integration of these systems into critical infrastructure remains a high-stakes gamble.
CISA’s Strategic Retreat from Infrastructure Support
Compounding the oversight challenges is the recent decision by CISA to discontinue six legacy cybersecurity assessment services, including the Cyber Resilience Review and the Ransomware Readiness Assessment. These services previously provided infrastructure operators with direct, hands-on guidance to identify and mitigate vulnerabilities. CISA officials stated that the decision was driven by a need to reduce redundancy and focus on newer frameworks, such as the Cross-Sector Cybersecurity Performance Goals (CPGs). However, industry observers and former agency officials have criticized the move, arguing that the CPGs do not provide the same level of interactive, standards-focused evaluation as the retired assessments.
Internal documents and reports suggest that the agency’s capacity to support its partners has been severely hampered by a loss of roughly one-third of its workforce since the beginning of the second Trump administration. The elimination of these assessments is viewed by many as a retreat from the agency’s core mission of providing direct, actionable assistance to critical infrastructure owners. Critics warn that this reduction in support, combined with the lack of independent AI verification, will likely increase the nation’s overall cyber risk profile.
The Escalating Threat Landscape
While policymakers debate the best approach to regulation—with respondents split between federal requirements, sector-specific mandates, and hybrid models—the threat landscape continues to evolve. More than eight in 10 survey respondents agreed that AI would likely improve cyber defense capabilities, such as vulnerability discovery and incident response. Yet, 57% of those surveyed cautioned that AI is currently providing more utility to attackers than to defenders. This near-term imbalance underscores the urgency of establishing effective oversight mechanisms that can keep pace with the rapid deployment of AI-driven tools.
Experts cautioned against focusing exclusively on the exotic risks of AI, such as loss-of-control scenarios, at the expense of addressing basic cybersecurity failures. The IST report highlighted that weak cybersecurity practices remain the most likely enabler for adversarial manipulation and the theft of sensitive model weights or weapons designs. The consensus among practitioners is that the primary bottleneck is not the technology itself, but the ability to deploy robust, scalable cyber defenses that can withstand the increased sophistication of modern threats.
Future Implications for National Security
The termination of CISA’s legacy assessments and the ongoing struggle to verify AI capabilities present a complex challenge for the future of national security. As the government continues to navigate the integration of AI, the lack of independent benchmarks and the erosion of direct support services for infrastructure operators create a widening gap in the nation’s defensive architecture. The path forward remains uncertain, with experts calling for a renewed focus on independent evaluation and the restoration of critical support functions that have been lost to budget and workforce constraints.
Unresolved questions persist regarding how the government will reconcile its reliance on industry expertise with the need for objective, third-party verification. As the debate over regulatory frameworks continues, the ability of federal agencies to adapt to these challenges will be a defining factor in the security of the nation’s critical infrastructure. The upcoming months will likely see increased pressure on policymakers to address these gaps, as the risks associated with unverified AI and weakened cyber defenses become increasingly difficult to ignore.