Published on: August 14, 2026
External Forces & Industry Knowledge
Article
Actuarial Profession
Non-country specific
Emerging Topics Community Newsletter

Monitoring AI After Deployment: Lessons from a New NIST Framework

Author: R. Dale Hall

Actuaries spend a lot of time and effort in developing new modeling systems. Our profession laboriously checks and rechecks model parameters, ensuring good fit and explainability. That persistence for model validation carries over in the emerging practice of building new models and systems that incorporate artificial intelligence (AI), and the opportunity to adhere to a new and growing world of AI risk management and governance frameworks. Inevitably, however, new data emerges and new models are trained, tested and ready to be implemented. What additional considerations should actuaries consider and what challenges need to be overcome in order to release our “Model 2.0?”

Keeping AI on Track

The Society of Actuaries Research Institute, as part of our involvement with the Center for AI Standards and Innovation (CASAI) at the National Institute of Standards and Technology (NIST), has been reviewing and discussing these challenges, especially with a focus of how AI is evolving across the insurance and financial services industries. A new NIST report titled “Challenges to the Monitoring of Deployed AI Systems” highlights a critical shift in how organizations must think about AI: Not as a one-time model development exercise, but as an ongoing lifecycle requiring continuous evaluation, monitoring and refinement. While much of the AI field has historically focused on pre-deployment testing, the report emphasizes that real-world performance often diverges from controlled testing environments, making post-deployment monitoring essential to ensuring reliability, safety and effectiveness.

A central idea is that AI systems are inherently dynamic. Once deployed, they interact with changing data, evolving user behavior, and shifting environments. As a result, models can experience performance degradation, unintended behaviors, or “drift” over time. Monitoring is not simply about validation. It is increasingly more about detecting when a system no longer behaves as expected and needs updates, recalibration or retraining. This aligns closely with actuarial practice, where models are periodically refreshed as experience emerges and assumptions are revisited.

To structure this complex problem, NIST identifies six categories of post-deployment monitoring:

  1. Functionality
  2. Operational performance
  3. Human factors
  4. Security
  5. Compliance
  6. Large-scale impacts

These categories collectively recognize that AI risk extends beyond predictive accuracy. For example, a model may perform well technically but fail in terms of user interaction, regulatory compliance, or broader societal outcomes. This view mirrors our profession’s familiar enterprise risk management approaches, where risks are assessed across operational, strategic and reputational dimensions rather than viewed in isolation.

Why Monitoring Remains Difficult

A key contribution of the report is its framework for understanding gaps and barriers in current monitoring practices. Gaps are defined as underexplored or insufficiently developed areas, while barriers are practical obstacles that hinder effective monitoring. This distinction is particularly useful for practitioners as gaps point to where more innovation and research are needed, while barriers highlight constraints that must be managed in implementation.

Among the most significant gaps is the lack of standardized methods and metrics for monitoring AI systems. Organizations often struggle to determine what should be measured and how to interpret results. This is compounded by the fact that monitoring is highly dependent on the context in which the model has been implemented. Metrics that are appropriate in one application may not generalize to another. The absence of widely accepted standards creates inconsistency and limits comparability across systems.

Barriers further complicate this landscape. The report identifies several cross-cutting challenges, including limited visibility into model behavior, insufficient information sharing across stakeholders, and the rapid pace of technological change. For example, organizations may not have access to the underlying data or model components needed to fully understand system performance, particularly when relying on third-party AI providers. At the same time, the speed at which AI systems evolve makes it difficult to maintain up-to-date monitoring frameworks or documentation.

Resource constraints also play a major role. Effective monitoring can require significant investment in data infrastructure, computational resources, and skilled personnel. Human-in-the-loop processes, while valuable for validation, can be difficult to scale as systems grow. Often, organizations spend a lot of time and resources on initial implementation but don’t fully account for the resources needed for effective monitoring. These challenges create tension between the need for rigorous oversight and the practical realities of implementation.

Lessons for Model Lifecycle Management

At a more granular level, the report highlights category-specific issues that are especially relevant to model lifecycle management. In functionality monitoring, for instance, organizations face difficulties in establishing performance baselines and detecting drift over time. This is directly analogous to actuarial model monitoring, where identifying when experience deviates materially from expectations is a core task. Similarly, the lack of high-quality “ground truth” data in live environments makes it harder to assess model accuracy post-deployment.

Human factors monitoring presents particularly significant gaps. There is limited understanding of how users interact with AI systems, how their behavior changes over time, and how feedback loops between users and models might end up influencing outcomes. AI systems do not operate in isolation but are built, developed and implemented in ways where human behavior is both an input and an outcome.

The Governance Questions Ahead

The report also raises a set of open questions that organizations need to resolve, including who is responsible for monitoring, what should be monitored, and how frequently monitoring should occur. These questions underscore the lack of consensus and point to the need for governance frameworks that clearly define roles, responsibilities and processes.

Overall, the report reinforces the idea that post-deployment monitoring is not to be discounted and is a central component of creating trustworthy AI systems and models. It creates a feedback loop in which insights from real-world performance inform model updates, improve pre-deployment testing, and ultimately enhance system reliability.

Closing Observations for Actuarial and Risk Management Practice

First, this report provides a valuable conceptual bridge between traditional actuarial model governance and modern AI systems. Actuaries are already accustomed to model validation, experience monitoring, and periodic assumption updates. The NIST framework extends this idea into a broader environment and the need for actuaries to be aware of the important role they play in ongoing future monitoring of the systems they construct.

Second, the emphasis on gaps and barriers highlights an opportunity for the actuarial profession to contribute meaningfully to AI risk management. Actuaries bring expertise in measurement under uncertainty, model lifecycle management, and risk-based decision-making—skills that are directly applicable to addressing the challenges identified in the report. By adapting these capabilities to AI monitoring, insurers can strengthen governance, improve model reliability, and better manage emerging risks in an increasingly AI-driven landscape.

Stay tuned for further updates from the SOA Research Institute’s work with the CASAI at the NIST. You can also read the latest issue of the Actuarial Intelligence Bulletin and visit Artificial Intelligence Research to stay up to date on the developments of the technology in the actuarial profession.

This article is provided for informational and educational purposes only. Neither the Society of Actuaries nor the respective authors’ employers make any endorsement, representation or guarantee with regard to any content, and disclaim any liability in connection with the use or misuse of any information provided herein. This article should not be construed as professional or financial advice. Statements of fact and opinions expressed herein are those of the individual authors and are not necessarily those of the Society of Actuaries or the respective authors’ employers.


R. Dale Hall, FSA, MAAA, CERA, is managing director of research for SOA Research Institute. Dale can be contacted at dhall@soa.org.

Author: R. Dale Hall
Published on: August 14, 2026
External Forces & Industry Knowledge
Article
Actuarial Profession
Non-country specific
Emerging Topics Community Newsletter
Back to Top