Evaluating Third-Party AI Use: How to Address Three Real-World Scenarios - Shared Assessments

Blogpost

Evaluating Third-Party AI Use: How to Address Three Real-World Scenarios

Third parties are rapidly integrating artificial intelligence (AI) into products, services, and business processes. Whether a legal team buys an AI-powered research tool, a software vendor uses generative AI to write code, or a critical supplier quietly adds AI capabilities to an existing platform, third-party risk management (TPRM) teams confront the same question: How do we evaluate the risk?

During a recent Best Practices Committee discussion with guest speaker Josh Harguess from Fire Mountain Labs, third-party risk practitioners worked through real-world situations that procurement analysts, contract managers, TPRM program analysts, cybersecurity analysts, and other professionals on the front lines of third-party risk contend with each day. The conversation revealed a few practical approaches and actions that help risk teams 1) evaluate AI more effectively and 2) engage business stakeholders earlier in the assessment process.

Scenario 1: Legal Wants to Purchase an AI Tool

Context: The legal department requests approval to use an AI-enabled platform that can answer legal questions, review contracts, and store documents. The solution is also positioned to support AI agents capable of performing tasks on a user’s behalf.

Response: While the discussion panel acknowledged the need to review the legal platform provider’s AI governance policies, model documentation, and certifications (e.g., ISO/IEC 42001), our experts asserted that policies and certifications are not a substitute for an independent assessment. Key actions in an effective assessment include:

  • Distinguish use cases: Rather than assessing the product as a whole (a natural inclination), start by examining individual use cases. Document storage, contract review, legal research, and autonomous task execution each present different risk profiles. Before evaluating the vendor, take time to understand exactly how the business intends to use the tool. The features being marketed by the vendor may not be the features the organization actually needs.
  • Deploy a phased approach: A phased approach reduces uncertainty. Rather than approving use of the entire platform, consider piloting a lower-risk function first. This gives users an opportunity to gain experience with the technology while providing risk teams with real-world performance data.
  • Prioritize data handling: Organizations should understand whether their information is used to train or fine-tune the model, whether customer data is segregated, and whether opt-out options are available. Most important, those representations and commitments should be addressed in the contract.
  • Validate output: For AI tools that generate work product, output validation is critical. An AI-powered legal research tool is only valuable if the information it produces is accurate and reliable. Testing outputs against known results routinely reveals issues that marketing materials and demonstrations neglect.
  • Establish ongoing monitoring: AI systems continually evolve. Model updates, retraining activities, and new capabilities change a tool’s risk profile over time. Evaluation should be ongoing, not a one-time approval exercise.

 

Scenario 2: A Software Development Vendor Uses AI to Write Code

Context: A software development vendor proposes using AI coding assistants to build and maintain a critical application.

Response: While the benefits of AI coding assistants (faster development and lower costs) are appealing, the difficulty lies in determining what risks accompany those gains. One risk is maintainability. If large portions of an application are generated by AI, the developers responsible for supporting the software may not fully understand the resulting code. That becomes problematic when systems fail, security vulnerabilities emerge, or enhancements are needed. Code quality also requires consideration. While AI coding tools can be useful accelerators, generated code is not automatically secure, efficient, or well-architected. Human expertise and involvement remain essential. To assess these types of risk when evaluating vendors that use AI-assisted development approaches, consider asking:

  • What portions of the application were generated with AI assistance?
  • Which components contain AI-generated code?
  • What human review processes are required before software is released?
  • How are security and quality testing performed?
  • Can the vendor provide a software bill of materials (SBOM)?
  • What knowledge transfer processes exist to ensure long-term support?

An important point: The goal here is not to discourage AI use. Rather, it is to ensure the vendor has controls in place that preserve software quality, maintainability, and accountability.

 

Scenario 3: A Critical Supplier Suddenly Adds AI Functionality

Context: A supplier updates its website, releases a new feature announcement, or adds a brief line in release notes stating that its product now includes AI functionality. No one in your organization is informed about this change, and the current contract with the vendor contains no AI-specific language.

Response: Start with a simple yet crucial question: Is it actually AI? Marketing claims about AI offerings often overstate the degree of actual autonomous decision-making (a practice known as “AI washing”). Ask suppliers to explain exactly what technology is being used and how it works. Vague answers should prompt additional scrutiny. If AI is involved, subsequent steps include:

  • Understand the exposure: A recommendation engine that suggests report formatting presents a substantially different level of exposure than a model that processes sensitive information or makes decisions that affect customers.
  • Temporarily disable the AI functionality: Determine whether the new AI functionality can be disabled while an assessment is completed. This provides valuable breathing room without disrupting business operations.
  • Assess broader AI governance practices: Ask who owns AI risk, how models are monitored, how performance is measured, and what controls exist to manage model changes over time.
  • Establish contractual protections: Vendor agreements negotiated before AI became a significant consideration rarely address notification requirements, model changes, data usage, or AI-specific responsibilities. As contracts are renewed, organizations should consider incorporating provisions that require suppliers to disclose significant AI-related changes before implementation.

 

Four Fundamental AI-Assessment Practices

Across the different scenarios the Best Practices Committee worked through, four fundamental AI-assessment practices repeatedly surfaced:

  1. Engage early: Risk teams are most effective when they are involved before a tool is selected or implemented. Early engagement allows risks to be evaluated alongside business benefits rather than after decisions have already been made.
  2. Focus on actual risk: Not every AI capability deserves the same level of scrutiny. Assess the use case, the data involved, and the potential impact. Avoid treating every AI feature as equally risky.
  3. Monitor continuously: AI systems change. Models are retrained, vendors introduce new functionality, and underlying technologies evolve. Ongoing monitoring is essential.
  4. Know when to enlist outside expertise: Asking the right questions is important. Understanding the answers is equally important. Organizations should recognize when specialized AI expertise is needed to support a thorough assessment.

 

Questions Every TPRM Team Should Ask

In addition to performing those actions, organizations should consider adding the following AI-focused questions to their assessment processes:

  • How does the vendor classify the use case under established frameworks such as the NIST AI Risk Management Framework or applicable regulatory requirements?
  • What level of human oversight exists?
  • What data was used to train the model?
  • How is bias tested and monitored?
  • How are model drift and performance degradation addressed?
  • For agentic AI, what actions can the system take and what controls limit those actions?
  • Has the solution been tested for prompt injection and other adversarial attacks?
  • What does AI-specific incident response look like?
  • Does the solution rely on another provider’s foundation model?
  • How will the vendor notify customers about model changes, retraining efforts, or major upgrades?

 

The Bottom Line: Ask the Right Questions at the Right Time

The challenge for third-party risk teams is no longer to determine when AI will appear in their vendor ecosystem; it already has. Now, the priority is to develop a practical approach to evaluating AI-enabled tools, understanding how they are used, and ensuring that governance evolves alongside technology.

The organizations that succeed won’t necessarily ask the most AI questions; they will be the ones that ask the right questions at the right time and then use the answers to make better risk decisions.

To learn more about Shared Assessments Committees, visit https://sharedassessments.org/committees/.