Worked example
This example demonstrates method and terminology. It is not a template conclusion for every customer-support system.
System
Section titled “System”Customer Support Copilot
- Drafts responses for support agents.
- Retrieves approved knowledge-base articles.
- Uses a third-party large language model.
- Can summarise customer account history.
- Cannot send a response without agent approval.
- Processes personal data.
- Operates in production.
The assessment includes:
- Application.
- Foundation-model service.
- Retrieval layer.
- Knowledge base.
- Prompt and response controls.
- Human approval workflow.
- Monitoring and incident process.
It excludes automated account changes because the copilot has no action-capable tool for them.
Active Profile lenses
Section titled “Active Profile lenses”- NIST AI RMF 1.0 Core.
- NIST AI 600-1 Generative AI Profile.
Priority drivers include:
- Personal data.
- Public/customer-facing output.
- Third-party model dependency.
- Retrieval.
- Generative output.
GOVERN example: GV-6.1
Section titled “GOVERN example: GV-6.1”Outcome focus: third-party AI risk policies.
Evidence:
- Supplier due-diligence record.
- Contract and data-processing terms.
- Model/service documentation.
- Change-notification clause.
- Incident-escalation route.
Gap:
- No documented process for evaluating a major foundation-model version change before production use.
Current Profile:
Partially achieved
Assurance:
Implemented
Target:
Achieved
Treatment:
- Require advance supplier notification where available.
- Evaluate material model changes before release.
- Maintain fallback and rollback arrangements.
MAP example: MP-2.2
Section titled “MAP example: MP-2.2”Outcome focus: knowledge limits and human oversight.
Evidence:
- System card describing limitations.
- Agent training.
- Mandatory approval workflow.
- Escalation procedure.
Test:
- Present unsupported product and policy questions.
- Confirm the copilot indicates uncertainty.
- Confirm source articles are visible.
- Confirm the agent can reject the draft.
- Confirm high-risk topics route to a specialist.
Result:
- Unsupported answers sometimes appear confident.
- Human rejection works.
- Escalation does not consistently trigger.
Current Profile:
Partially achieved
Finding:
Confidence presentation and high-risk escalation do not reliably reflect knowledge limits.
MEASURE example: ME-2.10
Section titled “MEASURE example: ME-2.10”Outcome focus: privacy risk examination.
Evidence:
- Data-protection impact assessment.
- Data-flow diagram.
- Prompt logging rules.
- Retention configuration.
- Access review.
Test:
- Use synthetic account histories containing unnecessary sensitive details.
- Verify the retrieval layer minimises context.
- Check whether output repeats irrelevant personal data.
- Verify logs do not retain disallowed content.
- Check deletion and access controls.
Initial result:
Failed — the retrieval layer passes complete account notes when only a small subset is needed.
Current Profile:
Not achieved
Treatment:
- Field-level retrieval controls.
- Context minimisation.
- Sensitive-data filtering.
- Retest before wider rollout.
MEASURE example: ME-2.5
Section titled “MEASURE example: ME-2.5”Outcome focus: valid and reliable demonstration.
Evaluation set:
- Common enquiries.
- Rare product issues.
- Policy questions.
- Ambiguous requests.
- Adversarial or misleading prompts.
- Different customer language patterns.
Metrics:
- Factual correctness.
- Source support.
- Appropriate abstention.
- Escalation accuracy.
- Material-error rate.
- Performance by enquiry type.
Result:
Average performance is strong, but policy questions have a materially higher error rate.
Current Profile:
Partially achieved
The assessor does not mark Achieved merely because the average score is high.
MANAGE example: MG-4.1
Section titled “MANAGE example: MG-4.1”Outcome focus: post-deployment monitoring.
Monitoring covers:
- Unsupported-answer rate.
- Agent rejection rate.
- Escalation rate.
- Privacy-filter events.
- Supplier model version.
- Customer complaints.
- Incident severity.
The plan also includes:
- User and agent feedback.
- Override.
- Deactivation.
- Incident response.
- Recovery.
- Change management.
Test:
A tabletop simulates a supplier update that increases unsupported policy answers.
Initial result:
- Detection works.
- Deactivation authority is unclear.
- Rollback time exceeds the intended objective.
Current Profile:
Partially achieved
Treatment:
- Name stop authority.
- Pre-approve rollback criteria.
- Maintain a tested fallback model or non-AI workflow.
Generative-AI risks
Section titled “Generative-AI risks”The assessment gives specific attention to:
- Confabulation.
- Data privacy.
- Human-AI configuration.
- Information integrity.
- Information security.
- Intellectual property.
- Value-chain and component integration.
Other NIST AI 600-1 risk families are considered and documented according to their relevance.
Current Profile summary
Section titled “Current Profile summary”Strengths
Section titled “Strengths”- System is inventoried and owned.
- Human approval exists.
- Supplier due diligence exists.
- Evaluation and monitoring have begun.
- Incident and rollback concepts are documented.
Material gaps
Section titled “Material gaps”- Privacy minimisation failure.
- Weak high-risk escalation.
- Policy-question reliability.
- Supplier-change assurance.
- Unclear deactivation authority.
- Slow rollback.
Target Profile
Section titled “Target Profile”The Target Profile sets Achieved for the material outcomes above, supported by:
- Passed privacy retest.
- Improved abstention and escalation.
- Risk-segmented performance thresholds.
- Supplier-change review.
- Named stop authority.
- Tested rollback and non-AI fallback.
Conclusion
Section titled “Conclusion”Customer Support Copilot has a partially achieved Current Profile. Foundational governance, ownership, human approval, supplier review and monitoring are present, but material gaps remain in privacy minimisation, knowledge-limit handling, policy-response reliability and recovery. Wider deployment remains conditional on remediation and successful retesting. A supplier model change, new action capability, new sensitive-data source, material incident or significant performance drift requires reassessment.
What the example demonstrates
Section titled “What the example demonstrates”- System scope comes before outcome assessment.
- Core and Generative AI Profile work together.
- Strong averages can conceal material risk.
- Human approval alone does not solve reliability or privacy.
- A documented practice may still be only partially achieved.
- Failed tests constrain the conclusion.
- Supplier risk remains part of the system Profile.
- The Target Profile drives treatment.
- Human judgement owns the conclusion.