Skip to content

Evidence, testing and findings

NIST AI RMF outcomes are outcome-oriented. A defensible assessment therefore needs more than a policy statement or a Yes answer. It needs evidence that the relevant practice exists and works for the selected system.

An owner states that the practice exists.

Useful for discovery, but unverified.

Shows what should happen:

  • Policy.
  • Procedure.
  • Architecture.
  • Standard.
  • Role description.
  • Test plan.

Shows the practice is configured or established:

  • Approved system record.
  • Training completion.
  • Supplier assessment.
  • Monitoring configuration.
  • Access or oversight workflow.
  • Release approval.

Shows what actually happened:

  • Logs and alerts.
  • Completed reviews.
  • Test results.
  • Incident records.
  • Appeal and override records.
  • Monitoring trends.
  • Change and decommissioning records.

Adds challenge or validation:

  • Independent review.
  • Internal audit.
  • External assessment.
  • Reperformance.
  • Red-team or resilience exercise.

Ask whether evidence is:

  • Relevant — does it address this outcome?
  • Scoped — does it belong to this system?
  • Current — does it reflect the assessed configuration and period?
  • Authentic — is provenance clear?
  • Complete — does it cover the material system boundary?
  • Approved — has an authorised reviewer accepted it?
  • Consistent — does it agree with tests, findings and other records?
  • AI policy and risk-tolerance statements.
  • Inventory and ownership records.
  • Role and escalation matrices.
  • Training and competence records.
  • Stakeholder-engagement procedures.
  • Supplier-governance standards.
  • Review and decommissioning procedures.
  • System and model cards.
  • Intended-use and prohibited-use statements.
  • Architecture and data-flow diagrams.
  • Impact and risk assessments.
  • Stakeholder maps.
  • Benefits, costs and alternatives analysis.
  • System limitations and assumptions.
  • TEVV plans and reports.
  • Metric definitions and thresholds.
  • Evaluation datasets and representativeness analysis.
  • Safety, security, privacy and fairness evaluations.
  • Explainability and interpretability assessments.
  • Production-monitoring results.
  • Drift and incident trends.
  • Risk-treatment plans.
  • Proceed, restrict, stop or decommission decisions.
  • Risk acceptance.
  • Incident and recovery records.
  • Supplier contingencies.
  • Change records.
  • Continual-improvement actions.

An evidence request should state:

  • Outcome ID.
  • Selected system.
  • Requested artefact or operating record.
  • Why it is needed.
  • Owner.
  • Due date.
  • Acceptance criteria.
  • Review status.

Avoid “provide evidence of compliance.” Ask for the exact record needed.

Example:

For ME-2.4, provide the current production-monitoring specification, enabled alert rules, three months of monitoring results, threshold-change history and evidence of response to one material alert for the selected system.

Testing determines whether an implemented practice is effective.

Include:

  • Objective.
  • System and outcome.
  • Preconditions.
  • Environment and sample.
  • Procedure.
  • Expected result.
  • Pass criteria.
  • Safety limits.
  • Actual result.
  • Exceptions.
  • Reviewer and date.
  • Obtain authorisation.
  • Prefer non-production or controlled environments.
  • Use synthetic or minimised data where possible.
  • Define stop conditions.
  • Avoid uncontrolled harmful outputs or actions.
  • Protect affected people.
  • Preserve evidence.
  • Record deviations.
  • Retest after remediation.
  1. Select a sample of deployed AI services.
  2. Trace each to the inventory.
  3. Verify owner, purpose, version, lifecycle and review dates.
  4. Search procurement and technical records for unregistered AI.
  5. Pass only if the inventory is materially complete and current.
  1. Review documented limitations.
  2. Present representative edge cases.
  3. Confirm operators recognise uncertainty.
  4. Verify escalation or override.
  5. Pass only if limitations are accurate and oversight works in practice.
  1. Select authorised threat scenarios.
  2. Test boundary controls and monitoring.
  3. Verify safe failure and recovery.
  4. Confirm alerts are attributable and actionable.
  5. Record bypasses as findings.
  1. Run a tabletop or controlled stop exercise.
  2. Verify authority, communication and dependencies.
  3. Confirm the system can be disabled safely.
  4. Confirm fallback and recovery objectives.
  5. Pass only if the process works within the documented objective.

Raise a finding when:

  • Current outcome is overstated.
  • Evidence is missing or rejected.
  • A test fails.
  • Exceptions are uncontrolled.
  • Ownership is unclear.
  • Scope excludes a material component.
  • Supplier dependency is unresolved.
  • Monitoring does not detect material risk.
  • The Target Profile has no credible treatment path.

Record:

  • Outcome ID and system.
  • Condition observed.
  • Expected outcome.
  • Evidence and test basis.
  • Risk and affected parties.
  • Root cause where known.
  • Immediate containment.
  • Recommendation.
  • Owner and target date.
  • Status and closure evidence.

Positive narrative must not override:

  • Failed tests.
  • Rejected or expired evidence.
  • Open material findings.
  • Incidents.
  • Complaints or appeals.
  • Known model or supplier limitations.
  • Monitoring that contradicts the claimed outcome.

Evidence may support multiple frameworks when it genuinely addresses each objective. Reuse should preserve:

  • Original provenance.
  • System scope.
  • Review date.
  • Mapping rationale.
  • Framework-specific interpretation.

A crosswalk is not evidence by itself.

  • Evidence names the selected system.
  • Design and operation are distinguished.
  • Dates and versions are current.
  • Supplier evidence covers the relevant dependency.
  • Tests have defined pass criteria.
  • Failed tests are visible.
  • Findings are linked.
  • Monitoring and reassessment are defined.