Evidence, testing and findings
GTSAF separates four things that are often incorrectly combined:
- Assessment answer: the assessor’s Yes, No or N/A determination.
- Implementation narrative: the explanation of how the control is intended to work.
- Verified evidence: reviewed artefacts or operating records supporting the claim.
- Control test: a procedure used to determine whether the control works.
An assessment is defensible only when these elements agree.
Narrative is not evidence
Section titled “Narrative is not evidence”A detailed implementation narrative is valuable because it explains:
- Control design.
- Operating process.
- Ownership.
- Frequency.
- Exceptions.
- Shared responsibility.
It does not become verified evidence because it contains words such as “policy”, “reviewed”, “approved”, “tested” or “logged”.
The GTSAF screen deliberately keeps Narrative sufficiency separate from Verified evidence.
What good evidence looks like
Section titled “What good evidence looks like”Evidence should be:
- Relevant: directly supports the control requirement.
- Scoped: applies to the selected system, environment or shared control.
- Current: reflects the assessed period and current implementation.
- Authentic: comes from a credible source and has not been improperly altered.
- Complete: covers the material parts of the requirement.
- Traceable: linked to the control, system, owner and period.
- Reviewable: another assessor can understand what it proves.
- Consistent: does not contradict tests, findings or other records.
Evidence types
Section titled “Evidence types”Depending on the control, evidence may include:
- Policies and standards.
- Approved procedures.
- Architecture and data-flow diagrams.
- Configuration exports.
- Identity and access records.
- API gateway or runtime configuration.
- Model and dataset documentation.
- Data lineage and provenance records.
- Supplier contracts and assurance reports.
- Approval records and decision logs.
- Board or committee minutes.
- Monitoring dashboards.
- Alert and incident records.
- Change records.
- Test results.
- Access reviews.
- Exception registers.
- Training records.
- User notices and transparency material.
- Impact assessments.
- Recovery exercises.
- Red-team or adversarial test records.
The control’s Required evidence advisory identifies the expected starting set.
Evidence linkage
Section titled “Evidence linkage”Each record should be linked to:
- Workspace.
- Framework: GTSAF.
- Control ID.
- Selected AI system or intake.
- Evidence owner.
- Relevant review period.
System-scoped AI analysis uses operational records explicitly linked to the selected system or intake. Another system’s evidence is not silently reused.
If an enterprise-wide artefact applies to several systems, document that relationship rather than assuming universal applicability.
Evidence workflow
Section titled “Evidence workflow”Typical evidence states include:
| State | Meaning |
|---|---|
| Requested | Evidence has been requested but not received |
| Received | The artefact is available for review |
| Reviewed | An assessor has evaluated it |
| Accepted | It sufficiently supports the relevant requirement for the stated scope |
| Rejected | It does not support the claim or is materially deficient |
| Expired | It is no longer current for the assessment period |
Evidence quality may be recorded as:
- Poor
- Fair
- Good
- Strong
Status and quality are separate. A reviewed artefact may still be Poor.
Evidence review questions
Section titled “Evidence review questions”For every important artefact, ask:
- What precise requirement does this prove?
- Does it apply to the selected system?
- Is it current?
- Is it approved where approval matters?
- Does it show design, implementation or operation?
- What period does it cover?
- Who produced and reviewed it?
- Is the population complete?
- Does it contain exceptions?
- Does it contradict another record?
- Can another assessor reproduce the conclusion?
Evidence by effectiveness layer
Section titled “Evidence by effectiveness layer”Design evidence
Section titled “Design evidence”Shows that a control has been designed to address the risk:
- Policy.
- Standard.
- Architecture.
- Procedure.
- Control specification.
- Contractual requirement.
Implementation evidence
Section titled “Implementation evidence”Shows that the design has been put in place:
- Configuration.
- Role assignments.
- Deployed rule sets.
- Approved workflow.
- Training completion.
- System inventory.
- Supplier onboarding records.
Operating evidence
Section titled “Operating evidence”Shows that the control works over time:
- Logs.
- Review records.
- Alerts.
- Decisions.
- Incident handling.
- Samples of transactions.
- Periodic access reviews.
- Completed recovery exercises.
- Repeated monitoring results.
A policy alone usually supports design, not operating effectiveness.
Control testing
Section titled “Control testing”Testing answers:
Does the control operate as designed under defined conditions?
Each GTSAF control includes a specific test procedure. The assessor can adapt the sample and safety limits to the system without changing the control objective.
A complete test record
Section titled “A complete test record”Record:
- Test objective.
- Control and system.
- Preconditions.
- Population.
- Sampling method.
- Sample size.
- Test steps.
- Expected result.
- Pass criteria.
- Safety boundaries.
- Tester.
- Test date.
- Actual result.
- Exceptions.
- Evidence captured.
- Conclusion.
- Retest requirement.
Test-result labels
Section titled “Test-result labels”Common results include:
- Passed
- Pass
- Effective
- Failed
- Fail
- Ineffective
- Exception
- Pending
- Not tested
Passed, Pass and Effective are positive operating signals.
Failed, Fail, Ineffective and Exception are adverse signals. A failed test caps control assurance at 25% until the failure is resolved and retested.
Bounded and safe testing
Section titled “Bounded and safe testing”GTSAF testing must be proportionate and authorised.
Do not:
- Run destructive tests against production without explicit approval.
- Expose personal or sensitive data unnecessarily.
- Attempt privilege escalation outside an agreed test boundary.
- Trigger harmful autonomous actions.
- Send live customer communications.
- Modify production models or retrieval data without rollback.
- Use real secrets in test prompts.
Use:
- Test environments.
- Synthetic or masked data.
- Read-only inspection.
- Rate limits.
- Approval gates.
- Rollback plans.
- Monitoring during execution.
- Predefined stop conditions.
Sampling
Section titled “Sampling”Choose a sample that reflects:
- Risk and criticality.
- Transaction volume.
- Environment diversity.
- User or customer groups.
- Time period.
- Supplier or model variants.
- Known exceptions.
- Recent changes.
Critical and Enhanced controls normally require stronger sampling than Low or ordinary baseline controls.
Findings
Section titled “Findings”A finding is a governed record of a control weakness or assurance exception.
Raise a finding when:
- A requirement is No.
- A Gate fails.
- Evidence is missing, rejected or expired.
- A test fails.
- The control operates inconsistently.
- Ownership is unclear.
- A supplier obligation is not demonstrated.
- The system boundary differs from the documented design.
- Monitoring is insufficient.
- An assessor cannot support an effectiveness conclusion.
Finding anatomy
Section titled “Finding anatomy”| Field | Purpose |
|---|---|
| Title | Concise statement of the weakness |
| Condition | What was observed |
| Criteria | What GTSAF or the organisation requires |
| Evidence | Records supporting the finding |
| Cause | Why the weakness exists |
| Consequence | Risk created by the weakness |
| Severity | Priority based on impact and exposure |
| Recommendation | Required corrective outcome |
| Owner | Accountable remediation party |
| Target date | Agreed completion date |
| Status | Open, in progress, pending validation, resolved or closed |
Findings versus remediation
Section titled “Findings versus remediation”A finding states the assurance problem.
A remediation item tracks the work required to resolve it.
One finding may require several remediation actions. A remediation action should not be closed until the relevant control has been retested or otherwise validated.
Evidence and findings in the final conclusion
Section titled “Evidence and findings in the final conclusion”The assessor conclusion must reconcile:
- Positive evidence.
- Missing evidence.
- Rejected evidence.
- Passing tests.
- Failed tests.
- Open findings.
- Compensating controls.
- Residual risk.
Do not write “effective” while leaving an unexplained failed test or open Critical finding.
Example: policy versus operation
Section titled “Example: policy versus operation”Control requirement:
Privileged non-human identities used by an AI agent must be scoped, reviewed and revocable.
Possible records:
- Narrative: “The agent uses least privilege.”
- Design evidence: access-control standard and approved identity architecture.
- Implementation evidence: actual service-account role and permission export.
- Operating evidence: quarterly access review and revocation logs.
- Test: attempt an unauthorised action and confirm it is denied; revoke the identity and confirm subsequent calls fail.
The narrative alone is insufficient. The combined evidence and test support an operating conclusion.
Evidence checklist before Assured
Section titled “Evidence checklist before Assured”Before accepting Assured, confirm:
- No applicable No answers.
- The selected system is correct.
- Evidence is linked to the control and system.
- Evidence is reviewed and current.
- Evidence supports the claimed effectiveness layer.
- Critical controls have a passing test.
- No failed tests remain unexplained.
- Rejected evidence has not been counted positively.
- Open findings are reflected in residual risk.
- Ownership is determined.
- The assessor conclusion names limitations and review triggers.