Evidence, testing and findings
ACRS is a capability-risk classification, but a defensible classification still needs evidence. The evidence should demonstrate the actual boundary used to select each dimension level.
Narrative is not evidence
Section titled “Narrative is not evidence”An assessor rationale explains judgement. It does not automatically prove:
- Effective permissions.
- Human approval.
- Fallback capacity.
- Kill-switch operation.
- Tenant isolation.
- Reversibility.
- Harm containment.
Gamut keeps rationale separate from objective evidence. Saving a rationale does not create an accepted or Strong evidence record.
System-linked records
Section titled “System-linked records”ACRS evidence records are linked to the selected scope through:
- Intake identifier.
- System identifier.
- System name.
- Framework.
- Dimension identifier.
This linkage applies to:
- Evidence.
- Evidence requests.
- Control tests.
- Findings.
When the system changes, only records belonging to the new scope should inform the assessment or AI analysis.
Evidence lifecycle
Section titled “Evidence lifecycle”A practical lifecycle is:
- Identify the claim the evidence must support.
- Create an evidence request where necessary.
- Collect or link the artefact.
- Verify system, environment and time-period relevance.
- Review quality and exceptions.
- Record the reviewed date.
- Accept, reject, expire or request correction.
- Reassess after material change.
Acceptance safeguards
Section titled “Acceptance safeguards”Gamut prevents an evidence request from being overstated as accepted, reviewed, closed or Strong unless it has:
- File or artefact references; and
- A reviewed date.
These are minimum integrity checks, not a complete quality judgement. The reviewer must still determine whether the artefact is authentic, current, complete, system-specific and sufficient.
Evidence quality criteria
Section titled “Evidence quality criteria”Assess:
| Criterion | Question |
|---|---|
| Scope | Does it relate to the selected system, environment, identity and workflow? |
| Authenticity | Is the source authoritative and protected from unauthorised change? |
| Currency | Does it cover the current deployed configuration and relevant period? |
| Completeness | Does it include the full path, population or exception set? |
| Independence | Was it produced or reviewed by an appropriately independent party? |
| Repeatability | Can the observed result be reproduced? |
| Traceability | Can the evidence be connected to a dimension claim, test or conclusion? |
| Exception visibility | Are failures, exclusions and limitations visible? |
Evidence by dimension
Section titled “Evidence by dimension”Operational Dependency
Section titled “Operational Dependency”Useful evidence:
- Business-impact analysis.
- Dependency maps.
- Fallback procedures.
- Capacity and staffing records.
- Recovery and substitution plans.
- Failover or continuity exercise results.
- Incident records.
Weak evidence:
- “The team can work manually.”
- Supplier SLA without fallback proof.
- Untested recovery documentation.
- A process map that omits integrations and data.
Action Autonomy
Section titled “Action Autonomy”Useful evidence:
- Deployed policy and approval configuration.
- Tool manifests.
- Approval, denial and action logs.
- Alternate-path and delegation tests.
- Kill-switch and rollback exercises.
- Enforced time, spend, step and retry limits.
Weak evidence:
- Prompt text saying “ask a human”.
- A screenshot of the intended workflow.
- Post-action notification.
- A policy document with no runtime enforcement.
Access Scope
Section titled “Access Scope”Useful evidence:
- Effective IAM and target-side policy.
- Token scope and lifetime.
- Credential-vault records.
- Denial and isolation tests.
- Data-flow and classification records.
- Access reviews.
- Attributable action traces.
Weak evidence:
- Role name without effective permissions.
- Tool schema treated as authorisation.
- An architecture diagram with no policy evidence.
- Intended least privilege with broad production credentials still active.
Harm Potential
Section titled “Harm Potential”Useful evidence:
- Impact, safety, rights and privacy assessments.
- Abuse cases.
- Incident and near-miss records.
- Affected-person analysis.
- Loss and service-disruption models.
- Appeal, redress, notification and recovery exercises.
- Independent challenge.
Weak evidence:
- Generic ethics principles.
- Accuracy benchmark alone.
- “No incidents have occurred.”
- Supplier assurance unrelated to the use context.
Testing principles
Section titled “Testing principles”Every test should define:
- Objective.
- Scope and environment.
- Preconditions.
- Authorisation.
- Test data and identities.
- Steps.
- Expected result.
- Pass criteria.
- Safety limits and stop conditions.
- Rollback or recovery.
- Actual result.
- Exceptions and evidence references.
- Tester and date.
Safe testing by dimension
Section titled “Safe testing by dimension”| Dimension | Preferred bounded test |
|---|---|
| Dependency | Tabletop, controlled failover, feature-flag disablement, shadow fallback or non-production restoration. |
| Action | Inert-tool action trace, approval-bypass test, retry/delegation test, kill-switch or rollback exercise. |
| Access | Synthetic-identity permission-path test, canary retrieval, target-side denial or tenant-isolation test. |
| Harm | Tabletop, simulation, synthetic population or bounded adverse-scenario exercise. |
Do not create real harmful outcomes merely to prove a risk. Tests involving production interruption, privilege, tenants, payments, code, safety, employment, health, legal position or public-service decisions require explicit authority and stronger safeguards.
Test result interpretation
Section titled “Test result interpretation”Passed
Section titled “Passed”A pass means the defined pass criteria were met in the tested scope. It does not prove:
- Every alternate path is safe.
- Future configurations remain safe.
- The dimension must be Low.
- Residual risk is accepted.
Failed
Section titled “Failed”A failed test should:
- Prevent reliance on the claimed boundary.
- Inform the dimension rationale.
- Create or update a finding.
- Identify immediate containment.
- Assign remediation and retest.
- Trigger route or confirmation review where material.
Partial or inconclusive
Section titled “Partial or inconclusive”Do not convert uncertainty into assurance. Record:
- What was tested.
- What remains unknown.
- Why the test was incomplete.
- The conservative scoring consequence.
- The next test or evidence request.
Findings
Section titled “Findings”Raise a finding when:
- A claimed Low or Medium boundary is not demonstrated.
- Effective capability exceeds intended capability.
- A contradiction cannot be resolved.
- A test fails.
- Evidence is missing, stale, rejected or system-mismatched.
- A severity-floor condition is ignored in governance decisions.
- The route is inconsistent with the deployed system.
- Monitoring cannot detect the assessed failure scenario.
- A material basis change has not been reassessed.
Finding content
Section titled “Finding content”A useful finding includes:
- Selected system and dimension.
- Observed condition.
- Expected boundary or control.
- Evidence and test references.
- Risk and credible consequence.
- Immediate containment.
- Root cause where known.
- Owner.
- Due date.
- Closure criteria.
- Retest requirement.
- Effect on ACRS confidence, confirmation or route.
Example
Section titled “Example”Action Autonomy: approval bypass through alternate tool path. The primary payment workflow requires approval, but the agent can invoke the settlement API directly through a secondary tool. The denial test failed. Until the alternate path is removed and retested, Action Autonomy remains High and the High ACRS route is authoritative. Disable the secondary tool in production immediately; closure requires policy-enforced approval on every settlement path and a passed bypass test.
Evidence and AI Assist
Section titled “Evidence and AI Assist”AI Assist may analyse the evidence metadata and statuses available to the assessment. It cannot declare an unsupported record accepted. Gamut limits accepted-evidence references to records that have actually been accepted or reviewed.
Recommendations, inferred facts and assessor narrative remain distinct from accepted evidence.
Completion checklist
Section titled “Completion checklist”- Evidence belongs to the selected system and dimension.
- Accepted records have artefact references and a reviewed date.
- Evidence supports the exact claimed boundary.
- Tests are authorised, bounded and reversible.
- Pass criteria were defined before interpreting the result.
- Failed and inconclusive tests affect the assessment.
- Findings include containment, owner, due date and closure evidence.
- Rationale is not presented as evidence.
- AI recommendations are not presented as evidence.