Skip to content

Assessments & control testing

An assessment evaluates its defined subject against a framework. Most subjects are AI systems; ISO/IEC 42001 uses an organisation-level AIMS scope and ATF uses a registered agent. Control testing goes a step further: it evidences that controls are not just designed, but operating effectively.

An assessment connects an AI system to a framework and works through its controls, domain by domain. For each control you record:

  • The current state, how well the control is met.
  • The rationale, why you reached that conclusion.
  • Evidence links, the proof that supports the conclusion.

Because an assessment records rationale and evidence per control, the result is defensible rather than just a score. See Run your first assessment for the steps.

The same AI system can be assessed against more than one framework, for example GTSAF for depth and the EU AI Act for regulatory readiness. Gamut keeps each assessment distinct while sharing the underlying system, evidence and findings, so work done once is reused across frameworks rather than repeated. See Frameworks overview for how routing selects frameworks.

  1. Make sure the AI system is registered and has been through intake and risk tiering, so it is routed to the right frameworks.
  2. Open the routed framework (for example GTSAF or the EU AI Act) for that system.
  3. Work through the controls, scoring each on the framework’s answer scale and recording your rationale. Each control’s auditor advisory explains what to assess and how to score it.
  4. Attach or request evidence, run control tests, and raise findings for gaps.
  5. Roll the result up with reporting.

Gamut applies a consistent scoring discipline across every framework, so an assessment score is honest about both quality and completeness.

  • Each framework uses its own answer scale. GTSAF uses Yes / No / N/A; ATF uses Fully met / Partially met / Not addressed; the EU AI Act and NAGF use Compliant / Non-compliant / N/A with a depth rating; ISO/IEC 42001 separates conformity from assurance; ISO/IEC 42005 uses maturity plus assurance; MAESTRO scores likelihood and impact. See Frameworks overview.
  • Scope-correct, with rollup. System assessments remain bound to one system, ATF to one agent and ISO/IEC 42001 to the approved AIMS scope. Portfolio views do not transfer evidence between those subjects.
  • Applicability is separate from depth. For GTSAF, system facts determine which conditional controls apply; ACRS and the governance weighting profile determine proportionate assessment rigour. A higher applicable count is not automatically a higher-risk conclusion.
  • Coverage-inclusive. Controls you have not assessed count as not-yet-met, so a partly-finished assessment cannot report a high score.
  • Justified N/A only. Marking an applicable control N/A requires a documented justification; a justified N/A leaves scope, while an unjustified one stays in scope and counts as not met.
  • Auditor advisory on every control. Each control carries built-in guidance on what to assess, what evidence to expect, how to test it and how to score it, so assessors apply the scale consistently.

Designing a control is not the same as operating it. A control test in Gamut captures the full testing record auditors expect:

  • Test objective and test procedure, what was tested and how.
  • Sample method and sample size, the basis for the conclusion.
  • A design effectiveness score and an operating effectiveness score, kept separate, because a well-designed control can still fail in operation.
  • A test result: pass, partial, fail or not_tested.
  • Exceptions: a count and a summary of what failed.
  • Tester and reviewer, test date and next test date, and evidence references.

This turns an assertion (“we review model outputs”) into proof (“here are the reviews, sampled this way, on this cadence, by these people, with these exceptions”). Control tests produce the artefacts that back an assessment and that auditors rely on when they test your governance.

Where an assessment or control test reveals a gap, raise a finding. A finding carries a severity, a root cause, a recommendation and a management response, and it links back to the control, the test and the system it relates to, then tracks through to remediation and closure. See Evidence & findings.

Gamut can assist assessment work with AI, for example by helping analyse context or draft narrative. All AI analysis is proxied server-side, so model provider keys are never exposed to the browser. See AI assistance & data handling.