Skip to content

MAESTRO

MAESTRO (Multi-Agent Environment, Security, Threat, Risk, and Outcome) is a threat-modelling framework designed for the distinctive risks of agentic AI. It examines an AI system layer by layer, then traces threats that propagate across layers, agents and trust boundaries.

Gamut implements the Cloud Security Alliance catalogue published on 6 February 2025 as CSA-2025-02-06-v1:

  • Seven canonical architectural layers.
  • A dedicated cross-layer threat section.
  • 50 layer-specific threats.
  • Five cross-layer threats.
  • Eight canonical agentic architecture patterns.
  • A six-step assessment method.
  • Likelihood × impact risk assessment.
  • A unique assessor playbook for every threat.
  • System-specific assessment records and workspace roll-up.
  • Evidence, tests, findings, treatment and monitoring records.
  • Validated, system-scoped AI assistance for each canonical threat.
If you need to…Read
Explain what MAESTRO is, how it differs from a control framework and what every label meansMethod, concepts and terminology
Understand all seven layers and every one of the 55 canonical threatsLayers and threat catalogue
Decompose a system and select the right agentic architecture patternSystem decomposition and architecture patterns
Complete an assessment and defend the scoringAssessment workflow and risk scoring
Know what evidence to request, how to test safely and how to treat findingsEvidence, testing and treatment
Explain what AI Assist reads, produces, secures and cannot decideAI Assist and security
Produce reports and connect MAESTRO to GTSAF, ACRS and ATFReporting and governance
Follow a complete realistic assessmentWorked example
Look up identifiers, scores, labels and formulasReference and glossary
PropertyGamut implementation
Catalogue versionCSA-2025-02-06-v1
Canonical layers7
Cross-layer section1
Layer-specific threats50
Cross-layer threats5
Total assessable threats55
Stable identifierMAE-L<layer>-<nn> or MAE-X-<nn>
Architecture patterns8
Workflow steps6
Risk dimensionsLikelihood 1–5 and impact 1–5
Raw riskLikelihood × impact, from 1 to 25
Severity bandsLow, Moderate, Elevated, High, Critical
Layer roll-upHighest scored threat in the layer
System roll-upHighest scored threat for the selected system
Workspace roll-upSystems assessed, worst case and distribution by worst-case band
SectionCanonical nameThreatsPrimary subject
Layer 1Foundation Models7Model behaviour, privacy, integrity, extraction and availability
Layer 2Data Operations5Data stores, pipelines, RAG, provenance, integrity and availability
Layer 3Agent Frameworks6Framework components, APIs, dependencies, validation and control evasion
Layer 4Deployment & Infrastructure6Images, orchestration, IaC, compute, networks and lateral movement
Layer 5Evaluation & Observability6Metrics, evaluators, telemetry, detection integrity and monitoring confidentiality
Layer 6Security & Compliance7The vertical security layer, especially AI agents performing security functions
Layer 7Agent Ecosystem13Agents, identities, tools, registries, discovery, markets and business integrations
Cross-layerCross-Layer Threats5Attack chains and cascading failures spanning two or more layers

Layer 6 is vertical: it cuts across the rest of the architecture. The cross-layer section is not an eighth architectural layer; it records threats that exploit relationships between layers.

  1. Select the AI system being assessed, or deliberately use workspace scope.
  2. Decompose its components, capabilities, tools, goals, constraints and interactions.
  3. Select the closest canonical architecture pattern.
  4. Confirm which assets and pathways exist in each layer.
  5. Tailor each applicable canonical threat into a system-specific scenario.
  6. Score likelihood and impact.
  7. Record current controls, treatment owner, target date, monitoring and reassessment triggers.
  8. Link evidence, tests and findings.
  9. Model cross-layer attack chains.
  10. Review the highest risks, generate reporting and obtain the authorised human risk decision.

A defensible threat record should answer:

  • Why is the threat applicable to this system?
  • Which actor, capability or failure can initiate it?
  • Which assets and trust boundaries form the attack path?
  • What is the credible business, safety, privacy or mission outcome?
  • What evidence demonstrates the current controls?
  • What bounded test was performed, with what pass criteria and safety limits?
  • Is the score inherent or residual?
  • Who owns treatment, by when?
  • Which signals will reveal attempted exploitation or control degradation?
  • Which changes require reassessment?
  • Which other layers or threats can form a cascade?

MAESTRO in Gamut is designed around:

  • Defence in depth: no single control is assumed to break every path.
  • Zero trust: identity, workload, data and agent claims are verified at each boundary.
  • Least privilege: agents, tools and service identities receive only necessary authority.
  • Fail-safe operation: uncertainty or control failure should reduce capability, not silently expand it.
  • Evidence over assertion: a confident narrative is not accepted as proof.
  • Safe testing: tests are bounded, authorised, reversible and non-destructive.
  • Human accountability: AI suggestions cannot approve evidence, accept risk or sign off the assessment.
  1. Method, concepts and terminology
  2. Layers and threat catalogue
  3. System decomposition and architecture patterns
  4. Assessment workflow and risk scoring
  5. Evidence, testing and treatment
  6. AI Assist and security
  7. Reporting and governance
  8. Worked example
  9. Reference and glossary