Evidence, testing and findings
Evidence and testing convert a legal answer from assertion into assurance. They must demonstrate the named system, regulated role, relevant version and assessment period.
Evidence quality
Section titled “Evidence quality”Review each item for:
- Relevance to the exact atomic obligation.
- System and entity scope.
- Version and environment.
- Currency and operating period.
- Authenticity and integrity.
- Completeness.
- Independent review.
- Contradictory information.
Generic supplier marketing, a policy with no implementation record, or an assessor narrative is not automatically accepted evidence.
Testing principles
Section titled “Testing principles”A defensible test records:
- Objective.
- Legal/technical requirement.
- System, role and environment.
- Population and sample.
- Steps.
- Expected result and pass criteria.
- Safety and privacy limits.
- Tester and date.
- Actual result and exceptions.
- Evidence captured.
- Retest requirement.
Use non-destructive environments and synthetic data where possible. Production or adversarial tests require explicit authority and safeguards.
Scope, role and AI literacy
Section titled “Scope, role and AI literacy”Evidence may include:
- Corporate and market-placement records.
- System/model inventory.
- Intended-purpose and architecture documents.
- Contracts and role matrices.
- Geographic output/use analysis.
- Exclusion memorandum.
- Training needs analysis, role-specific curriculum and comprehension results.
Test by tracing one system from supplier through deployed use and asking whether each factual role and EU nexus is supported by independent records.
Prohibited practices
Section titled “Prohibited practices”Evidence should address the actual feature and use, not merely policy language:
- Product requirements and prohibited-use controls.
- Model and feature configuration.
- Data-source records.
- User journeys and decision logic.
- Marketing and sales claims.
- Access restrictions.
- Red-team or misuse tests.
- Legal analysis for any narrow exception.
Test each of the eight Article 5 points. Where a dangerous capability could be enabled through an alternate endpoint, tool or configuration, include that path.
Any positive or uncertain result requires escalation and blocks confirmation.
High-risk classification
Section titled “High-risk classification”Inspect:
- Product legislation and conformity route.
- Annex III use-case mapping.
- Intended purpose and actual workflow.
- Affected decisions and persons.
- Profiling analysis.
- Article 6(3) exception memorandum.
- Article 6(4) documentation and registration, where relevant.
Test classification by reconstructing the decision from facts without relying on the existing label. Independently challenge industry-based assumptions.
Risk management
Section titled “Risk management”Evidence:
- Continuous lifecycle risk-management plan.
- Hazard, misuse and fundamental-rights scenarios.
- Risk estimates and acceptance criteria.
- Testing and validation plans.
- Residual-risk decisions.
- Change and monitoring integration.
Test whether a material model, data or intended-purpose change updates hazards, mitigations, validation and residual-risk approval.
Data governance and quality
Section titled “Data governance and quality”Evidence:
- Data provenance, collection and preparation records.
- Relevance and representativeness analysis.
- Bias and error analysis.
- Data quality thresholds.
- Special-category handling.
- Dataset versioning and lineage.
- Remediation and exception records.
Test traceability from a sampled output or failure back to the data version and quality controls.
Technical documentation and instructions
Section titled “Technical documentation and instructions”Evidence:
- Annex IV technical file.
- Architecture and model description.
- Performance and limitation results.
- Risk, data, logging, oversight and cybersecurity sections.
- Version/change history.
- Instructions for use.
Test whether the file matches the released system and whether a material change triggers controlled updates before release.
Logging
Section titled “Logging”Evidence:
- Logging design and schema.
- Event samples.
- Integrity, access and retention controls.
- Clock synchronisation.
- Privacy minimisation.
- Retrieval and investigation procedures.
Test that a representative decision or incident can be reconstructed, and that deployer-controlled logs meet the applicable minimum retention period.
Human oversight
Section titled “Human oversight”Evidence:
- Named oversight roles.
- Competence and training.
- Decision authority.
- Interface and alert design.
- Override, stop and rollback mechanisms.
- Workload and automation-bias controls.
- Exercise results.
Test a consequential scenario from abnormal output to detection, comprehension, intervention and safe recovery. A nominal human presence is insufficient if intervention is too late or lacks authority.
Accuracy, robustness and cybersecurity
Section titled “Accuracy, robustness and cybersecurity”Evidence:
- Declared metrics and tolerances.
- Validation across relevant groups and conditions.
- Robustness, drift and failure testing.
- Secure development and vulnerability management.
- Prompt-injection, poisoning, evasion and extraction testing where relevant.
- Incident and recovery records.
Test the actual deployed configuration and foreseeable conditions, including interfaces and dependencies.
Provider and conformity duties
Section titled “Provider and conformity duties”Evidence:
- Quality-management system.
- Technical-document retention.
- Conformity assessment.
- EU declaration of conformity.
- CE marking.
- EU database registration.
- Corrective-action and withdrawal process.
- Authority-cooperation process.
Test a release gate using a missing or inconsistent mandatory artefact. The system should fail closed.
Deployer duties
Section titled “Deployer duties”Test the twelve atomic duties separately, including:
- Following instructions.
- Competent oversight.
- Input-data fitness.
- Monitoring and abnormal-operation escalation.
- Incident communication.
- Protected logs and minimum six-month retention.
- Workplace notification.
- Public-authority registration checks.
- DPIA/FRIA linkage.
- Notice to affected persons.
- Post-remote-biometric safeguards.
Importer, distributor and value chain
Section titled “Importer, distributor and value chain”Evidence:
- Pre-market checklists.
- Provider and representative records.
- CE/declaration/instruction checks.
- Storage and transport controls.
- Stop, recall and corrective-action authority.
- Substantial-modification assessment.
- Information and assistance rights.
Test whether a missing conformity artefact actually blocks release, and whether a material change causes provider-role and conformity reassessment.
Fundamental-rights impact assessment
Section titled “Fundamental-rights impact assessment”Evidence should cover:
- Deployer process and intended purpose.
- Duration and frequency.
- Affected natural-person categories and groups.
- Specific risks to fundamental rights.
- Human oversight.
- Mitigation and governance.
- Complaint, contestability and redress.
- Residual risk and approval.
- Authority notification where required.
- Review triggers.
Test the FRIA against a credible affected-person scenario. Confirm it changes design or deployment where the analysis identifies unacceptable or poorly controlled impact.
Article 50 transparency
Section titled “Article 50 transparency”Inspect the actual user experience:
- Notice copy and placement.
- First-interaction timing.
- Accessibility.
- Machine-readable marking.
- Deepfake or public-interest disclosure.
- Emotion/biometric notices.
- Language and channel variants.
- Exception analysis.
Test each trigger across representative interfaces, API outputs and content transformations. Verify that downstream processing does not strip required markings.
Monitoring and serious incidents
Section titled “Monitoring and serious incidents”Evidence:
- Post-market monitoring plan.
- Telemetry and complaint sources.
- Trend and threshold analysis.
- Serious-incident definition and triage.
- Provider/deployer escalation.
- Authority reporting decisions and timelines.
- Corrective action.
Run a tabletop incident from detection through legal classification, containment, preservation, notification and follow-up.
Depending on role, inspect:
- Model/version and market-placement record.
- Systemic-risk classification.
- Annex XI technical documentation.
- Annex XII downstream information.
- Copyright policy and rights-reservation controls.
- Public training-content summary.
- Open-source exemption analysis.
- Evaluation, systemic-risk mitigation and cybersecurity.
- Serious-incident reporting.
- Authorised representative mandate.
- Supplier evidence for deployers/API consumers.
The voluntary GPAI Code can be useful evidence, but code adherence is not conclusive proof for every fact or obligation.
Findings
Section titled “Findings”Raise a finding when evidence or testing shows:
- Legal non-compliance.
- Unsupported classification or N/A.
- Failed operating control.
- Stale or rejected evidence.
- Inadequate scope or role determination.
- Unresolved contradiction.
- Missing notification or authority-cooperation capability.
A useful finding includes condition, requirement, evidence, affected system and role, consequence, root cause, severity, recommendation, owner, target date, status and retest.
Evidence-to-conclusion rules
Section titled “Evidence-to-conclusion rules”- Accepted evidence can support a claim; it does not erase a failed test.
- A passing test does not cure an absent legal process outside the tested scope.
- An open adverse finding blocks depth 3 for the affected item.
- Supplier evidence must be mapped to the customer’s role and implementation.
- Evidence from one system cannot silently assure another.
- Risk acceptance does not change a non-compliant answer.