Purpose and scope of this paper. This is an architectural comparison paper that advances the proposition that execution-gated generation architectures are inherently better aligned with emerging regulatory principles than conventional generation-first architectures. It is intentionally an argument, not a neutral survey. The analysis draws on published regulatory text, architectural documentation, and independent expert review. It does not constitute legal advice, a formal compliance assessment, or a regulatory certification of any kind. Institutions should engage qualified legal counsel for formal compliance determinations. The numerical alignment scores are illustrative architectural assessments developed for comparative analysis, not official regulatory measurements. Where the paper states that COMPAiSS is "designed to support" or "architected to facilitate" a regulatory principle, this reflects an architectural assessment; deployment context, institutional governance, and regulatory determination govern actual compliance.
The Central Thesis - A Different Evaluation Methodology
Conventional AI evaluation and COMPAiSS evaluation are not answering the same question.
Conventional AI Evaluation
"Is the answer accurate?"
Success is defined as factual correctness. The source of the information is secondary to its correctness. A well-grounded, well-scored hallucination is considered a better outcome than a correct answer that lacks source attribution.
COMPAiSS Evaluation
"Was the answer produced exclusively from authorized institutional sources - and is it accurate relative to those sources?"
Success requires authorization first, then accuracy within that authorized corpus. A factually correct answer drawn from an unauthorized source is a governance failure, regardless of its accuracy.
This distinction is not a feature difference. It is a methodological one. Conventional AI systems optimize for the first question. COMPAiSS is architected around the second. For regulated public institutions - government services, universities, hospitals, professional regulators - the second question is the one that matters. When a citizen asks about their benefit eligibility, or a student asks about academic standing, the institution's obligation is not to provide an accurate answer from any available source. It is to provide an authorized answer from its own institutional sources. An unauthorized truth carries the same governance risk as a hallucinated one, because the institution cannot be held accountable for information it did not authorize.
This difference explains why the regulatory alignment comparison that follows consistently favors an authorization-first architecture. The EU AI Act, the Canadian Treasury Board Directive, PIPEDA, and Quebec Law 25 were all written in response to the governance failures of generation-first systems. Their traceability, transparency, accountability, and human oversight requirements map directly and naturally to an architecture that enforces authorization before generation - not because COMPAiSS was designed to satisfy those regulations, but because both are responses to the same underlying institutional governance problem.
Traditional AI Evaluation Metrics
Conventional AI evaluation frameworks emphasize performance metrics applicable across general-purpose use cases:
- Factual accuracy
- Retrieval precision and recall
- Response latency
- Robustness to adversarial inputs
- Hallucination rate
- User satisfaction
- Context faithfulness
These metrics are necessary and well-established. This paper does not propose replacing them.
Proposed Additional Dimensions for Institutional AI
This paper proposes that AI systems operating in regulated institutional environments require additional governance-oriented evaluation dimensions:
- Authorization Fidelity - the degree to which a system remains within an institution's authorized knowledge boundary
- Governance Preservation Rate - the proportion of responses that avoid unauthorized institutional claims
- Evidence Fidelity - the degree to which each claim can be traced to a specific parsed source
- Audit Traceability - the completeness and determinism of the per-response governance record
- Institutional Accountability - the degree to which human authority governs system scope before inference
- Safe Refusal Rate - the proportion of out-of-scope queries that produce appropriate refusals rather than unauthorized responses
Conceptual Model
The Architectural Inversion: Governance Before vs. After Generation
The paper's central architectural distinction illustrated as a process model. Governance moves from after generation to before generation - changing not just when controls are applied, but what kind of controls are structurally possible.
Generation-First Architecture (Conventional RAG)
1 Question received
↓
2 Query embedding and retrieval
↓
3 Candidate ranking and reranking
↓
4 Context assembly
↓
5 Generation - model inference runs
↓
6 Answer delivered to user
↓
7 Governance review, monitoring, logging
Authorization-First Architecture (COMPAiSS)
1 Question received
↓
GATE Authorization check against greenlist
↓
GATE Source validation - does evidence exist?
↓
4 Live parsing of authorized sources
↓
5 Generation - only if gates pass
↓
RECEIPT Compliance receipt generated at runtime
↓
7 Answer delivered with audit trail
In the generation-first model, governance is applied to outputs the model has already produced. In the authorization-first model, governance determines whether the model is permitted to produce output at all. This distinction is not a matter of degree - it is a difference in the structural position of the governance control relative to the inference step. All downstream properties of the two architectures - traceability, accountability, audit trail completeness, adversarial robustness, procurement defensibility - flow from this single architectural difference.
The generation-first model asks: How do we manage the risks of outputs the model has already generated? The authorization-first model asks: How do we govern whether the model is permitted to generate at all? For regulated institutions, the second question is the one that most regulatory frameworks are ultimately designed to answer.
Proposed Framework
An Expanded Evaluation Framework for Authorization-First Institutional AI
This framework does not replace traditional AI evaluation. It proposes additional governance-oriented dimensions appropriate for AI systems operating in regulated institutional environments.
Traditional AI evaluation frameworks were developed for general-purpose AI systems operating across broad domains. They optimize for accuracy, efficiency, and user satisfaction - all essential properties. However, they were not designed for a class of AI systems whose primary obligation is not to maximize the range of questions they can answer, but to ensure that every answer they do provide falls within the governance boundaries of a specific regulated institution.
This paper proposes that authorization-first institutional AI systems warrant evaluation across two complementary sets of dimensions:
Dimension Set A - Traditional Performance
Retained from Conventional AI Evaluation
Factual accuracy within authorized scope. Retrieval precision for authorized sources. Response latency and throughput. Robustness to adversarial inputs. User comprehension and satisfaction. Multilingual accuracy and consistency. These dimensions remain essential and should be evaluated alongside the governance dimensions below.
Dimension Set B - Governance Performance
Proposed for Authorization-First Systems
Authorization fidelity across query types. Governance preservation rate under adversarial testing. Evidence fidelity - claim-to-source traceability. Audit trail completeness and determinism. Safe refusal rate for out-of-scope queries. Human oversight integration depth. Procurement documentation completeness.
The key methodological proposition is that for regulated institutional AI, a system that scores highly on Dimension Set A but poorly on Dimension Set B is not a high-performing system - it is a high-performing system with an unacceptable governance profile. Conversely, a system that scores highly on both sets is genuinely well-suited for regulated institutional deployment. The COMPAiSS Version 8 Proof-of-Concept Validation Framework represents one early operationalization of Dimension Set B, including the Governance Preservation Rate, Constitutional Fidelity Rate, and Evidence Integrity Rate. Independent validation, standardization, and broader empirical testing of these dimensions is a priority for future research.
The Regulatory Landscape - June 2026
What the Frameworks Actually Require
The current state of both frameworks before drawing architectural comparisons.
European Union
EU AI Act (Regulation 2024/1689)
Entered into force August 2024. Prohibited practices enforced from February 2025. GPAI transparency requirements in force from August 2025. Full high-risk system requirements apply August 2026. Penalties up to €35 million or 7% of global turnover. Education, employment, essential services, and public-sector AI are classified as high-risk categories under Annex III, with mandatory conformity assessments, technical documentation, human oversight, traceability, and registration in the EU AI database.
Canada
Federal Governance Framework (2026)
AIDA (Bill C-27) died in Parliament in January 2025 and has not been re-introduced. Binding instruments in 2026 are: the Treasury Board Directive on Automated Decision-Making (in force since 2019 for federal institutions), PIPEDA, Quebec's Law 25 (the most demanding provincial requirement, including rights to human review of automated decisions), OSFI guidance for financial institutions, and the Canada-EU MOU on AI signed December 2025 signalling alignment with EU principles. Federal procurement is increasingly used as a de facto standard-setting mechanism.
Why AIDA's absence strengthens the compliance case
The failure of AIDA to pass does not create a regulatory vacuum. It creates a compliance patchwork that is in practice more demanding than a single statute, because institutions must satisfy multiple overlapping instruments simultaneously: federal privacy law, provincial privacy regimes, Treasury Board directives, sector guidance, and procurement standards. An architecture designed to facilitate compliance with the underlying principles of traceability, authorization, and human oversight is well positioned to satisfy all of these simultaneously - not as a series of separate compliance exercises, but as a natural consequence of how the system is built.
Part 1 - European Union AI Act
COMPAiSS vs. Conventional RAG Implementations: EU Principles
Eight core obligations under the EU AI Act. The comparison is against typical conventional RAG implementations. Sophisticated enterprise RAG deployments that include deterministic retrieval filters, provenance tracking, observability stacks, and human review workflows can narrow some of these gaps through supplementary governance controls - though not through architectural equivalence.
COMPAiSS
COMPAiSS's execution gate is designed to structurally limit output to authorized institutional sources, which appears likely to reduce the system's risk profile in impact assessments. In a completed Algorithmic Impact Assessment conducted for a Service Canada deployment, the system was assessed at Level 1 impact - the lowest tier under the Treasury Board framework - reflecting that its architecture provides information within a defined, authorized scope rather than making autonomous decisions. This assessment outcome is consistent with an architecture in which the output space is explicitly bounded by institutional authorization.
Conventional RAG Implementations
Conventional RAG implementations for institutional contexts - student services, government benefits, healthcare information - typically fall into EU AI Act high-risk categories under Annex III. Each pipeline stage introduces probabilistic uncertainty that is difficult to bound in advance, making impact assessment less predictable and potentially requiring classification at a higher risk tier. Additional governance controls are generally required to achieve comparable risk bounds. Sophisticated enterprise RAG implementations with source restrictions and confidence thresholds can narrow this gap, but typically cannot eliminate it without supplementary architecture.
COMPAiSS
Article 13 requires that high-risk AI systems be designed so their operation is sufficiently transparent for deployers to interpret outputs. COMPAiSS is designed to support this at the architectural level: every response is traceable to a specific, named, authorized source URL that was actually fetched and parsed for that query. The compliance receipt generated per interaction documents the gate decision, sources retrieved, pages parsed, assets injected, and token usage - providing deployers with a detailed operational record without supplementary tooling. Deployers can directly inspect the source of any claim by visiting the cited institutional URL.
Conventional RAG Implementations
Conventional RAG implementations typically provide citations indicating which documents were retrieved, but the generation step can blend, summarize, or infer from retrieved context in ways that make the specific origin of individual claims unclear. A statement in the AI's response may draw on multiple retrieved passages, on parametric training knowledge, or on inference - and distinguishing between these sources in a standardized, auditable form typically requires additional logging and attribution tooling beyond the core pipeline. Meeting Article 13's transparency requirements generally requires supplementary observability infrastructure.
COMPAiSS
Article 14 requires that high-risk AI systems be designed to allow deployers to implement meaningful human oversight. COMPAiSS is designed to embed human governance at two architecturally enforced points. First, the greenlist is a human-curated, human-maintained authorization list: institutional administrators control which sources are authorized, can add or remove URLs, and conduct regular audits through the greenlist dashboard. Second, the execution gate is a structural checkpoint that reflects human-defined institutional scope before any inference occurs. Changes to what the system can address require explicit human authorization - oversight is pre-inference, not post-generation.
Conventional RAG Implementations
Conventional RAG implementations require human oversight to be implemented as a governance layer on top of the generation architecture: monitoring dashboards, human-in-the-loop review queues, escalation workflows, or post-generation quality controls. This is human oversight as remediation rather than as structural constraint. Sophisticated enterprise deployments can add approval workflows and source restrictions that partially close this gap, but these controls typically operate on the retrieval or output layer rather than as a pre-inference authorization gate. The generation step itself remains primary, with human oversight as a secondary control.
COMPAiSS
Articles 12, 18, and 19 establish requirements for logging, record-keeping, and automatically generated audit trails. COMPAiSS is designed to generate a compliance receipt for every interaction as part of the response pipeline. Each receipt records the gate decision and rule applied, the sources retrieved and their relevance scores, the pages actually parsed, the authoritative assets injected, whether inference was permitted or blocked, and the token usage. This creates a forensically traceable record for every response - not a secondary log, but a governance artifact produced at runtime. Auditors can trace any generated claim back through the pipeline to its specific source document.
Conventional RAG Implementations
Conventional RAG deployments can be instrumented to produce logs, but logging is typically not architecturally native to the generation pipeline. Production deployments frequently require separate observability tools to capture retrieval and generation steps with sufficient fidelity for regulatory audit. When a hallucination or scope error occurs, post-hoc log analysis often cannot definitively establish which retrieved content contributed to a specific generated claim, because the generation step blends context probabilistically. Meeting Articles 12, 18, and 19 for high-risk systems typically requires substantial additional observability infrastructure beyond the core RAG pipeline.
COMPAiSS
Article 15 requires appropriate levels of accuracy, robustness, and cybersecurity, including resistance to adversarial inputs. COMPAiSS's architecture is designed to reduce certain classes of adversarial risk: because the gate decision is based on a deterministic lookup against the greenlist rather than model inference, adversarial language inside the prompt is less likely to redirect the system to unauthorized sources. However, the architecture does not eliminate all security risks - parsing pipelines can be attacked, authorized sources can contain compromised content, translation layers introduce their own vulnerabilities, and downstream LLM behavior remains probabilistic within the authorized scope. The architecture appears to materially reduce certain adversarial attack surfaces without claiming to eliminate them.
Conventional RAG Implementations
Conventional RAG implementations are subject to a broader range of adversarial vulnerabilities: prompt injection through retrieved documents, jailbreaking via creative phrasing, retrieval poisoning, and model exploitation through crafted inputs. Robustness requires dedicated adversarial testing programs, regular red-teaming, guardrail models, content moderation layers, and prompt safety classifiers - each of which is an additional system component that can itself fail or be circumvented. The probabilistic nature of generation means accuracy is expressed as a rate rather than a structural bound, which is a different kind of claim when presented to regulators evaluating Article 15 compliance.
COMPAiSS
Article 10 requires that training, validation, and testing datasets meet quality criteria and that data governance is maintained. COMPAiSS's architecture is designed to limit the data governance burden for institutional answers: the epistemic environment provided to the model at inference time consists of content fetched from greenlist-authorized, live-parsed institutional sources. The institution controls what that content is, audits it through the greenlist dashboard, and updates it as sources change. There is no opaque or unauditable data pool contributing to institutional answers beyond what the institution itself has authorized. The data governance scope is defined by the greenlist rather than by a broad training corpus.
Conventional RAG Implementations
Conventional RAG implementations inherit data governance obligations from both the base LLM's training data and the retrieval corpus they operate over. Establishing the provenance, quality, and representativeness of retrieval corpora for regulated institutional use cases requires significant data auditing work. When base model parametric knowledge contributes to a generated response alongside retrieved content, establishing which parts of the output originate from which data source is technically complex. Article 10 compliance typically requires a formal data governance program covering both the training data and the retrieval corpus, with separate documentation for each.
COMPAiSS
Article 11 requires providers of high-risk AI systems to produce technical documentation sufficient to demonstrate conformity. COMPAiSS's deterministic, bounded architecture makes technical documentation more tractable than for probabilistic systems: the execution gate logic, greenlist authorization mechanism, parsing pipeline, evidence requirements, and compliance receipt generation are architecturally stable and fully documentable. The COMPAiSS Version 8 Proof-of-Concept Validation Framework provides a formal evaluation methodology with structured performance metrics, an inter-rater reliability standard (Cohen's Kappa ≥ 0.70), and a typed failure taxonomy - all of which are directly mappable to Article 11 documentation requirements.
Conventional RAG Implementations
Producing Article 11-compliant technical documentation for a conventional RAG deployment is substantially more complex, because the system's behavior is not fully deterministic and can shift with model updates, retrieval corpus changes, or embedding model updates. Documentation must account for the interaction of multiple AI model stages, each with their own performance characteristics, failure modes, and update cycles. The documentation burden is ongoing: changes to any component of the RAG stack may require documentation updates and potentially new conformity assessments. Enterprise RAG vendors typically require dedicated compliance engineering to maintain this documentation at the level Article 11 requires for high-risk systems.
COMPAiSS
Articles 17 and 27 require a quality management system and, for public-sector deployers, a fundamental rights impact assessment before deployment. COMPAiSS's architecture is designed to make both more tractable. The quality management system is substantially built into the architecture: the greenlist audit cycle, the compliance receipt trail, the execution gate's pass/fail decision, and the structured failure taxonomy (Type A utility refusals, Type B governance violations, Type C pipeline defects) provide a built-in quality framework with measurable, auditable metrics. The fundamental rights impact assessment is facilitated by the bounded scope: the output space is explicitly defined by institutional authorization rather than by the full range of model capability.
Conventional RAG Implementations
Building a quality management system for a conventional RAG deployment requires constructing quality controls at every stage of the probabilistic pipeline: retrieval quality monitoring, generation faithfulness scoring, post-generation output review, incident reporting workflows, and continuous performance monitoring. The fundamental rights impact assessment for a general-purpose RAG deployed in public services, healthcare, or education must account for a wider and less predictable output space. Each of these elements requires organizational investment that COMPAiSS's architecture is designed to reduce by constraining the output space to an authorized institutional scope from the start.
Part 2 - Canadian Regulatory Framework
COMPAiSS vs. Conventional RAG Implementations: Canadian Principles
Seven principles drawn from the Treasury Board Directive on Automated Decision-Making, PIPEDA, Quebec Law 25, and the Canada-EU MOU on AI (December 2025). Sophisticated enterprise RAG implementations with supplementary governance layers can close some of these gaps; the comparison reflects typical implementations without such supplementation.
COMPAiSS
The Treasury Board Directive requires federal institutions to conduct Algorithmic Impact Assessments prior to deploying automated decision systems. A completed AIA for a Service Canada deployment assessed the COMPAiSS implementation at Level 1 impact - the lowest tier - because the architecture provides information rather than making administrative decisions, and the execution gate is designed to prevent outputs beyond the authorized institutional scope. This assessment outcome reflects an architecture specifically suited to the information-provision use case that falls at the lower end of the Directive's impact scale. The AIA completion constitutes documented procurement history with a federal institution.
Conventional RAG Implementations
A conventional RAG deployment for public-facing government services is more likely to attract a higher impact tier under the Treasury Board Directive, particularly where the information provided affects citizens' understanding of their benefits, eligibility, or obligations. The probabilistic output space and potential for cross-jurisdictional information to influence citizens' decisions elevates assessed risk. Meeting the Directive's obligations at a higher tier requires mandatory human review checkpoints, more extensive documentation, and ongoing monitoring mechanisms. Sophisticated enterprise RAG deployments with source restrictions can narrow this gap but typically cannot achieve the same bounded risk profile as an authorization-first architecture.
COMPAiSS
Canadian frameworks require that individuals be informed about how automated systems work and, under Quebec Law 25, have the right to be informed of automated decisions affecting them. COMPAiSS is designed to support transparency structurally: users receive answers sourced from named, publicly accessible institutional URLs, with the ability to verify any claim at its source. When the system cannot answer - because no authorized source supports the query - it explicitly communicates this and redirects to institutional resources rather than generating a plausible but unsupported response. The system's constraints are visible to users through its behavior, not disclosed separately in fine print.
Conventional RAG Implementations
Conventional RAG implementations cite sources, but the relationship between a cited source and a specific generated claim is often non-trivial to establish for explanation purposes. The model may cite a document that informed the response generally while specific phrasing or factual claims were generated from parametric knowledge or inference. For Quebec Law 25's right-to-explanation requirements and for PIPEDA transparency obligations in contexts involving personal information, this blended provenance is challenging to satisfy without additional explainability tooling and user-facing disclosure infrastructure.
COMPAiSS
Both the Treasury Board Directive and Quebec Law 25 require that meaningful human intervention be available in automated processes that affect individuals. COMPAiSS's greenlist governance structure is designed to make this requirement native: institution administrators actively control what the system can respond to, can remove or update authorized sources in real time, and conduct regular audits. The system's structural refusal of out-of-scope queries ensures that edge cases are escalated to human processes rather than addressed by model inference. Human authority over institutional scope is a precondition for inference, not a dashboard override applied after generation.
Conventional RAG Implementations
Meaningful human intervention in conventional RAG deployments for regulated public services typically requires human review queues for flagged outputs, escalation mechanisms when confidence thresholds are not met, and ongoing output monitoring. These mechanisms represent human intervention as a reactive control rather than a pre-inference structural constraint. For Quebec Law 25's right to request human review of automated decisions, RAG deployments require explicit process infrastructure beyond the AI pipeline. Enterprise implementations with approval workflows can partially close this gap, but the underlying generation step typically remains primary rather than secondary to human authorization.
COMPAiSS
COMPAiSS's greenlist architecture is designed to provide structural bias mitigation for institutional use cases by constraining outputs to authorized institutional sources. Because the system draws from the institution's own published policies - which must by law be non-discriminatory - the bias introduced by the AI is bounded by the source material rather than by broad training corpus characteristics. Multilingual capability is governed pre-inference: translation normalizes the question before authorization, and the authorized answer is then translated back, designed to ensure users in different languages receive the same institutionally authorized content rather than a corpus-shaped approximation. Independent bias testing across languages and demographic groups would strengthen this assessment.
Conventional RAG Implementations
Conventional RAG implementations can introduce bias through multiple channels: base model training biases, retrieval ranking biases that favor certain document types or phrasings, and generation biases in how retrieved information is synthesized. In multilingual contexts, different languages may retrieve different documents or generate differently weighted responses, potentially introducing differential outcomes across user populations. Bias mitigation requires active testing programs across languages and demographic groups, continuous monitoring, and regular model evaluation - a non-trivial ongoing investment that sophisticated enterprise RAG deployments typically address through dedicated testing programs.
COMPAiSS
The Treasury Board Directive requires that automated decision systems maintain audit trails sufficient to reconstruct decisions and demonstrate compliance. The Canada-EU MOU on AI, signed December 2025, commits Canada to aligning with EU principles including traceability throughout the model lifecycle. COMPAiSS is designed to generate a per-interaction compliance receipt tracing the gate decision, sources consulted, pages parsed, and inference authorization status. Institutional auditors can trace any generated output to specific parsed source content - or confirm that a query was appropriately refused and why. This creates the kind of deterministic audit trail that both the Treasury Board Directive and the Canada-EU MOU's traceability commitment appear to require.
Conventional RAG Implementations
Producing audit trails for conventional RAG that satisfy both the Treasury Board Directive and the EU traceability principles in the December 2025 Canada-EU MOU requires observability infrastructure that is not standard in typical RAG deployments. When an audit requires establishing which retrieved content produced which specific claim, the probabilistic generation step means this tracing is often approximate rather than definitive. A conventional RAG system that generates an incorrect or out-of-scope response may produce logs, but those logs cannot always definitively establish whether the error originated in retrieval, generation, or their interaction. Enterprise observability tooling can improve this, but cannot provide the same claim-level traceability as a runtime governance receipt.
COMPAiSS
PIPEDA and Quebec Law 25 require that personal information be collected only as necessary and used only for the purpose for which it was collected. COMPAiSS is designed to process user queries against a defined, institution-controlled greenlist without building user profiles, retaining personal data for model training, or passing information to external data pools. The greenlist contains only publicly authorized institutional URLs, not personal information. The architecture's data minimization is structural: it uses no more data than is necessary for authorization and inference, and retains no data beyond what is captured in the per-interaction compliance receipt. The deployment context and specific implementation may affect these properties and should be confirmed for each institutional deployment.
Conventional RAG Implementations
Conventional RAG deployments in public-sector or institutional contexts must carefully manage query data, conversation history, user context, and retrieval logs under PIPEDA and Law 25. Enterprise RAG platforms may maintain conversation history, use queries to improve retrieval systems, or share telemetry with model providers - each of which triggers privacy obligations requiring explicit governance. Data residency requirements under Law 25 impose additional constraints on where data may be processed. Building a RAG deployment that satisfies PIPEDA and Law 25 simultaneously requires explicit privacy architecture decisions that are not default in standard enterprise RAG implementations.
COMPAiSS
Federal and provincial procurement frameworks increasingly require institutions to demonstrate how unauthorized AI responses are structurally prevented - not merely managed after delivery. COMPAiSS provides an architecturally defensible answer to this requirement: the execution gate is a documented, testable mechanism designed to prevent unauthorized inference. The compliance receipt provides a per-interaction audit trail directly compatible with Treasury Board documentation requirements. The completion of an Algorithmic Impact Assessment for a Service Canada deployment demonstrates that the formal federal procurement evaluation pathway has been navigated. This constitutes verified institutional procurement history rather than a theoretical compliance claim.
Conventional RAG Implementations
Conventional RAG vendors typically cannot demonstrate structural prevention of unauthorized outputs at procurement. Their compliance story is probabilistic: hallucination rates have been reduced, monitoring is in place, human review is available. Procurement officers who ask "how does the system structurally prevent unauthorized responses?" cannot receive a structural answer from a generation-first architecture, because generation itself is the first step. The response must be "we manage unauthorized responses after they occur." This procurement vulnerability becomes an increasing liability as federal and provincial procurement guidance tightens around AI governance requirements. Sophisticated enterprise implementations with approval workflows partially close this gap but do not fully resolve the underlying architectural claim.
Summary
Illustrative Alignment Scores: All Principles
Scoring Methodology and Limitations
These are illustrative architectural alignment assessments, not official regulatory scores, compliance certifications, or legal determinations. No numerical scoring rubric exists under either the EU AI Act or the Canadian regulatory framework. The scores reflect the authors' assessment of how naturally each architecture is designed to support each regulatory principle - a score of 9–10 indicates the principle is addressed structurally by the architecture; a score of 4–5 indicates the principle can be addressed through supplementary governance controls but is not addressed by the core architecture. These assessments are based on published regulatory text, architectural documentation, and cross-model review by GPT-4, Gemini, and Copilot. They represent a reasoned architectural argument, not a measured compliance outcome. Institutions should engage qualified legal counsel for formal compliance assessments. The comparison is against typical conventional RAG implementations; sophisticated enterprise RAG deployments with deterministic source controls, provenance tracking, and observability infrastructure can close some of these gaps through supplementary governance layers.
| # |
Principle |
Framework |
COMPAiSS |
Conventional RAG |
Gap |
| EU-01 | Risk Classification and Impact Assessment | EU AI Act Art. 9, 27 |
9.0 |
4.0 |
+5.0 |
| EU-02 | Transparency and Information to Deployers | EU AI Act Art. 13 |
9.0 |
4.0 |
+5.0 |
| EU-03 | Human Oversight | EU AI Act Art. 14 |
9.5 |
5.0 |
+4.5 |
| EU-04 | Record-Keeping and Audit Logs | EU AI Act Arts. 12, 18, 19 |
9.5 |
3.5 |
+6.0 |
| EU-05 | Accuracy, Robustness, Cybersecurity | EU AI Act Art. 15 |
8.0 |
4.5 |
+3.5 |
| EU-06 | Data Governance | EU AI Act Art. 10 |
8.5 |
4.0 |
+4.5 |
| EU-07 | Technical Documentation | EU AI Act Art. 11 |
9.0 |
4.0 |
+5.0 |
| EU-08 | Accountability and Quality Management | EU AI Act Arts. 17, 27 |
9.0 |
4.0 |
+5.0 |
| CAN-01 | Algorithmic Impact Assessment and Risk Tiering | Treasury Board Directive |
9.0 |
4.5 |
+4.5 |
| CAN-02 | Transparency and Explainability | Treasury Board; PIPEDA; Law 25 |
9.0 |
4.0 |
+5.0 |
| CAN-03 | Meaningful Human Intervention | Treasury Board; Law 25 Art. 12 |
9.5 |
4.5 |
+5.0 |
| CAN-04 | Bias Mitigation and Equitable Access | Treasury Board; Canadian Human Rights Act |
8.0 |
4.5 |
+3.5 |
| CAN-05 | Auditability and Accountability | Treasury Board; PIPEDA; Canada-EU MOU |
9.5 |
3.5 |
+6.0 |
| CAN-06 | Privacy and Data Minimization | PIPEDA; Quebec Law 25 |
8.5 |
5.0 |
+3.5 |
| CAN-07 | Procurement Defensibility | Federal Procurement; Treasury Board |
9.5 |
3.5 |
+6.0 |
| Overall Illustrative Architectural Alignment (15 principles, mean) |
8.9 / 10 |
4.2 / 10 |
+4.7 |
Cost Implications
The Compliance Cost of Generation-First Architecture
The structural cost difference between achieving regulatory alignment through architecture versus achieving it through supplementary governance controls.
The alignment assessments above reveal a structural cost asymmetry with direct procurement implications. COMPAiSS's alignment across these principles is a consequence of its architecture. A conventional RAG deployment aiming to achieve comparable alignment must invest in supplementary systems at each gap point - and each of those investments recurs.
Conventional RAG: Compliance Supplementation Required
Observability and audit logging infrastructureAdditional layer
Human oversight workflows and staffingOngoing cost
Guardrail models and content moderationAdditional layer
Adversarial testing and red-teaming programsOngoing cost
Technical documentation maintenanceOngoing cost
Bias testing across languages and groupsOngoing cost
Privacy architecture and data residency controlsAdditional layer
Illustrative annual operating cost (100K queries)$90K - $200K
COMPAiSS: Compliance by Architecture
Audit logsBuilt-in
Human oversightGreenlist (built-in)
Output constraintExecution gate (built-in)
Adversarial scope controlStructural (built-in)
Technical documentationDeterministic (tractable)
Multilingual equityPre-inference (built-in)
Privacy / data minimizationArchitectural (built-in)
Illustrative annual operating cost (100K queries)$15K - $30K
Cost estimates are illustrative deployment scenarios based on typical institutional deployments at approximately 100,000 annual queries. Conventional RAG figures reflect inference costs, vector database hosting, observability tooling, human review staffing, and compliance program overhead. COMPAiSS figures reflect inference costs for authorized queries only (approximately 60% of total queries, since unauthorized queries do not incur primary model inference costs), greenlist maintenance, and compliance receipt storage. Actual costs vary significantly by institution size, query volume, hosting environment, existing infrastructure, and governance requirements. These figures should be treated as directional indicators, not procurement-ready estimates. Detailed cost modelling for specific institutional contexts is available on request.
Limitations and Future Validation
What This Paper Claims - and What It Does Not
Scholarly papers distinguish between scope and limitations. This section addresses both explicitly.
This paper advances a testable architectural hypothesis: that authorization-first AI architectures are structurally better positioned to satisfy the governance requirements of regulated institutional environments than conventional generation-first architectures. It supports this hypothesis through architectural analysis, regulatory text mapping, and illustrative alignment assessment. It does not prove it through independent empirical measurement.
The following limitations should be understood by any reader using this paper to inform procurement, policy, or research decisions:
Architectural Assessment, Not Legal Opinion
This paper assesses how an architecture is designed to support regulatory principles. It does not constitute legal advice, a compliance certification, or a regulatory determination. Actual compliance is determined by regulators evaluating organizations, governance systems, processes, and deployment contexts - not software architecture alone.
Alignment Scores Are Illustrative
The numerical alignment scores are expert judgments based on architectural analysis, not measurements produced by a validated scoring instrument. No official regulatory scoring rubric exists under either the EU AI Act or the Canadian framework. These scores should be understood as reasoned assessments, not objective measurements.
The RAG Comparison Is Against Typical Implementations
Sophisticated enterprise RAG implementations with deterministic source controls, provenance tracking, full observability stacks, and human review workflows can narrow some of the governance gaps identified here through supplementary controls. The comparison does not represent a ceiling on what conventional RAG architectures can achieve with significant additional investment.
Cost Estimates Are Illustrative Scenarios
The deployment cost figures are illustrative scenarios based on typical institutional deployments at approximately 100,000 annual queries. They are not audited cost reports. Actual costs vary significantly by institution size, query volume, hosting environment, existing infrastructure, and governance requirements.
Security Claims Are Directional
The paper describes how COMPAiSS's architecture is designed to reduce certain adversarial attack surfaces. It does not claim to eliminate all security risks. Parsing pipelines can be attacked, authorized sources can contain compromised content, translation layers introduce vulnerabilities, and LLM behavior within authorized scope remains probabilistic.
Independent Validation Is Needed
The authorization fidelity construct, governance preservation rate, and expanded evaluation framework proposed here are preliminary. They have not yet been independently validated, standardized, or subjected to peer review outside the cross-model review process described in this paper.
The next stage of validation for the architectural hypothesis advanced here should include: independent third-party security assessment of the execution-gated architecture; comparative benchmarking against sophisticated enterprise RAG implementations with full governance supplementation; longitudinal deployment studies measuring governance outcomes in live institutional environments; independent regulatory review of the proposed alignment assessment methodology; and empirical measurement of authorization fidelity, governance preservation rate, and evidence fidelity across diverse query types and institutional contexts.
The Testable Hypothesis
The central claim of this paper can be stated as a testable hypothesis: AI systems whose architecture enforces authorization as a precondition for generation will, in regulated institutional deployment contexts, demonstrate measurably higher governance preservation rates, more complete audit trails, and lower compliance supplementation costs than systems whose architecture generates first and applies governance controls afterward. This hypothesis is empirically falsifiable. The authors invite independent researchers, regulators, and procurement professionals to test it.
Conclusion
Overarching Themes from the Comparative Analysis
Five structural observations that emerge consistently across both the European and Canadian regulatory frameworks.
Authorization Before Accuracy
Regulatory frameworks for public institutions are not primarily concerned with whether AI answers are correct. They are concerned with whether AI answers are authorized. A factually accurate answer drawn from an unauthorized source is a governance failure. This is the central methodological distinction between conventional AI evaluation and the evaluation framework appropriate for regulated institutional AI.
Determinism vs. Probability
Both regulatory frameworks require traceability, accountability, and predictable behavior. An execution-gated architecture is designed to satisfy these requirements structurally. A probabilistic generation-first architecture is designed to satisfy them through supplementary controls. Regulators are increasingly distinguishing between these two approaches in high-risk institutional contexts.
Pre-Inference vs. Post-Generation Governance
Every major regulatory principle - human oversight, transparency, accountability, record-keeping - can be addressed either before inference or after generation. The regulatory frameworks do not explicitly require pre-inference governance, but they appear to reward architectures that provide it structurally rather than reactively, because structural controls are more traceable, more predictable, and less dependent on secondary systems functioning correctly.
Procurement Inversion
As regulatory requirements tighten, the burden of proof in procurement shifts from "prove this system is safe" toward "demonstrate how unauthorized responses are structurally prevented." This is a question that a generation-first architecture cannot fully answer - because generation itself precedes any structural constraint. An authorization-first architecture is designed to answer it directly.
The Compliance Tax
Conventional RAG deployments in regulated institutional contexts carry a compliance tax: the ongoing cost of governance infrastructure required to satisfy regulatory requirements that an authorization-first architecture is designed to satisfy structurally. This tax manifests as infrastructure costs, staffing costs, monitoring costs, and organizational overhead. For institutions operating at scale, this tax is material and grows with deployment volume and regulatory scrutiny.
The architectural conclusion
Both the EU AI Act and the Canadian framework were developed in response to governance failures observed in generation-first AI deployments in regulated contexts. Their core requirements - traceability, human oversight, transparency, record-keeping, accountability, data governance, procurement defensibility - map more naturally to an architecture where authorization precedes generation than to one where governance is applied after the fact.
This analysis advances the proposition that this alignment is not coincidental. An architecture designed around the question "Was this answer produced exclusively from authorized institutional sources?" is structurally better positioned to satisfy regulatory frameworks that ask the same question - not because it was engineered to satisfy those regulations, but because both are responses to the same underlying institutional governance problem: the need for AI systems whose outputs can be traced, bounded, authorized, and defended.
Future Research Agenda
Questions for the Field
This paper proposes a framework. It does not complete one. The following questions identify where empirical investigation, independent validation, and field-wide collaboration are needed to advance authorization-first institutional AI from an architectural proposition to an established evaluation category.
-
01
How should authorization fidelity be measured? The construct proposed here requires a validated measurement instrument. What query types, adversarial conditions, and institutional contexts are necessary to produce a reliable authorization fidelity score? How should the instrument handle edge cases where authorized sources partially address a query?
-
02
Can architectural alignment scoring be standardized? The illustrative alignment scores in this paper are expert judgments. Is it possible to develop a reproducible, methodologically rigorous scoring instrument that could be applied consistently across AI systems and regulatory frameworks by independent evaluators?
-
03
How do authorization-first systems compare with advanced enterprise RAG under common benchmarks? The comparison in this paper is against typical conventional RAG implementations. A rigorous head-to-head comparison against sophisticated enterprise RAG deployments with full governance supplementation - using the expanded evaluation framework proposed here - would substantially strengthen or challenge the architectural hypothesis.
-
04
What independent testing methodologies are appropriate for governance-oriented AI evaluation? Traditional AI benchmarks are not designed to assess authorization fidelity or governance preservation. What evaluation protocols, adversarial testing regimes, and inter-rater reliability standards are appropriate for assessing Dimension Set B metrics in live institutional deployments?
-
05
Can execution-gated architectures measurably reduce regulatory risk in practice? The architectural hypothesis predicts that authorization-first deployments will produce fewer governance violations, lower compliance supplementation costs, and more complete audit trails than generation-first deployments in comparable institutional contexts. Longitudinal deployment studies across multiple institutions would test this prediction empirically.
-
06
How should procurement frameworks evaluate authorization-first AI? Current procurement frameworks for AI in regulated institutions were developed for generation-first systems. What modifications to procurement criteria, vendor evaluation processes, and compliance documentation requirements are needed to appropriately assess authorization-first architectures - and to prevent generation-first vendors from claiming equivalent governance properties through supplementary controls?
-
07
How should the expanded evaluation framework evolve as AI architectures develop? The governance landscape is changing rapidly. As hybrid architectures emerge that incorporate elements of both generation-first and authorization-first approaches, how should the evaluation framework adapt? Are there authorization-first properties that can be retrofitted into existing enterprise RAG deployments, and if so, at what governance cost?
These questions define a research agenda rather than a conclusion. The proposition that authorization-first AI represents a distinct architectural category warranting its own evaluation methodology is offered here as a starting point for a broader field-wide conversation - among AI researchers, regulatory scholars, institutional AI practitioners, procurement professionals, and the regulated institutions that depend on AI systems they can genuinely govern.