The Provenance Gap – Why Enterprise AI Is Creating a New Evidence Challenge

By Christian Siegers, of KPMG Netherlands

One of the more interesting shifts taking place in enterprise AI is the growing recognition that building AI systems is only part of the challenge. As organisations begin deploying generative and agentic AI in production environments, attention is gradually moving towards understanding how these systems behave once they are running.

The first wave of enterprise AI largely focused on capabilities: foundation models, retrieval architectures, copilots, orchestration frameworks and increasingly sophisticated agents. Those conversations remain important, but operational adoption brings different questions to the foreground. How do we govern systems that make decisions at runtime, understand why an agent selected a course of action and establish trust in outcomes assembled dynamically from multiple sources?

These questions shift the discussion away from intelligence and towards governance, exposing a challenge that remains underappreciated. The challenge is not simply visibility. It is evidence.

Enterprise architecture was designed around predictable evidence

Traditional enterprise systems were built around predictable execution. Applications processed inputs through predefined business logic, while decision rules and processing chains could be documented, tested and reviewed before deployment. This shaped governance: architects assessed requirements, risk functions evaluated controls, engineers tested the logic and operational teams monitored whether the system behaved as expected.

The evidence explaining an outcome was largely embedded in the solution itself. Business rules, decision tables, source data and process logic provided a stable basis for investigation. When something unexpected occurred, organisations could compare the execution path with the intended logic and identify the error. For many conventional applications, this model still works. AI challenges that assumption.

AI is changing how evidence is assembled

Generative AI systems do not simply execute predefined rules. They interpret instructions, retrieve information, recognise patterns and combine sources into newly generated outputs. Agentic systems extend this behaviour further by formulating plans, selecting tools, comparing alternatives, adjusting execution paths and initiating actions across enterprise systems.

Part of the resulting behaviour emerges while the system is running, so the supporting evidence is assembled dynamically as well. A model may retrieve information unknown at design time, an agent may encounter several versions of a document, and a system may combine policies, conversation history, business data and tool responses into a conclusion that appears in no individual source.

Explanation therefore requires more than tracing predefined logic. It increasingly means reconstructing the context that was assembled, the sources selected and the relationships between them. The question is not only whether the system followed its intended process, but whether the evidence underlying the outcome was appropriate.

Visibility and evidence are not the same

This is one reason why AI observability has received so much attention. As AI systems become increasingly dynamic, organisations need visibility into prompts, model interactions, retrieval results, tool calls, latency, token consumption, workflow behaviour and generated outputs. Without it, failures become difficult to investigate, costs difficult to manage and agent behaviour difficult to evaluate.

From an architecture perspective, this investment makes sense. Observability provides the technical records required to reconstruct what happened during execution: which model was invoked, which documents were retrieved, which tools were selected and how the system arrived at its final response. That is an essential capability, yet visibility and evidence are not the same.

An execution trace may show which documents an agent retrieved, but not why they were authoritative, current or applicable. It may show that a policy was used without demonstrating that it was the approved version or that conflicting information was resolved correctly. Observability can reconstruct the path taken; it cannot always establish whether the basis of the outcome was defensible.

Enterprise architecture has repeatedly seen its governance boundary expand. It moved from individual applications and infrastructure towards data quality, ownership, metadata and lineage, and then towards distributed cloud services and identities. AI is driving the next stage. An enterprise AI outcome rarely comes from a model alone; it emerges from prompts, retrieval mechanisms, enterprise knowledge, orchestration, tools, policies and runtime decisions.

Yet much of today’s AI governance discussion remains centred on models: performance, safety, responsible AI, security and regulatory classifications. Those concerns matter, but governing the model does not govern the complete evidence chain. Organisations must also know where supporting information originated, why particular sources were selected and whether their authority remained valid when used.

This reflects a broader shift from design-time towards runtime governance. Reviewing a solution before deployment cannot fully control how it behaves when it interprets context, selects tools and makes autonomous choices. Evidence follows the same path: in AI-enabled environments, its relevance and authority may change during execution as policies are superseded, sources conflict or semantically relevant content proves non-authoritative.

Context is becoming part of the control model

This is why the growing attention to context engineering is so interesting. Organisations are developing retrieval capabilities, memory services, semantic models, knowledge graphs and context platforms to connect AI systems to enterprise knowledge. Better context is likely to improve outcomes, but relevance alone does not establish trust.

An agent needs more than information that appears related to its task. It needs information that is current, authoritative and applicable. A mature context platform should therefore preserve where information came from, who owns it, which version was used, when it was valid and how it supports the conclusion produced.

This matters when context is assembled from multiple sources. Vector search may identify a document as semantically relevant without proving that it represents the authoritative business position. A knowledge graph may connect two concepts without explaining whether the relationship remains valid. Context engineering must preserve the evidence associated with the information it assembles. Without it, context remains information; with it, context becomes part of the control model.

Runtime decisions require runtime evidence

Evidence must therefore remain connected to conclusions and actions at runtime. More logging may improve visibility, but volume does not create meaning. Unless technical events stay connected to business meaning, source authority and decision relevance, the organisation may still be unable to explain why an outcome was reasonable.

The evidence chain is therefore broader than a collection of runtime events. It includes the origin of information, the transformations applied, the authority under which it was used and the relationship between supporting claims and the resulting decision. This is where the challenge intersects with semantic architecture, knowledge architecture, data governance and decision governance. No single observability product can resolve a concern that spans the architecture.

From observability to provenance

This broader evidence challenge has a name: provenance. It is sometimes treated as another form of logging, tracing or data lineage, but these capabilities answer different questions. Observability explains what happened, traceability connects the events involved, and lineage shows where data originated and how it changed. Provenance connects an outcome to the evidence that makes it credible.

Provenance is not a replacement for observability, but an extension of the evidence chain that observability begins. A provenance-aware architecture connects conclusions, recommendations and actions to supporting sources, policies, transformations and authority. It enables an organisation to reconstruct both the technical path and the basis on which the outcome was considered trustworthy.

This does not necessarily require a separate platform. Provenance is better understood as architectural capabilities spanning context services, orchestration, observability, data governance, policy enforcement and decision controls. The objective is not to capture every technical detail, but to preserve the evidence needed to understand and defend an outcome.

Provenance is becoming an enterprise capability

Supporting capabilities often become strategic enterprise capabilities. Data evolved from a technical resource into a business-critical asset requiring ownership, governance and architecture. Lineage became a foundation for compliance, impact analysis and data quality, while observability moved from an operational concern to an essential capability for distributed technology landscapes. Provenance may follow the same trajectory.

The concept itself is not new. Scientific research depends on source references, legal processes on chains of custody, and data governance on lineage and metadata. Agentic AI brings comparable requirements into a much broader range of enterprise activities.

As AI generates recommendations, supports decisions and performs actions, organisations must increasingly demonstrate why an outcome was reasonable. That requires explicit relationships between conclusions, evidence, source authority, policies and ownership. Which source is authoritative when versions differ? How long does evidence remain valid? Can the basis of a decision still be reconstructed after its context changes? These are becoming fundamental architecture and governance questions.

Looking beyond observability

The AI industry is investing heavily in observability, and rightly so. Without visibility, organisations cannot operate complex AI ecosystems effectively. Observability remains essential for quality, safety, performance, cost and operational control, but it may represent only the first stage of a broader architectural shift.

Enterprise architecture first focused on governing systems and later recognised the need to govern data. AI may be introducing a third challenge: governing evidence. As agents become more autonomous and outcomes less directly connected to predefined logic, explaining how a conclusion was produced remains important. Explaining why it deserves trust may become even more important.

This is why provenance may become one of the more significant architectural discussions of the coming years. Not because organisations need another governance framework, but because AI increasingly relies on evidence assembled at runtime. The challenge is no longer simply to explain what happened; it is to establish why we should believe it. Provenance may therefore become a foundational capability for trustworthy enterprise AI.