PROVENANCE

Status: Established interdisciplinary concept

Terminological origin: Established terminology across data management, scientific workflows, knowledge representation, information systems, and related fields

Relations: related to → Data Lineage · Traceability · Semantic Provenance · Semantic Preservation · Semantic Drift · Knowledge Representation · Coherence / Incoherence · QualComputing

Concise definition

Provenance concerns the traceable history through which an entity, item of data, representation, or claim came to exist in its present form, including its origins and the processes involved in its derivation.

In information systems, provenance is broader than source attribution. It may describe not only where information originated, but also how it was generated, used, transformed, revised, combined, or derived, and which agents or activities participated in that history.

Provenance therefore addresses not only:

Where did this information come from?

but also:

How did it become what it is now?

Conceptual scope

Provenance has become an important concept across database research, scientific workflows, knowledge representation, information systems, and artificial intelligence.

In database systems, provenance can identify source data contributing to a result and reconstruct derivational relationships between inputs and outputs. In scientific workflows, it can document processes and intermediate states involved in producing data. In knowledge representation, provenance can associate informational entities with activities, agents, evidence, and derivational relations.

The World Wide Web Consortium (W3C) PROV framework provides one of the principal standards for representing provenance. Its core architecture distinguishes Entities, Activities, and Agents, together with relations describing generation, use, derivation, attribution, revision, and related processes.

Provenance should therefore not be reduced to citation or source identification.

It concerns the history through which information becomes what is subsequently interpreted or used.

Key distinctions

Source vs. provenance. A source identifies an origin. Provenance may additionally describe the derivational history connecting that origin with a later representation.

Provenance vs. data lineage. The concepts substantially overlap. Data lineage often emphasises derivational paths connecting outputs with contributing inputs, while provenance may encompass a broader history involving activities, agents, attribution, and contextual information. The boundary is not universally standardised.

Provenance vs. traceability. Traceability is the broader capacity to follow relations among objects, processes, states, or transformations. Provenance constitutes one important form of informational traceability.

Semantic provenance vs. provenance semantics. Semantic provenance is established terminology for provenance enriched with semantic or domain-specific information. Provenance semantics can instead concern the formal interpretation and propagation of provenance information.

Provenance vs. semantic preservation. Provenance describes derivational history. Semantic preservation concerns whether specified semantic properties survive transformation. A transformation may be completely traceable without every property relevant to interpretation necessarily being preserved.

Across disciplines

Data management and scientific workflows. Provenance supports the reconstruction of derivational histories, contributing to reproducibility, validation, auditing, and understanding how informational products were produced.

Knowledge representation and the Semantic Web. Provenance can be formally represented through ontologies and graphs connecting entities, activities, agents, and derivational relations. Semantic provenance further demonstrates that provenance can incorporate domain-specific semantic information.

Formal and computational systems. Provenance can participate directly in computational semantics, including formal accounts of how contributions from source data propagate through transformations and queries.

Artificial intelligence. Provenance becomes particularly significant when information passes through retrieval, summarisation, generation, tool use, or interactions between artificial and human agents. In such systems, identifying a final source may represent only one part of reconstructing the informational trajectory.

Conceptual Issues / Points of Debate

The relationship between provenance and lineage remains terminologically variable across research communities.

More importantly, traceability should not be confused with epistemic or semantic adequacy. Knowing precisely how a representation was produced does not by itself establish that it is true, reliable, unbiased, or appropriate for a particular use.

Contemporary provenance frameworks are capable of representing rich transformations and can be extended with domain-specific semantic information. The relevant limitation is therefore not that provenance is intrinsically unable to represent meaning-related properties.

A different question arises when information undergoes successive transformations:

A derivational history may remain fully traceable even when properties relevant to interpretation change along the trajectory.

Existing research on semantic preservation, factuality, modality, and semantic drift already addresses significant parts of this problem. These fields must therefore be distinguished from, rather than absorbed into, provenance.

A further distinction is particularly important:

Information loss does not necessarily imply semantic failure.

Legitimate abstraction, translation, or summarisation may transform information substantially while preserving what matters for the intended use.

Conversely, a small transformation may sometimes have substantial interpretive consequences.

The magnitude of a transformation and the magnitude of its interpretive consequences should therefore not be assumed to coincide.

BSI perspective

From a BSI perspective, provenance becomes particularly significant when information moves across heterogeneous representational systems and undergoes successive transformations before reaching interpretation or use.

BSI does not propose a new definition of provenance. Nor does it claim that existing provenance systems are limited to source attribution. Contemporary frameworks already provide rich resources for representing derivational histories, transformations, agents, activities, and semantic information.

The BSI research question begins after these capacities have been acknowledged.

A derivational history may remain traceable while distinctions or relations important for interpretation change during transformation.

This leads to a simple diagnostic question:

What happened to meaning along the way?

More precisely:

Across a derivational trajectory, which distinctions and relations relevant to interpretation are preserved, altered, suppressed, introduced, or transferred?

The question connects provenance with a broader BSI investigation into semantic coherence and incoherence.

The objective is not to preserve every feature of an original representation. Translation, abstraction, compression, and reformulation necessarily transform information.

The relevant question is whether the distinctions and relations necessary for adequate interpretation remain appropriately represented.

From a QualComputing perspective, provenance therefore contributes to a broader concern:

Provenance tells us how information arrived. It does not, by itself, tell us whether what mattered for interpretation survived the journey.

BSI is investigating this relation between derivational history, semantic transformation, interpretive relevance, and semantic coherence.

At this stage, this is a research direction rather than a proposed redefinition of provenance.

Sources

Buneman, P., Khanna, S., & Tan, W. C. (2001). “Why and Where: A Characterization of Data Provenance.” In Database Theory, ICDT 2001, Lecture Notes in Computer Science, 1973, 316–330. Digital Object Identifier (DOI): 10.1007/3-540-44503-X_20.

Source contribution: Foundational database-provenance research distinguishing different dimensions of the origins and derivation of data.

BSI relevance: Establishes that provenance cannot be reduced to source citation and provides an essential historical basis for BSI’s investigation of informational trajectories.

Simmhan, Y. L., Plale, B., & Gannon, D. (2005). “A Survey of Data Provenance in e-Science.” ACM SIGMOD Record, 34(3), 31–36. DOI: 10.1145/1084805.1084812.

Source contribution: Surveys provenance as derivation history in scientific computing and its roles in reproducibility, validation, and reuse.

BSI relevance: Supports the distinction between provenance as informational history and narrower forms of source attribution.

Green, T. J., Karvounarakis, G., & Tannen, V. (2007). “Provenance Semirings.” Proceedings of the Twenty-Sixth ACM SIGMOD-SIGACT-SIGART Symposium on Principles of Database Systems, 31–40.

Source contribution: Provides a formal algebraic framework for representing how provenance information propagates through relational queries.

BSI relevance: Demonstrates the formal computational richness of provenance and helps distinguish provenance semantics from semantic provenance.

Sahoo, S. S., Sheth, A. P., & Henson, C. A. (2008). “Semantic Provenance for eScience: Managing the Deluge of Scientific Data.” IEEE Internet Computing, 12(4), 46–54. DOI: 10.1109/MIC.2008.86.

Source contribution: Develops semantic provenance through the combination of provenance information and domain-specific ontologies.

BSI relevance: Establishes semantic provenance as existing terminology and prevents BSI from claiming novelty merely by associating provenance with semantic information.

World Wide Web Consortium. (2013). PROV-DM: The PROV Data Model and PROV-O: The PROV Ontology. W3C Recommendations, 30 April 2013.

Source contribution: Provides a standard conceptual and ontological framework for representing entities, activities, agents, derivations, attribution, and related provenance structures.

BSI relevance: Establishes the expressive breadth of contemporary provenance and provides a reference point for distinguishing representational capability from semantic evaluation.

Gulla, J. A., Solskinnsbakk, G., Myrseth, P., Haderlein, V., & Cerrato, O. (2010). “Semantic Drift in Ontologies.” Proceedings of the Sixth International Conference on Web Information Systems and Technologies, 13–20. DOI: 10.5220/0002788800130020.

Source contribution: Examines semantic drift associated with changes in concepts and conceptual relations.

BSI relevance: Provides an established neighboring concept and helps distinguish semantic change from semantic incoherence.

Yao, J., Qiu, H., Zhao, J., Min, B., & Xue, N. (2021). “Factuality Assessment as Modal Dependency Parsing.” Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, 1540–1550. DOI: 10.18653/v1/2021.acl-long.122.

Source contribution: Models factuality through relations among events, informational sources, and degrees of certainty.

BSI relevance: Demonstrates that modality and epistemic attribution already have computational treatments and therefore helps delimit the legitimate scope of any BSI extension.

Related Concepts

Related to → Data Lineage. Traceability . Semantic Provenance. Knowledge Representation. Semantic Preservation. Semantic Drift. Factuality. Evidence Attribution. Coherence / Incoherence. QualComputing

Distinguished from → Source Attribution. Factual Correctness. Semantic Preservation. Semantic Coherence

Last revised

September 2026.