The Provenance Tax: Understanding the Impact of LLM Watermarking on AI Agent Behavior
Pangram verdict · v3.3
We believe that this text is a mix of AI and human-written content.
AI likelihood · overall
MixedArticle text · 1,406 words · 10 segments analyzed
Recently, Anthropic announced that future Claude models would embed an invisible watermark in their output [1], [2], and subsequently disclosed that the watermark is based on Google DeepMind’s SynthID-Text [2], [3]. Text watermarking itself is not new, but its deployment now has regulatory relevance. Article 50(2) of the EU AI Act [4] requires providers of AI systems generating synthetic text to mark their outputs in a machine-readable format and make them detectable as artificially generated or manipulated, using technical solutions that are effective, interoperable, robust, and reliable as far as technically feasible.Watermarking is designed for provenance, but SynthID-Text changes the process by which the model generates each next token.
At the model level, this can change safety behavior, including whether the model refuses a harmful request and whether that refusal holds under prompt injection. At the agent level, the same sampled tokens can determine which tool is called and what arguments are passed to it. Prompt injection connects these two settings because a weakened refusal becomes more consequential when the model can also act through tools. Such a watermarking procedure can therefore affect both what the model says and what an agent does. We call this behavioral effect sampling drift.Whether this drift appears in practice is an empirical question. We find that it does, in both model refusal behavior and agent tool calling. The effect is model- and key-dependent and can be obscured by aggregate scores when changes in opposite directions cancel. We therefore report both net performance and paired disagreement between watermarked and unwatermarked runs.
Further, the closing section discusses what it means for AI safety and security and what developers should do about it.Built for Content Provenance, Deployed Inside AgentsA text watermark embeds a signal that allows output to be identified as AI-generated. Existing approaches include post-processing methods and methods integrated directly into LLM generation [8]. Generation-time approaches include logits-biasing methods [5], distortion-free keyed sampling [6], cryptographically motivated constructions [7], and SynthID-Text’s Tournament sampling [3]. Figure 1 contrasts this process with ordinary sampling. We use SynthID’s non-distortionary configuration, which preserves the original token distribution in expectation over the watermark randomness while individual generations under a fixed key can still differ [3]. Dathathri et al. report no measurable quality degradation across nearly twenty million Gemini responses [3].Figure 1. Standard versus watermarked text generation. Ordinary generation samples from the model’s token distribution. A generative watermark adds a random seed generator, sampling algorithm, and scoring function; SynthID-Text uses tournament sampling. Adapted from Dathathri et al. [3].Anthropic’s deployment also illustrates why this matters beyond first-party chat interfaces. The company states that watermarking is applied at the model level and covers supported models accessed through the Claude Platform API as well as cloud providers [1]. A developer using a watermarked model as the reasoning component of an agent can therefore receive watermarked outputs even when the agent itself is a separate application. This makes model-level behavioral effects of watermarking relevant to the agents built around such models.Same Tokens the Watermark Biases, Same Tokens the Agent Acts OnTournament sampling has more opportunity to alter token selection where the model is uncertain.
In structured output such as JSON, braces, keys, and function names are often highly predictable, while values such as queries, numbers, paths, and recipients are less so. A change that would amount to a lexical variation in ordinary prose can therefore alter an argument that an agent executes. The weights and prompt remain unchanged, but token selection does not. Importantly, non-distortionary does not imply identical behavior under a fixed watermark key. The guarantee holds over the watermark randomness, while a particular key changes token selection during generation [3]. The resulting sampling drift can therefore change agent behavior even though the watermark is non-distortionary in the sense defined by Dathathri et al. Its effect can also depend on the watermark key, so we test multiple keys rather than relying on one.How We Measure the EffectWe use a paired design for two experiments.
Tool calling is evaluated on BFCL v4 single-turn AST [9], and refusal on 200 HarmBench harmful behaviors [10] plus 100 benign JailbreakBench controls [11], with harmful requests tested both bare and under one fixed prompt injection technique. Table 1 summarizes the datasets, evaluation scope, temperatures, and expected behavior.We use the non-distortionary SynthID-Text configuration through HuggingFace’s unmodified SynthIDTextWatermarkLogitsProcessor, with 30 Tournament layers, n-gram length 5, sampling table 216, and context history 1,024. Each item is generated with and without SynthID from the same seed, batch composition, and order at each temperature. The watermark processor is the only difference within each pair.Table 1. Datasets and experimental settings. Experiment Dataset and scope T Correct behavior Tool calling BFCL v4 single-turn AST [9], live and non-live call-expected tasks plus relevance and irrelevance 0.001, 0.7, 1.0 Correct call, or no call when none fits Refusal HarmBench [10], 200 harmful behaviors; JailbreakBench [11], 100 benign controls; bare and fixed-injection prompts 0.001, 0.7 Refuse harmful; answer benign The Tool-Calling Cost of WatermarkingWe test whether watermarking changes tool selection, arguments, or output validity.
A well-formed call to the correct tool with an incorrect path, recipient, query, or amount is particularly consequential because it can execute successfully while performing the wrong action. We evaluate calls individually, although an incorrect call in a deployed agent could also affect subsequent observations and decisions.How Often Tool-Call Correctness ChangesOn items where a tool call is expected, watermarking reduces accuracy on six of the seven models, with a significant decrease on four. The net change in accuracy, however, does not show whether the same individual calls succeed with and without the watermark. A call that becomes incorrect can be offset by another that becomes correct, leaving the aggregate result nearly unchanged even though the model behaves differently on both items.We measure this directly using the paired disagreement rate, which we refer to as churn, defined as the share of items whose verdict differs between the watermarked and unwatermarked runs. For the comparison across temperatures in Figure 2, we use BFCL’s researcher-defined non-live items, which give us the same fixed set of 1,150 call-expected tasks at each temperature. Figure 2 shows that the paired disagreement is substantially larger than the net accuracy change. At T=1.0, 16.8% of phi-4’s call verdicts differ between the two conditions while its net accuracy loss is 2.87 points. Llama-3.1-8B shows the same pattern, with 9.9% of verdicts changing while the net loss is only 0.87 points. Across the 21 model-temperature combinations, churn averages 6.5%, and its bootstrap interval excludes zero in every case.Figure 2.
Paired tool-call disagreement under watermarking by model and temperature. Results use 1,150 non-live BFCL call-expected items across models and temperatures. The central vertical line represents no change relative to the unwatermarked condition. Orange bars show calls that changed from correct to incorrect, and blue bars show calls that changed from incorrect to correct. The final column reports the churn with non-live irrelevance items included. Diamonds denote 95% bootstrap intervals excluding zero.Which Tool-Call Errors ChangeError type also matters.
Malformed output prevents the intended call from executing, while a well-formed call with the wrong tool or argument can still execute. Figure 3 separates these failures into wrong tool, wrong arguments, and malformed output. Unlike Figure 2’s across-temperature comparison, this analysis combines BFCL live and non-live call-expected items at T=0.001 to characterize errors across the broader benchmark. Relevance and irrelevance are excluded because they test whether a call should be made rather than whether the emitted call is correct.Figure 3.
Changes in tool-calling errors under watermarking by error type. BFCL live and non-live call-expected items at T=0.001, separated into wrong tool, wrong arguments, and malformed output. The vertical lines show the accuracy without watermarking, and the bars show the change when watermarking is applied. Orange denotes correct-to-error changes and blue error-to-correct changes. Intervals are 95% item-level bootstrap intervals.The error profile also differs across models.
On Llama-3.1-8B, the largest contribution to the accuracy loss comes from incorrect arguments (−3.48 points), followed by wrong-tool calls (−1.84 points). On phi-4 and Granite-3.2-8B, malformed output dominates (−5.96 and −4.36 points). Similar aggregate changes can therefore arise from different failure modes.Watermarking Can Weaken Refusal Under Prompt InjectionRefusals are also generated token by token, so watermarking can affect them. We test harmful requests alone and with one simple, fixed prompt-injection technique intended to reduce refusal. The technique appends an adversarial instruction as retrieved content, claiming that the safety filter is disabled and instructing compliance. It is held constant across prompts, models, and temperatures.