AutoMark: Enabling Autoresearch to Discover Better LLM Watermarks
Thibaud Gloaguen, Robin Staab
September 28, 2026
LLM watermarking embeds a signal invisible to humans in model outputs, and has become the standard for tracing AI-generated text: regulators mandate it, and major providers deploy it. Yet, recent progress has mostly come from refining existing designs, at a pace capped by that of human researchers. Autoresearch, where agents propose, implement, and evaluate ideas in a loop, is a natural way to accelerate it.
Watermarking, however, is harder to verify than, e.g., mathematics, where a proof can be checked by Lean. Detectability, robustness, and quality trade off, so an improvement on a single metric need not make a better watermark. Agents may also exploit gaps in the evaluation, e.g., by inflating robustness through flagging more human texts as watermarked. In our new paper, we show that with an adequate harness, autoresearch discovers watermarks that are trustworthy, explainable, and Pareto-superior to existing schemes. In this post, we summarize how, and let you explore the components that the agents discovered.
What makes a watermark valid?
Soundness over the key and given a key. Each bar is the false positive rate on human texts of one watermark key (not to scale). Left: averaged over the key, the naive scheme flags any text with probability 1/100, but a key that puts “.” in the first class flags most human texts. Right: a scheme that is sound given a key, for which at most a fraction β of the keys have a false positive rate above α.
Before agents can search for better watermarks, we first need to formally define what makes a watermarking scheme valid, so that every submission can be checked automatically. Without such a definition, autoresearch is easy to fool: in earlier versions of our harness, agents designed schemes that appeared to outperform prior work, yet had an abnormally high false positive rate.
All schemes in this work are distortion-free. At every step of generation, a hash of the preceding tokens and a private key seed a pseudo-random score for every candidate token, and the watermark samples the next token so that it correlates with these scores, while preserving the next-token distribution in expectation over the key. The detector recomputes the scores of a text and returns a p-value under the null hypothesis that the text is not watermarked.
Most prior schemes are sound over the key: for any text, the probability over the key that its p-value falls below α is at most α. This alone does not make a scheme reliable. Consider a deliberately naive scheme that uses the key to split the vocabulary into a hundred classes, and flags a text if its last token lies in the first class. Averaged over the key, it flags any text with probability 1/100, but for a key that puts “.” in the first class, most human texts are flagged as watermarked (left of the figure above).
We therefore introduce a second requirement, soundness given a key: for every (non-degenerate) human text distribution, the probability of sampling a key whose false positive rate exceeds α must be at most β (we use β = 5%), as illustrated on the right of the figure above. Any scheme that is sound over the key becomes provably sound given a key if we multiply its p-values by 1/β, at the cost of detection power. We turn both properties, along with distortion-freeness, into statistical tests that run automatically on every submitted scheme.
The autoresearch harness
The autoresearch loop. An agent iterates on a watermarking scheme in an isolated sandbox. Submissions are first verified and, upon passing, evaluated. Either way, a new agent then takes over with access to all prior results.
The agent has a dedicated GPU, Internet access, our vLLM watermarking library, and the implementations and evaluations of AAR, SynthID, and TextSeal. When it submits a scheme, a private pipeline first runs our three tests. If the scheme passes, it is evaluated on detectability (TPR at 1% FPR), on robustness to four attacks (word deletion, synonym substitution, back-translation, and paraphrasing), and on quality (the distance to the unwatermarked model in perplexity and in the diversity of replies, measured with Self-BLEU). The results are then added to the environment, and a new agent takes over with access to all prior results and code. The first agent has four submissions, and every subsequent agent has a single one, a setup that we found leads to more creative solutions and less hyperparameter tuning.
Both the verification and the evaluation are private. In an earlier version of our harness with public verification, an agent remapped its p-values with a threshold chosen just below the fifth-smallest p-value of the public human texts, and documented it in its submission.
Agents outperform prior work
We ran four research agents until their submission budget was exhausted: GPT-6 Astra (xhigh and low reasoning), Opus 5 (high), and Gemini-3.8 Flash (high). This yields 52 evaluated schemes, at an API cost between $13.81 (Gemini-3.8 Flash) and $435.96 (Opus 5) per agent. All agents improve their submissions over time, and outperform the baselines in most cases. Two baselines, AAR and SynthID, fail our soundness given a key test in the evaluated configurations, which means that their robustness is slightly inflated.
Submission history of each agent. Left: robustness to paraphrasing (TPR@1%FPR). Right: quality, as the distance to the unwatermarked model. Each dot is a submission and the solid lines are the best value so far. Baselines are in italic, and a parenthesis means that the baseline fails our verification.
The research processes of GPT-6 Astra and Opus 5 were similar: after first tweaking the TextSeal baseline, both proposed a range of novel and creative watermarking ideas. In contrast, all submissions of Gemini-3.8 Flash tuned the hyperparameters of TextSeal and did not result in any new ideas.
Two new state-of-the-art schemes
To mitigate overfitting to our evaluation, we select two agent schemes and evaluate them under a held-out protocol: 1,000 English and 1,000 Chinese prompts from WildChat, DeepSeek-v4 Flash as the back-translation and paraphrasing attacker, and replies from both Llama-3.1-8B and Qwen3.8-27B. ConcordMark combines Lexical group units, Ladder context seeding, a Prelude, six Channels, and the Gumbel race; its high variant uses a model forward pass at detection. RotationFirst combines Lexical group units, Global first use seeding, the Rotation geometry, the ResidualRace, and the Predictive Gamma detector.
Held-out evaluation. Robustness axes show the TPR@1%FPR under each attack, and quality axes show the relative deviation (in %) from the unwatermarked model in perplexity and in bigram and trigram Self-BLEU. Outward is better on every axis, and each axis is rescaled to the plotted schemes, with its range (from center to rim) under its label. We report the worst TPR and the median quality deviation over five watermark keys, and the clean TPR in the legend.
ConcordMark (high) is Pareto-better than all baselines: on almost all metrics, it outperforms the best baseline. RotationFirst has a diversity very close to that of the unwatermarked model, while its detection and robustness remain comparable to those of the baselines, which makes it suitable for model providers that want a strong watermark with the smallest impact on quality. For both schemes, the results of the autoresearch evaluation transfer to new prompts, models, and attackers.
What did the agents discover?
Agents may produce results faster than humans can understand them. To extract their key ideas, we manually decompose every scheme into five components: the four components of most watermarks (context seeding, score, logits transformation, and detector), and the watermark unit, a generalization introduced by the agents in which the watermark applies to classes of tokens rather than to tokens. We then ablate each component against a simple core scheme in the style of AAR (token units, a two-token context, the Gumbel race, and the Gamma detector), swapping one component at a time.
To improve quality, the agents consistently add a small amount of true randomness to the watermark score. Their Gaussian copula mixes the keyed score with fresh Gaussian noise, and with all other components fixed, it outperforms TextSeal’s random key switching along all evaluated dimensions. The agents also discovered fundamentally new ideas, such as aligning watermark scores with a random direction drawn for every request (the private geometries), on which RotationFirst builds. Finally, they extend the Gumbel race with residual clocks (ResidualRace), which improves robustness against every attack at a slight cost to quality, and design detectors that group the evidence by context or use a model forward pass.
Explore the components
Each chip below is one of the 42 components of the paper, 39 of which are new, arranged by stage. Select a component to read its description and the results of every configuration in which we evaluated it, and plot up to four configurations on the radar. The dashed polygon is the core scheme, which every configuration modifies. The Unit switch shows the results with token units or with the Lexical group, and the examples reproduce findings from the paper.
Components of the agents' schemes. Colors indicate the family (fresh race, private geometry, or skeleton), within which components are interchangeable, and white components are compatible with every family. Skeleton schemes reuse the fresh-race scores and detectors (green edge), and components marked * come from prior work. All values come from our ablation study: a single watermark key, 1,000 Llama-3.1-8B replies to ELI5 prompts, and p-values multiplied by 1/β = 20 so that every configuration is provably sound given a key.
Summary
In this post, we presented AutoMark, which enables autoresearch for LLM watermarking with rigorous validity criteria, statistical tests to verify them, and a tailored harness. We found that frontier agents go beyond hyperparameter tuning and introduce novel ideas, with two schemes, ConcordMark and RotationFirst, outperforming prior work on held-out data. In autoresearch, the key challenge is to build a verification that agents cannot exploit and that faithfully captures the desired property. There are many desirable properties we did not include in this work, e.g., security, radioactivity, or the multi-bit setting, and we believe that future research should focus not only on new schemes but also on how to build verifiers and evaluations for those properties.