ADP 前沿学习

← 板块一 · 研究前沿

Semantic Differentiation for Tackling Challenges in Watermarking Low-Entropy Constrained Generation Outputs

Le, Ritter, Goyal · cs.CR,cs.LG · 2026-08-14 · 原文 · 证据态 未学

We demonstrate that while the current approaches for language model watermarking are effective for open-ended generation, they are inadequate at watermarking LM outputs for constrained generation tasks with low-entropy output spaces. Therefore, we devise SeqMark, a sequence-level watermarking algorithm with semantic differentiation that balances output quality, watermark detectability, and imperceptibility. It improves on the shortcomings of token-level watermarking algorithms that cause under-utilization of the sequence-level entropy available for constrained generation tasks. Moreover, we identify and improve upon the problem of region collapse, a different failure mode associated with prior sequence-level watermarking algorithms. This occurs because the pseudorandom partitioning of semantic space for watermarking in these approaches causes all high-probability outputs to collapse into either invalid or valid regions, leading to a trade-off in output quality and watermarking effectiveness. Instead, SeqMark differentiates the high-probable output subspace and partitions it into valid and invalid regions, ensuring the even spread of high-quality outputs among all the regions. On va

🔮 让 ChatGPT 全网深度追问

讲义

讲义·推断 依据「原文」自动生成的结构化摘要(推断),非原文表述;以原文为准。

1. 人话版

We demonstrate that while the current approaches for language model watermarking are effective for open-ended generation, they are inadequate at watermarking LM outputs for constrained generation tasks with low-entropy output spaces.

Therefore, we devise SeqMark, a sequence-level watermarking algorithm with semantic differentiation that balances output quality, watermark detectability, and imperceptibility.

2. 领域脉络

本文类目:cs.CR、cs.LG,属于其所在研究脉络的最新进展。

3. 机制拆解

It improves on the shortcomings of token-level watermarking algorithms that cause under-utilization of the sequence-level entropy available for constrained generation tasks.

This occurs because the pseudorandom partitioning of semantic space for watermarking in these approaches causes all high-probability outputs to collapse into either invalid or valid regions, leading to a trade-off in output quality and watermarking effectiveness.

4. 证据与数字

摘要未给出量化结果——留意原文的实验与数据。

5. 反例与边界

Moreover, we identify and improve upon the problem of region collapse, a different failure mode associated with prior sequence-level watermarking algorithms.

6. 跨领域连接与意外收获

横跨 2 个类目(cs.CR、cs.LG),关注其在你兴趣板块间的迁移面。

7. 可复用方法

把本文机制与你手头项目对照,找一个两周内能验证的最小实验。

8. 术语表

精读时把不熟的术语记入此处,作为下次回忆的锚点。

主动回忆

先合上内容自己复述,再点「显示」核对,然后如实自评。评分即时进 FSRS 排程。