Semantic Differentiation for Tackling Challenges in Watermarking Low-Entropy Constrained Generation Outputs
Le, Ritter, Goyal · cs.CR,cs.LG · 2026-08-14 · 原文 · 证据态 未学
We demonstrate that while the current approaches for language model watermarking are effective for open-ended generation, they are inadequate at watermarking LM outputs for constrained generation tasks with low-entropy output spaces. Therefore, we devise SeqMark, a sequence-level watermarking algorithm with semantic differentiation that balances output quality, watermark detectability, and imperceptibility. It improves on the shortcomings of token-level watermarking algorithms that cause under-utilization of the sequence-level entropy available for constrained generation tasks. Moreover, we identify and improve upon the problem of region collapse, a different failure mode associated with prior sequence-level watermarking algorithms. This occurs because the pseudorandom partitioning of semantic space for watermarking in these approaches causes all high-probability outputs to collapse into either invalid or valid regions, leading to a trade-off in output quality and watermarking effectiveness. Instead, SeqMark differentiates the high-probable output subspace and partitions it into valid and invalid regions, ensuring the even spread of high-quality outputs among all the regions. On va
讲义
讲义·推断 依据「原文」自动生成的结构化摘要(推断),非原文表述;以原文为准。
1. 人话版
We demonstrate that while the current approaches for language model watermarking are effective for open-ended generation, they are inadequate at watermarking LM outputs for constrained generation tasks with low-entropy output spaces.
Therefore, we devise SeqMark, a sequence-level watermarking algorithm with semantic differentiation that balances output quality, watermark detectability, and imperceptibility.
2. 领域脉络
本文类目:cs.CR、cs.LG,属于其所在研究脉络的最新进展。
3. 机制拆解
It improves on the shortcomings of token-level watermarking algorithms that cause under-utilization of the sequence-level entropy available for constrained generation tasks.
This occurs because the pseudorandom partitioning of semantic space for watermarking in these approaches causes all high-probability outputs to collapse into either invalid or valid regions, leading to a trade-off in output quality and watermarking effectiveness.
4. 证据与数字
摘要未给出量化结果——留意原文的实验与数据。
5. 反例与边界
Moreover, we identify and improve upon the problem of region collapse, a different failure mode associated with prior sequence-level watermarking algorithms.
6. 跨领域连接与意外收获
横跨 2 个类目(cs.CR、cs.LG),关注其在你兴趣板块间的迁移面。
7. 可复用方法
把本文机制与你手头项目对照,找一个两周内能验证的最小实验。
8. 术语表
精读时把不熟的术语记入此处,作为下次回忆的锚点。
主动回忆
先合上内容自己复述,再点「显示」核对,然后如实自评。评分即时进 FSRS 排程。
讲义·推断 依据「原文」自动生成的结构化摘要(推断),非原文表述;以原文为准。
1. 人话版
We demonstrate that while the current approaches for language model watermarking are effective for open-ended generation, they are inadequate at watermarking LM outputs for constrained generation tasks with low-entropy output spaces.
Therefore, we devise SeqMark, a sequence-level watermarking algorithm with semantic differentiation that balances output quality, watermark detectability, and imperceptibility.
2. 领域脉络
本文类目:cs.CR、cs.LG,属于其所在研究脉络的最新进展。
3. 机制拆解
It improves on the shortcomings of token-level watermarking algorithms that cause under-utilization of the sequence-level entropy available for constrained generation tasks.
This occurs because the pseudorandom partitioning of semantic space for watermarking in these approaches causes all high-probability outputs to collapse into either invalid or valid regions, leading to a trade-off in output quality and watermarking effectiveness.
4. 证据与数字
摘要未给出量化结果——留意原文的实验与数据。
5. 反例与边界
Moreover, we identify and improve upon the problem of region collapse, a different failure mode associated with prior sequence-level watermarking algorithms.
6. 跨领域连接与意外收获
横跨 2 个类目(cs.CR、cs.LG),关注其在你兴趣板块间的迁移面。
7. 可复用方法
把本文机制与你手头项目对照,找一个两周内能验证的最小实验。
8. 术语表
精读时把不熟的术语记入此处,作为下次回忆的锚点。