Appropriateness (1-5)
Does this paper fit in the event? Both empirical and theoretical results are welcome.
5: Certainly.
4: Probably.
3: Unsure.
2: Probably not.
1: Certainly not.
Clarity (1-5)
For the reasonably well-prepared reader, is it clear what was done and why? Is the paper well-written and well-structured?
5 = Very clear.
4 = Understandable by most readers.
3 = Mostly understandable to me with some effort.
2 = Important questions were hard to resolve even with effort.
1 = Much of the paper is confusing.
Originality / Innovativeness (1-5)
How original is the approach? Does this paper break new ground in topic, methodology, or content? How exciting and innovative is the research it describes?
Note that a paper could score high for originality even if the results do not show a convincing benefit.
5 = Surprising: Noteworthy new problem, technique, methodology, or insight.
4 = Creative: Relatively few people in our community would have put these ideas together.
3 = Somewhat conventional: A number of people could have come up with this if they thought about it for a while.
2 = Rather boring: Obvious, or a minor improvement on familiar techniques.
1 = Significant portions have actually been done before or done better.
Soundness / Correctness (1-5)
First, is the technical approach sound and well-chosen? Second, can one trust the claims of the paper — are they supported by proper experiments, proofs, or other argumentation?
5 = The approach is very apt, and the claims are convincingly supported.
4 = Generally solid work, though I have a few suggestions about how to strengthen the technical approach or evaluation.
3 = Fairly reasonable work. The approach is not bad, and at least the main claims are probably correct, but I am not entirely ready to accept them (based on the material in the paper).
2 = Troublesome. There are some ideas worth salvaging here, but the work should really have been done or evaluated differently, or justified better.
1 = Fatally flawed.
Meaningful Comparison (1-5)
Does the author make clear where the problems and methods sit with respect to existing literature? Are the references adequate? Are any experimental results meaningfully compared with the best prior approaches?
5 = Precise and complete comparison with related work. Good job given the space constraints.
4 = Mostly solid bibliography and comparison, but I have some suggestions.
3 = Bibliography and comparison are somewhat helpful, but it could be hard for a reader to determine exactly how this work relates to previous work.
2 = Only partial awareness and understanding of related work, or a flawed empirical comparison.
1 = Little awareness of related work, or lacks necessary empirical comparison.
Thoroughness (1-5)
Does this paper have enough substance, or would it benefit from more ideas or results?
Note that this question mainly concerns the amount of work; its quality is evaluated in other categories.
5 = Contains more ideas or results than most publications in this conference; goes the extra mile.
4 = Represents an appropriate amount of work for a publication in this conference. (most submissions)
3 = Leaves open one or two natural questions that should have been pursued within the paper.
2 = Work in progress. There are enough good ideas, but perhaps not enough results yet.
1 = Seems thin. Not enough ideas here for a full-length paper.
Impact of Ideas or Results (1-5)
How significant is the work described? If the ideas are novel, will they also be useful or inspirational? If the results are sound, are they also important? Does the analysis in the paper bring new insights into the nature of the problem?
5 = Will affect the field by altering other people’s choice of research topics or basic approach.
4 = Some of the ideas or results will substantially help other people’s ongoing research.
3 = Interesting but not too influential. The work will be cited, but mainly for comparison or as a source of minor contributions.
2 = Marginally interesting. May or may not be cited.
1 = Will have no impact on the field.
Recommendation (1-5)
There are many good submissions competing for slots at this event; how important is it to feature this one? Will people learn a lot by reading this paper or seeing it presented?
In deciding on your ultimate recommendation, please think over all your scores above. But remember that no paper is perfect, and remember that we want a conference full of interesting, diverse, and timely work. If a paper has some weaknesses, but you really got a lot out of it, feel free to fight for it. If a paper is solid but you could live without it, let us know that you’re ambivalent. Remember also that the author has a couple of weeks to address reviewer comments before the camera-ready deadline.
Should the paper be accepted or rejected?
5 = Exciting: I’d fight to get it accepted
4 = Worthy: I would like to see it accepted
3 = Borderline: I’m ambivalent about this one
2 = Mediocre: I’d rather not see it in the conference
1 = Poor: I’d fight to have it rejected
Reviewer Confidence (1-5)
5 = Positive that my evaluation is correct. I read the paper very carefully and am very familiar with related work.
4 = Quite sure. I tried to check the important points carefully, and checked for uncited prior work. It’s unlikely, though conceivable, that I missed something that should affect my ratings.
3 = Pretty sure, but there’s a chance I missed something. Although I have a good feel for this area in general, I did not carefully check the paper’s details, e.g., the math, experimental design, or novelty.
2 = Willing to defend evaluation, but it is fairly likely that I missed some details, didn’t understand some central points, or can’t be sure about the novelty of the work.
1 = Not my area, or paper is very hard to understand. My evaluation is just an educated guess.
Detailed Comments
The paper presents an investigation of how back-translation attack reduce watermark detectability. This seems to be especially the case when the back-translation attack is done by passing trough a low-resource language w.r.t. to language with more resources available.
This is an actual underexplored intersection, watermark robustness for Italian LLMs, with a commendably broad experimental sweep across six models, four watermarking families, and three resource tiers for the pivot language. The core empirical finding, that lower-resource pivot languages more thoroughly destroy the watermark signal, is intuitive but valuably confirmed here, and the SIR/SemaMark vs. KGW/KTH robustness contrast is a nice secondary result.
However there are a few issues I’d like to raise:
On the watermarking efficacy the main body of the paper reports only a KDE estimate of z-score. No AUROC is reported, nor FPR@1%TPR or other filtering estimates. So while a reduction of efficacy after the back-translation is plainly visualized in the KDE, we have no quantitative measure of how bad this becomes, and how good the watermarking detection really was in the start.
On fluency, we only have an estimate using perplexity. I think this is a deeply flawed measure. Perplexity increases naturally as watermarking is applied, as a forceful shift in model’s output distribution is applied. This is also true for KTH which was presented as a watermarking techniques where “the generated text adheres strictly to the original language model distribution”, however your Table 3 directly contradicts this, as is constantly one of the worst offender per model in terms of perplexity increase, even without back-translating. In general, with watermarking, an increase in perplexity is expected, as the model isn’t generating its own training distribution as is being forced to produce certain tokens w.r.t. the rest. This increases perplexity naturally, but that this also means that human-perceived fluency goes down, is a stretch that needs to be proven, not assumed. This is even more the case for the judging of the Machine translated text: it is unclear from 3.4 which model computes perplexity for back-translated text. If it is the original generative LLM, this comparison is confounded, since the text was produced by X-ALMA, not that LLM. The authors should clarify. If that is the case, the authors should also explain why anything different than a higher perplexity should be expected.
The case of Modello Italia and Velvet is also really strange, and should’ve been investigated further. I speculate that the low perplexity could have been for outlier generation or degenerate generations, and error analysis on specifically these models should have been done.
In general, for both KDE plots and perplexity values, no confidence intervals nor statistical testing has been reported, so I think some of the conclusion on which watermarking technique seem more robust are overstated.
Finally, even if we accept for true that a higher perplexity implicates a lower fluency, I would again argue that if that is the case, then this type of attack are not of concern. If back-translating creates obviously low quality text, we don’t need watermarking to detect them, it will be plainly obvious to humans and the problem of detection resolves itself. More broadly, the practical threat model deserves more discussion: a malicious actor can simply use an unwatermarked open-weight LLM directly, making the back-translation attack a niche concern in the current landscape of MGT.
I think that for fluency estimation, even a small human acceptability of the text before and after each backtranslating would have sufficed and seriously strengthen the paper results.
Questions for Authors
The box below can be used to ask specific questions of the authors. If your conference allows authors to respond to reviews - before making final acceptance decisions - then these questions will help guide the authors in drafting their response to your review. If your conference is not running this type of process, your questions may never receive answers. Nonetheless, you still may wish to use this area to posit hypothetical questions to the authors (e.g., to identify some major problems you have with the work) - and to separate these questions from the main body of your comments.
Confidential Comments for Committee
You may wish to withhold some comments from the authors, and include them solely for the committee’s internal use. For example, you may want to express a very strong (negative) opinion on the paper, which might offend the authors in some way. Or, perhaps you wish to write something which would expose your identity to the authors. If you wish to share comments of this nature with the committee, this is the place to put them.