Appropriateness (1-5)
Does this paper fit in the event? Both empirical and theoretical results are welcome.
5: Certainly.
4: Probably.
3: Unsure.
2: Probably not.
1: Certainly not.
Clarity (1-5)
For the reasonably well-prepared reader, is it clear what was done and why? Is the paper well-written and well-structured?
5 = Very clear.
4 = Understandable by most readers.
3 = Mostly understandable to me with some effort.
2 = Important questions were hard to resolve even with effort.
1 = Much of the paper is confusing.
Originality / Innovativeness (1-5)
How original is the approach? Does this paper break new ground in topic, methodology, or content? How exciting and innovative is the research it describes?
Note that a paper could score high for originality even if the results do not show a convincing benefit.
5 = Surprising: Noteworthy new problem, technique, methodology, or insight.
4 = Creative: Relatively few people in our community would have put these ideas together.
3 = Somewhat conventional: A number of people could have come up with this if they thought about it for a while.
2 = Rather boring: Obvious, or a minor improvement on familiar techniques.
1 = Significant portions have actually been done before or done better.
Soundness / Correctness (1-5)
First, is the technical approach sound and well-chosen? Second, can one trust the claims of the paper — are they supported by proper experiments, proofs, or other argumentation?
5 = The approach is very apt, and the claims are convincingly supported.
4 = Generally solid work, though I have a few suggestions about how to strengthen the technical approach or evaluation.
3 = Fairly reasonable work. The approach is not bad, and at least the main claims are probably correct, but I am not entirely ready to accept them (based on the material in the paper).
2 = Troublesome. There are some ideas worth salvaging here, but the work should really have been done or evaluated differently, or justified better.
1 = Fatally flawed.
Meaningful Comparison (1-5)
Does the author make clear where the problems and methods sit with respect to existing literature? Are the references adequate? Are any experimental results meaningfully compared with the best prior approaches?
5 = Precise and complete comparison with related work. Good job given the space constraints.
4 = Mostly solid bibliography and comparison, but I have some suggestions.
3 = Bibliography and comparison are somewhat helpful, but it could be hard for a reader to determine exactly how this work relates to previous work.
2 = Only partial awareness and understanding of related work, or a flawed empirical comparison.
1 = Little awareness of related work, or lacks necessary empirical comparison.
Thoroughness (1-5)
Does this paper have enough substance, or would it benefit from more ideas or results?
Note that this question mainly concerns the amount of work; its quality is evaluated in other categories.
5 = Contains more ideas or results than most publications in this conference; goes the extra mile.
4 = Represents an appropriate amount of work for a publication in this conference. (most submissions)
3 = Leaves open one or two natural questions that should have been pursued within the paper.
2 = Work in progress. There are enough good ideas, but perhaps not enough results yet.
1 = Seems thin. Not enough ideas here for a full-length paper.
Impact of Ideas or Results (1-5)
How significant is the work described? If the ideas are novel, will they also be useful or inspirational? If the results are sound, are they also important? Does the analysis in the paper bring new insights into the nature of the problem?
5 = Will affect the field by altering other people’s choice of research topics or basic approach.
4 = Some of the ideas or results will substantially help other people’s ongoing research.
3 = Interesting but not too influential. The work will be cited, but mainly for comparison or as a source of minor contributions.
2 = Marginally interesting. May or may not be cited.
1 = Will have no impact on the field.
Recommendation (1-5)
There are many good submissions competing for slots at this event; how important is it to feature this one? Will people learn a lot by reading this paper or seeing it presented?
In deciding on your ultimate recommendation, please think over all your scores above. But remember that no paper is perfect, and remember that we want a conference full of interesting, diverse, and timely work. If a paper has some weaknesses, but you really got a lot out of it, feel free to fight for it. If a paper is solid but you could live without it, let us know that you’re ambivalent. Remember also that the author has a couple of weeks to address reviewer comments before the camera-ready deadline.
Should the paper be accepted or rejected?
5 = Exciting: I’d fight to get it accepted
4 = Worthy: I would like to see it accepted
3 = Borderline: I’m ambivalent about this one
2 = Mediocre: I’d rather not see it in the conference
1 = Poor: I’d fight to have it rejected
Reviewer Confidence (1-5)
5 = Positive that my evaluation is correct. I read the paper very carefully and am very familiar with related work.
4 = Quite sure. I tried to check the important points carefully, and checked for uncited prior work. It’s unlikely, though conceivable, that I missed something that should affect my ratings.
3 = Pretty sure, but there’s a chance I missed something. Although I have a good feel for this area in general, I did not carefully check the paper’s details, e.g., the math, experimental design, or novelty.
2 = Willing to defend evaluation, but it is fairly likely that I missed some details, didn’t understand some central points, or can’t be sure about the novelty of the work.
1 = Not my area, or paper is very hard to understand. My evaluation is just an educated guess.
Detailed Comments
The paper presents SENSEI, a complete, openly licensed subtitle post-editing ecosystem, that combines a two stage ASR-MT cascade with a specifically created editing interface for corrections. The core selling point, being fully open sourced, bilingual, and with a focus on monitoring important features, are clearly explained and makes it genuinely different from already available commercial platforms and existing open-source editors, which are largely single-track and desktop oriented. It’s worth noting, however, that the ASR-MT cascade itself is not new to this paper, and the stated novel contributions are actually narrower than its current three-component structure and consists, essentially, of the SENSEI interface.
The weakest point is the evaluation. Everything that is quantitative concern the ASR-MT cascade and even that is mostly recycled stuff from the IWSLT paper [4]. The interface itself, has no empirical evaluation at all. Claims about reduced cognitive load, reduced error rates, and reduced editing time are all plausible given the design choices, but have not been tested at all. A small pilot study with professional or semi-professional subtitlers, on 1-2 shared videos, evaluating error rates and completion times would’ve been sufficient.
On its positioning as an open-source editor, the paper would benefit from a comparison table between SENSEI and other available tools, at least on some of its most important axis: licence, bilingual use, live compliance flagging, self-hosting, collaborative access, etc.
Overall, the paper present a useful fully openly licensed interface, but overstates on the claims of contribution by presenting the ASR-MT cascade as equally novel as the interface.
Separately, the paper’s claims that the interface reduces cognitive load, error rates, and editing time but without presenting any baseline, measurement, or comparison of any kind, not even against manual/traditional post-editing workflows.
Questions for Authors
- Is there any usability data, even anecdotal, comparing post-editing times or error rate when using SENSEI vs other tools?
- What are the minimum hardware requirements to self-host and use the full SENSEI architecture?
- on the “External Asset Import” of section 3.5.: Does Sensei handle cases where the SRT is a certified monolingual subtitling of the video, and there is the need to only generate other languages subtitles? Or does one need to start from scratch?
Confidential Comments for Committee
You may wish to withhold some comments from the authors, and include them solely for the committee’s internal use. For example, you may want to express a very strong (negative) opinion on the paper, which might offend the authors in some way. Or, perhaps you wish to write something which would expose your identity to the authors. If you wish to share comments of this nature with the committee, this is the place to put them.