Attn Mask for Non-causal Models

Open roshansh-cmu opened this issue 3 years ago • 2 comments

We are examining non-NLP applications of the cosformer self-attention, and would need to use attention masking for the padded tokens in the batch. Is there a way to incorporate this ? Because the code does not explicitly compute the attention weights on which masking is traditionally applied.

Mar 09 '22 23:03 roshansh-cmu

We are examining non-NLP applications of the cosformer self-attention, and would need to use attention masking for the padded tokens in the batch. Is there a way to incorporate this ? Because the code does not explicitly compute the attention weights on which masking is traditionally applied.

Can you provide some examples before/after masking?

Mar 12 '22 09:03 Doraemonzzz

for example, Swin-transformer mask Uploading image.png…

Sep 07 '23 06:09 npzl