Reset attention mask across doc boundary

#14

by jimmyhbx - opened Apr 18

Apr 18

Hi,

Thanks for sharing this great model. I am wondering if we want to continue pretrain Llama 3 and also reset the attention mask to avoid cross doc attention, should we also reset the position ID in the rotary embedding?

Johnhy42112

Apr 18

This comment has been hidden

suzhu001

Apr 19

I think there is no need to reset the position ID in the pre-training stage, since docs longer than 8k are limited.

The implementation of cross doc attention-mask is available in Flash-attention, https://github.com/Dao-AILab/flash-attention/issues/432#issuecomment-1698610752.

suzhu001

Apr 30

I just realized that resetting the position ID after the doc boundary is identical to not resetting. Because the rotary embedding refers to relative position.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.

Tap or paste here to upload images

· Sign up or log in to comment