Viewing a single comment thread. View all comments

Lopsided-Factor-780 t1_j7743wv wrote

Question from a noob:
When they say H_Fuse is fed into the decoder model, such that Y = Decoder(H_Fuse), how is it fed in? Is it fed in like the encoder output in an encoder-decoder transformer with cross-attention? Or something else?

Also, if there is a separate encoder and decoder component, are they trained together or separately?

3