Stack
A stack contains identically sized layers that differ from classical deep learning models, as shown in Figure I.1. A stack runs from bottom to top. A stack can be an encoder or a decoder.
Figure I.1: Layers form a stack
Transformer stacks learn and see more as they rise in the stacks. Each layer transmits what it learned to the next layer just as our memory does.
Imagine that a stack is the Empire State Building in New York City. At the bottom, you cannot see much. But you will see more and farther as you ascend throught the offices on higher floors and look out the windows. Finally, at the top, you have a fantastic view of Manhattan!