Leitura de longo prazo · 《Attention Is All You Need》(2017)
The 2017 paper "Attention Is All You Need" poses an architectural question: When processing a sentence, can attention primarily associate different positions, replacing the way recurrent networks pass information step-by-step from the input end? The paper proposes the Transformer and compares translation quality and training efficiency on two machine translation tasks. This introduction is based on excerpts actually read from the publicly available paper.
The first concept is aggregating information by relevance. Attention constructs a query for the current position, compares it with keys at other positions, and then aggregates their values according to the resulting weights. Keys and values are vector representations within the model, not manually filled semantic labels; this weighted process allows a position to directly utilize information from other positions, rather than relying solely on the state passed from the previous position.
The second concept is multi-head attention. The model does not compute only one set of associations but computes multiple sets in parallel across different representation subspaces, then combines the results. This allows simultaneous attention to information at different positions and aspects; however, it does not guarantee that each head corresponds to a fixed linguistic rule, nor should the division of labor shown in diagrams be taken as the inevitable division of labor learned by the model.
For example, suppose a sentence contains "After the company acquired the factory, it expanded its production capacity." A reader needs to link "it" with the preceding object and also pay attention to relationships such as "acquired" and "expanded." This scenario is merely intended to help understand why multiple sets of associations are useful; it is not a specific test case reported in the paper, nor does it prove that the model will necessarily understand this sentence correctly.
Parallel training also has clear boundaries. The decoder in the paper still generates outputs token by token, masking positions that have not yet been generated, so predictions can only rely on outputs already known. Therefore, "not using recurrent networks" does not mean "all text can be generated simultaneously at once." When reading descriptions of today's AI models, distinguishing between parallel computation during training and the actual generation order is more useful than simply remembering an architecture name.
Recommended reading: Vaswani et al.'s "Attention Is All You Need" (2017). To understand why the Transformer became an important sequence modeling method, you can start by reading Section 3.2 on attention and multi-head attention, then compare it with the encoder and decoder structures in Section 3.1; the focus is on how information is connected and the boundaries of parallel computation.
Fontes:proceedings.neurips.cc