Transformer Paper: How Attention Replaces Step-by-Step Computation
The 2017 paper "Attention Is All You Need" poses an architectural question: When processing a sentence, can attention primarily associate different positions, replacing the way recurrent networks pass information step-by-step from the input end?
Write-upStandard
The 2017 paper "Attention Is All You Need" poses an architectural question: When processing a sentence, can attention primarily associate different positions, replacing the way recurrent networks pass information step-by-step from the input end? The paper proposes the Transformer and compares translation quality and training efficiency on two machine translation tasks. This introduction is based on excerpts actually read from the publicly available paper.
The first concept is aggregating information by relevance. Attention constructs a query for the current position, compares it with keys at other positions, and then aggregates their values according to the resulting weights. Keys and values are vector representations within the model, not manually filled semantic labels; this weighted process allows a position to directly utilize information from other positions, rather than relying solely on the state passed from the previous position.
The second concept is multi-head attention. The model does not compute only one set of associations but computes multiple sets in parallel across different representation subspaces, then combines the results. This allows simultaneous attention to information at different positions and aspects; however, it does not guarantee that each head corresponds to a fixed linguistic rule, nor should the division of labor shown in diagrams be taken as the inevitable division of labor learned by the model.
For example, suppose a sentence contains "After the company acquired the factory, it expanded its production capacity." A reader needs to link "it" with the preceding object and also pay attention to relationships such as "acquired" and "expanded." This scenario is merely intended to help understand why multiple sets of associations are useful; it is not a specific test case reported in the paper, nor does it prove that the model will necessarily understand this sentence correctly.