The Night Before the Birth of the Transformer: Why Eight Researchers at Google Dared to Completely Abandon Recurrent Neural Networks?

CN
1 day ago
Throw away the crutches and keep only the attention.

Author: Robonaissance

Translated by: Deep Tide TechFlow

Deep Tide's Introduction: Before 2017, everyone believed that recurrent neural networks were essential for machine translation and that the attention mechanism was merely a supplementary tool. Eight researchers at Google bet on a radical idea: to throw away the crutches and keep only the attention. The theoretical foundation of this gamble is actually a sentence William James wrote in "The Principles of Psychology" in 1890.

In 1985, Jeffrey Moran and Robert Desimone inserted a microelectrode into a monkey's visual cortex and then waited.

This monkey was trained to do a simple yet strange task. Two stimuli appeared within the receptive field of a single neuron, one at the location the animal was prompted to pay attention to, and the other at the location it was told to ignore. This cell could see both. The question was whether this cell cared.

It cared a great deal. When the monkey attended to one of the two stimuli, the response of neurons in the V4 area was roughly as if the other stimulus did not exist at all. The response to the ignored object, in the author's words, was significantly reduced. Cells in the preceding striate cortex showed no such behavior. Attention is not a metaphor for what the animal is thinking. It is a measurable event in a single cell, and its mechanism is subtraction. The brain is not amplifying the important stuff. It is down-regulating everything else.

Ninety-five years ago, without microelectrodes or monkeys, William James made the same statement in the same order. In Chapter 11 of "The Principles of Psychology," he wrote that everyone knows what attention is, and then still provided a definition: the mind occupies one of several possible objects simultaneously. Then came the most important sentence. James wrote that attention means to withdraw from certain things in order to effectively deal with others.

Withdrawal. Not amplification. The oldest definition in the field and the first single-cell recording correspondingly agree that attention is what you subtract.

Keep this idea in mind. It will return about two thousand words later, wearing different symbols.

Everyone agrees on the cost

By 2016, machine translation had solidified. You take a sentence, run it word by word through a recurrent neural network, accumulating a hidden state that carries everything seen so far, and then generates the translation word by word at the other end using that state. The mainstream variants are LSTM and gated recurrent units. They work effectively. Google had put them into production.

This architecture has a characteristic that is continuously discussed in the field but no longer regarded as a problem. The hidden state at position t is computed as a function of the hidden state at position t minus one. This is not an implementation detail. This is a definition. Position five must wait for position four to exist to be computed, and position four must wait for position three.

The introduction of the Transformer paper bluntly states the consequences — for a section that should have been polite. The authors wrote that this inherent sequential nature precludes parallelization within the training samples, and this issue becomes critical for longer sequence lengths, as memory limits how many samples you can batch process to compensate.

Think about what this means when the hardware in the room is a GPU. A GPU is a machine that can do thousands of things simultaneously. A recurrent network is a machine that does things one at a time. Running the latter on the former is like hiring an orchestra and giving them a piece written for solo violin. Everyone showed up. Most are waiting.

The field knows. There are fixes. The paper cites factorization tricks and conditional computation, both of which improve efficiency, one of which also improves quality. Then it provides the sentence that sets everything else in motion: however, the fundamental constraints of sequential computation remain.

This is what it looks like when a field prices its costs into its worldview. No one thinks recurrence is free. Everyone believes it is necessary.

Attention arrives as a helper

Attention entered machine translation in 2014 as a fix.

The problem it fixed is specific and worth stating precisely because the first form of this fix determined how the field would think about attention for the next three years. In the encoder-decoder model of the time, the encoder read the entire source sentence and compressed it into a single fixed-length vector. The decoder then generated the translation solely from that vector. Every word of a forty-word sentence had to survive in a numerical array that would not change in size regardless of how long the sentence was.

At that time, Dzmitry Bahdanau at Jacobs University Bremen, in collaboration with Kyunghyun Cho and Yoshua Bengio from Montreal, named this bottleneck. Their proposal was to stop imposing forced compression. Let the encoder retain a state for each input position and let the decoder compute a set of weights for all these states at each output step and take a weighted sum. The model learns to see where to look. Their title clearly stated the resulting capability: neural machine translation with joint learning of alignment and translation. Alignment, the statistical translation system that required handcrafted machines to approximate, fell out of the network as a byproduct of learning.

This mechanism worked and spread. A year later, Minh-Thang Luong, Hieu Pham, and Christopher Manning published a set of improvements that made it cheaper and more general. By 2016, attention had become the standard configuration for almost every competitive translation system.

And in every such system, it sat atop the recurrent network.

This is the detail where the entire story turns, marked by a single sentence in the introduction of the Transformer paper. The authors point out that the attention mechanism has become an indispensable part of sequence modeling, allowing for dependence modeling regardless of distance. However, except in a few cases, these mechanisms were used in conjunction with recurrent networks.

Read it again and notice what it does not say. It does not say that attention was underestimated. It says that for three years, the field had a mechanism to directly relate any two positions in one step, regardless of the distance between them, yet tethered it to an architecture that could not do that. Attention is the passenger. Recurrence is the car.

No one asked whether this car would bear weight.

The Questioner

Almost no one, anyway.

Jakob Uszkoreit did not intend to study language. He stumbled into it. His father, Hans Uszkoreit, was a computational linguist who spent fifteen months in an East German prison for protesting the Soviet invasion of Czechoslovakia in his teens, escaped to the West after his release, studied computer science and linguistics in Berlin, and eventually found his way to an artificial intelligence lab in Menlo Park, where his son was born. The family returned to Germany. Jakob went to college there and interned at Google’s Mountain View office, landing in the translation team. He describes this as finally entering the family business. He never completed his PhD.

In 2012, he joined a team at Google building a system that could answer questions directly on the search page. Apple had just released Siri, and Google’s leadership decided this was an emergency. In retrospect, Uszkoreit thinks this panic was unwarranted. But it committed resources to a team devoted to machines that could engage in similar conversations, which is where he hit a wall.

Recurrent networks, even with long short-term memory, could not piece together long paragraphs. The classic demonstration is a sentence where earlier stated facts determine the meaning of phrases stated later, while the model progresses from left to right and must still carry the earlier facts by the time it reaches the later phrases. LSTMs made this longer span compared to ordinary recurrence possible. They did not allow it to work at the scale Uszkoreit wanted. His self-assessment of the toolkit at that time was that the approach he applied was, in his words, "basically a band-aid."

Around 2014, he began developing an alternative he called self-attention: having the model translate a word by directly referencing any other part of the paragraph, weighting it according to how clarifying each part was for the word at hand. He was skeptical for two reasons. The first was that it might simply work better. The second, less obvious but ultimately more significant, was that self-attention viewed many inputs simultaneously rather than sequentially, which meant it was shaped just like the parallel processing chips that the machine learning wave was producing in large quantities. The mechanism and hardware were built for each other. No one had designed it that way.

The reaction was lukewarm. Abandoning recurrence meant discarding an architecture that the whole field had spent years perfecting; Uszkoreit says people raised their eyebrows at the suggestion. Doubters included his father, with whom he did not see entirely eye to eye at the dinner table, by his own account.

He persuaded a few colleagues to experiment anyway. The results were promising enough to publish at a small scale in 2016, using only short text spans. The paper was by Ankur Parikh, Oscar Täckström, Dipanjan Das, and Jakob Uszkoreit, titled "A Decomposable Attention Model," and published at EMNLP 2016. It performed natural language inference using attention without any recurrence at all.

Then his collaborators moved forward. The technology was good enough to deploy, so they went ahead and deployed it across Google Search and eventually Ads. By standard measures, it was a success. Uszkoreit was the only one who thought that a small experiment hinted at a larger one.

This is why when "Attention Is All You Need" appeared a year later, its introduction acknowledged that attention was almost always used in conjunction with recurrent networks, while the citation attached to that phrase—except in a few cases—was that 2016 paper. The exception to the rule, sitting in the bibliography at reference twenty-seven, is the work of those about to break it.

A Café, A Hallway, A Coffee Machine

The Transformer was assembled from overheard conversations. This is not embellishment; it’s what the participants describe.

During lunch at the Google cafeteria in 2016, Uszkoreit heard Illia Polosukhin complaining about the team building direct answers for search pages, which had to return results in milliseconds but were not succeeding. Uszkoreit suggested self-attention. Polosukhin sometimes worked with Ashish Vaswani, who came to Google Brain from the University of Southern California with a PhD in machine translation and was looking for a big problem. Vaswani’s office was right next to Polosukhin’s. He heard about the idea and joined.

The three of them wrote a design document. The name in the title was chosen early on, the theory being that this mechanism transformed the information passing through it, also because Uszkoreit had two Transformers toys as a child. The document ended with a picture of the six of them shooting lasers at each other in the mountains. The opening sentence tells the reader that the authors are terrific.

Niki Parmar joined from Google Search, where she had been building model variants with Uszkoreit. Llion Jones heard about self-attention from colleague Mat Kelcey and joined; Kelcey later heard the project briefing and told Jones he doubted it would work, something he now describes as the most incorrect prediction of his life. Łukasz Kaiser came from a separate effort at Google Brain related to language models, bringing his intern Aidan Gomez, an undergraduate who entered a position reserved for PhD students by talking his way in and wouldn’t discover it for months. Kaiser and Gomez decided to merge their project into the other.

Polosukhin left Google to start his own company in early 2017; this is why his affiliation does not appear in the paper's author line.

Then the group hit a wall. They built a self-attention translation model and measured it against standard benchmarks, finding it roughly on par with the then-best LSTM systems. On par. Not leading. A radical architecture discarding years of prior accumulated in the field and returning parity was not the result. This was an expensive way of staying in place.

The wall collapsed because one person walked through a door.

Noam Shazeer had been at Google since 2000 and had a reputation tracing back to the company’s early ad systems. He had spent five years on deep learning and had developed an interest in language models, which he felt were far from achieving the fluency he thought could be realized. Walking down the hallway of Building 1965, he passed Kaiser’s workspace and overheard Vaswani and Parmar intensely discussing self-attention. He found recurrent networks annoying. The proposal to replace them seemed both right and interesting to him.

Uszkoreit’s account of Shazeer’s contribution is worth detailing, as it describes something that is rarely talked about in the field. Mechanisms that are theoretically sound, he said, often require a few experienced people to implement them very carefully to show any signs of working. Shazeer did not debug the team’s code. He read the idea and wrote his own implementation from scratch, occasionally checking in with Kaiser, disappearing for a while, and then returning with a working version. His colleagues described what he did with words like magic and alchemy. Jones simply called him a wizard.

The specific additions Shazeer made are the subject of the next three articles in the series. The paper’s credit footnotes assign the scaling dot-product attention, multi-head attention, and non-parametric position encoding to him—these are the anchor texts for parts 2, 3, and 4. One person, in a sprint, produced three of the four components for the breakdown of this series.

What the Bet Really Was

The abstract states the position in one sentence. The authors proposed a new, simple network architecture based solely on the attention mechanism, completely discarding recurrence and convolutions.

The word that matters there is completely.

Removing Recurrence

Removing recurrence itself isn’t the radical part. Others had been trying too. Section 2 of the paper lists them: Extended Neural GPU, ByteNet, and ConvS2S, all of which replace recurrence with convolution to compute representations at all positions in parallel. The authors acknowledge that the goal of reducing sequential computation is the basis of those efforts.

But convolutions come at the cost of parallelism, and the paper has a precise description of that cost. In those models, the number of operations required to relate any two positions grows with distance: ConvS2S is linear, ByteNet is logarithmic. Still not free. Connecting two words at the ends of long sentences remains expensive, and the paper cites the standard conclusion: longer paths make long-range dependencies harder to learn.

Self-attention compresses this problem down to constant order. Any position to any other position, in one step, regardless of distance. Table 1 presents a side-by-side display of the three layer types, with the important column being maximum path length: recurrent is n-th order, convolution is log n-th order, self-attention is 1-st order.

This is the bet. It’s not saying attention is useful—this had already been believed in the industry. The bet is that a constant path length is valuable enough to throw away all the other structural priors accumulated in the industry and still come out on top.

The paper does not hide what it is giving up. Within two sentences after the claim of constant operations, there’s an acknowledgment that is rarely retained in later summaries: the constant cost comes at the expense of reduced effective resolution because attention averages over the weighted positions. The existence of multi-head attention is meant to counteract this. The paper tells you on the second page that its core mechanism causes blurriness, and one of its most celebrated components is actually a patch. Part 3 will return to this issue.

There is a second cost, stated in Table 1 and defended in Section 4, which the paper treats as a reasonable trade-off and that the next decade will treat as a core problem of AI infrastructure. The complexity of each layer of self-attention is n squared times d. The paper’s defense is that for the input from sentence lengths in translation, n is typically less than d, which was true in 2017 but is no longer true when someone wants to feed a model a book. Section 7 will collect this bill.

Twelve Hours

The deadline was May 19, for submission to the December conference. The last two weeks were spent in Building 1965, partly because some team members had desks there, partly because the espresso machine there was better than the one in Building 1945. Intern Gomez described a period of continuous debugging, with little sleep for anyone, systematically removing components to see if the model would work without them. Most of what is now called the Transformer is a residue of that process: those parts that could not be removed.

Then the numbers came out.

The base model trained on a single machine with 8 NVIDIA P100 GPUs for 100,000 steps. Clock time: 12 hours. This model surpassed all previously published English-to-German models and ensembles. The large model trained on the same 8 GPU machine for three and a half days reached 28.4 BLEU, exceeding the previous best result (including ensembles) by more than two points. Uszkoreit opened a bottle of champagne he had kept in his mountain expedition truck to celebrate.

Now look at Table 2, the paper’s quietest yet most impactful table. It lists the training costs in floating point operations. At that time, competitive ensemble systems ranged between 1.1×10²¹ and 7.7×10¹⁹. The Transformer large model is at 2.3×10¹⁹, and the base model is an order of magnitude lower than that. The new optimal result was produced by the cheapest system in the table.

A field that had spent years trading scale for quality was shown a model that traded structure for quality. And because this structure maps cleanly to parallel hardware, the same architecture would later become an ideal vehicle for spending compute when there was compute worth spending on. The 2017 results read as an efficiency story. In hindsight, it is a capacity story dressed as an efficiency story.

The paper was submitted with about two minutes to spare. Parmar says the numbers for English-to-French came in about five minutes before submission; she was sitting in the small kitchen waiting for the last number. This perhaps explains a small inconsistency that still exists in the published paper: the abstract and Table 2 report the English-to-French score as 41.8, while Section 6.1 reports it as 41.0. A self-contradictory number resides in one of the most influential papers in modern machine learning because the experiments were completed after the body was written.

The title was finalized a few nights before the deadline. Welshman Jones pointed out that the team rejected the industry convention of choosing a single technique, and the Beatles had already written the line for them. He said the idea came to him in about five seconds, and he didn’t think anyone would use it.

They did.

Back to the Monkey

The Transformer calculates a weighted sum of all other positions for each position, with weights produced by softmax. Softmax does more than just ranking. It suppresses. Increasing one score causes all other scores in the distribution to drop, because the sum is fixed at 1. This core mechanism of the architecture is competitive; its output is defined both by what drives it toward zero and what it elevates.

This is precisely what Moran and Desimone found in V4 neurons when two objects are in their receptive field. This is precisely what James described as withdrawing from certain things to effectively deal with others.

Claiming that the authors of 2017 implemented James would be an exaggeration. They did not. The queries, keys, and values come from information retrieval rather than psychology; Part 2 will trace that lineage correctly. But the words they chose are not coincidental; the field has been honing the concepts behind it for a century. When Anne Treisman and Garry Gelade published the feature integration theory in 1980, they argued that features are registered early in parallel, and attention is the operation that binds them into objects. Parallel registration, then selective binding. This is a description of the Transformer layer, written by researchers of human vision thirty-seven years ago.

This convergence raises the last question of whether it is a true fact about intelligence or a coincidence of vocabulary. Part 8 will attempt to answer.

However, there is a better ending that occurred during the poster presentation at a conference in December 2017. The room was full for four hours. Security eventually had to clear it out. At some point that evening, a man walked up to the poster and told the authors he was impressed with the work; that man was Sepp Hochreiter.

Hochreiter co-invented long short-term memory networks. He came to congratulate the people who had just rendered it obsolete.

The bet: recurrence is not essential for sequence modeling. Architectures built entirely on attention, with no recurrence or convolution layers, can match and exceed optimal levels, and the constant path length between any two positions is more valuable than all the discarded structural priors.

The payoff: won, and faster than the authors expected. Within about two years, this architecture had replaced recurrent models in natural language processing. All cutting-edge language models built since are descendants of it. The paper's stated ambition in the conclusion was to extend this method to other modalities and make generation less sequential. It has birthed an industry.

The status: settled, with an asterisk. This bet was won on a trade-off that the paper explicitly made and priced correctly for 2017: trading the quadratic cost of sequence length for constant path length. Today, all long-context issues in AI are interest payments on that trade-off. Section 7 will collect it.

免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。

Share To
APP

X

Telegram

Facebook

Reddit

CopyLink