In 2014 an intern in Montreal, Dzmitry Bahdanau, lets a translation network glance back at the sentence it is translating, as he did at school. Three years later eight people at Google throw away the network and keep the glance. Their paper lists its authors in random order, and all eight leave. The chapter works through words as coordinates, attention as every word asking every other how much it matters, and the single task these machines are trained on. Then a laboratory in San Francisco does the obvious thing, with help from Nairobi.
This chapter has not been serialised yet. It will appear on Substack first, and this page will carry its opening when it does.
- In this chapter
- Dzmitry Bahdanau, Tomas Mikolov, Jakob Uszkoreit, Noam Shazeer, Llion Jones, Alec Radford, Ilya Sutskever, Dario Amodei, Sam Altman
- The idea it builds
- Words as coordinates; attention as every word asking every other how much it matters; the whole machine trained only to guess the next word.