
From AI Winters to the Transformer
The history of artificial intelligence is a cycle of euphoria and disillusionment: twice, grand promises were followed by long periods of disappointment that went down in research history as AI winters. 2017 marks the turning point — the moment when architecture and hardware finally came together.
The Beginning
In 1950, Alan Turing poses the question of whether machines can think, introducing the Turing test. At the Dartmouth conference in 1956, John McCarthy coins the term "artificial intelligence" — AI research is born as a field of its own. Expectations are high from the very start.
First Euphoria
Frank Rosenblatt's perceptron (1958) demonstrates that machines can learn from examples, triggering a wave of optimism. Leading researchers promise thinking machines within a generation. The computers of the era, however, cannot keep pace with these announcements.
❄️ First AI Winter
Minsky and Papert's critique of the perceptron (1969) and the devastating Lighthill report (1973) cause funding in the US and the UK to dry up. The promises were decades ahead of the available computing power. Entire branches of research lie fallow.
Expert Systems
Rule-based AI enjoys a commercial boom: expert systems save companies millions, and Japan launches its ambitious "Fifth Generation" project. Knowledge is encoded in thousands of hand-written if-then rules — an approach with a built-in limit.
❄️ Second AI Winter
Expert systems do not scale: every rule has to be written and maintained by hand, and the systems become brittle and expensive. The market for specialized LISP hardware collapses. "AI" becomes a dirty word — many researchers prefer to call their work "machine learning" instead.
The Quiet Years
Away from the headlines, backpropagation matures, and in 1997 Deep Blue beats chess world champion Kasparov. In the background, the real preconditions for the later breakthrough take shape: Moore's law, programmable GPUs, and the web as a nearly inexhaustible source of data.
The Deep Learning Spring
AlexNet wins the ImageNet competition by a wide margin — trained on off-the-shelf GPUs. For the first time, enough data and enough computing power are available at the same time. Within a few years, deep learning becomes the dominant paradigm.
⚡ "Attention Is All You Need"
Google researchers replace recurrence with self-attention: the Transformer. What matters is not just accuracy but parallelizability — the architecture can finally make full use of GPUs and therefore scale. From here, a direct path leads to BERT (2018), the GPT series, and ChatGPT (2022).
Why (Probably) No Winter This Time
For the first time, AI delivers broad everyday value and real products instead of mere promises — from translation and research to code assistance. There is still a share of hype today. Serious assessment means separating value from buzz.
The Common Thread
Both AI winters arose from the same gap: promises and available computing power had drifted apart. The Transformer was the moment an architecture could fully exploit the existing hardware for the first time. That is why 2017 was followed not by a winter, but by a climate change.
To see how we put this technology to practical use in businesses today, take a look at our services in AI & LLM Integration.