Avançar para o conteúdo principal

Generative AI exists because of the transformer (February 02, 2025)

 

February 02, 2025


https://ig.ft.com/generative-ai/



Over the past few years, we have taken a gigantic leap forward in our decades-long quest to build intelligent machines: the advent of the large language model, or LLM.

This technology, based on research that tries to model the human brain, has led to a new field known as generative AI — software that can create plausible and sophisticated text, images and computer code at a level that mimics human ability.

Businesses around the world have begun to experiment with the new technology in the belief it could transform media, finance, law and professional services, as well as public services such as education. The LLM is underpinned by a scientific development known as the transformer model, made by Google researchers in 2017.

“While we’ve always understood the breakthrough nature of our transformer work, several years later, we’re energised by its enduring potential across new fields, from healthcare to robotics and security, enhancing human creativity, and more,” says Slav Petrov, a senior researcher at Google, who works on building AI models, including LLMs.

LLMs’ touted benefits — the ability to increase productivity by writing and analysing text — are also why it poses a threat to humans. According to Goldman Sachs, it could expose the equivalent of 300mn full-time workers across big economies to automation, leading to widespread unemployment.

As the technology is rapidly woven into our lives, understanding how LLMs generate text means understanding why these models are such versatile cognitive engines — and what else they can help create.

To write text, LLMs must first translate words into a language they understand.

First a block of words is broken into tokens — basic units that can be encoded. Tokens often represent fractions of words, but we’ll turn each full word into a token.

In order to grasp a word’s meaning, work in our example, LLMs first observe it in context using enormous sets of training data, taking note of nearby words. These datasets are based on collating text published on the internet, with new LLMs trained using billions of words.

Eventually, we end up with a huge set of the words found alongside work in the training data, as well as those that weren’t found near it.

As the model processes this set of words, it produces a vector — or list of values — and adjusts it based on each word’s proximity to work in the training data. This vector is known as a word embedding.

A word embedding can have hundreds of values, each representing a different aspect of a word’s meaning. Just as you might describe a house by its characteristics — type, location, bedrooms, bathrooms, storeys — the values in an embedding quantify a word’s linguistic features.

The way these characteristics are derived means we don’t know exactly what each value represents, but words we expect to be used in comparable ways often have similar-looking embeddings.

A pair of words like sea and ocean, for example, may not be used in identical contexts (‘all at ocean’ isn't a direct substitute for ‘all at sea’), but their meanings are close to each other, and embeddings allow us to quantify that closeness.

By reducing the hundreds of values each embedding represents to just two, we can see the distances between these words more clearly.

We might spot clusters of pronouns, or modes of transportation, and being able to quantify words in this way is the first step in a model generating text.

I
we
go
to
work
by
train
the
they
swim
run
walk
with
on
in
car
bus
college
school

But this alone is not what makes LLMs so clever. What unlocked their abilities to parse and write as fluently as they do today is a tool called the transformer, which radically sped up and augmented how computers understood language.

Transformers process an entire sequence at once — be that a sentence, paragraph or an entire article — analysing all its parts and not just individual words.

This allows the software to capture context and patterns better, and to translate — or generate — text more accurately. This simultaneous processing also makes LLMs much faster to train, in turn improving their efficiency and ability to scale.

Research outlining the transformer model was first published by a group of eight AI researchers at Google in June 2017. Their 11-page research paper marked the start of the generative AI era.

A key concept of the transformer architecture is self-attention. This is what allows LLMs to understand relationships between words.

Self-attention looks at each token in a body of text and decides which others are most important to understanding its meaning.

Before transformers, the state of the art AI translation methods were recurrent neural networks (RNNs), which scanned each word in a sentence and processed it sequentially.

With self-attention, the transformer computes all the words in a sentence at the same time. Capturing this context gives LLMs far more sophisticated capabilities to parse language.

In this example, assessing the whole sentence at once means the transformer is able to understand that interest is being used as a noun to explain an individual’s take on politics.

If we tweak the sentence . . .

. . . the model understands interest is now being used in a financial sense.

And when we combine the sentences, the model is still able to recognise the correct meaning of each word thanks to the attention it gives the accompanying text.

For the first use of interest, it is no and in that are most attended.

For the second, it is rate and bank.

This functionality is crucial for advanced text generation. Without it, words that can be interchangeable in some contexts but not others can be used incorrectly.

Effectively, self-attention means that if a summary of this sentence was produced, you wouldn’t have enthusiasm used when you were writing about interest rates.

This capability goes beyond words, like interest, that have multiple meanings.

In the following sentence, self-attention is able to calculate that it is most likely to be referring to dog.

And if we alter the sentence, swapping hungry for delicious, the model is able to recalculate, with it now most likely to refer to bone.

The benefits of self-attention for language processing increase the more you scale things up. It allows LLMs to take context from beyond sentence boundaries, giving the model a greater understanding of how and when a word is used.

The dog chewed the bone because it was delicious.

In a quaint little town nestled amidst rolling hills and lush meadows, there lived a faithful canine with a had a red collar, that adorned his neck like a crown. This charming creature, whose name was Rex, was a beloved member of the Johnson family. With a gleaming coat of golden fur and eyes that sparkled with warmth and affection, he won the hearts of everyone who crossed his path. The red collar became a symbol of his loyalty and an emblem of the countless adventures that awaited him.From the day Rex trotted into the Johnsons' lives, he brought an abundance of joy and laughter. His days were filled with frolics in the nearby park, chasing butterflies, and playing fetch with the children. In the afternoons, he would faithfully accompany Mr. Johnson on his walks around the neighborhood, sniffing the scents of the world with curious enthusiasm. The townsfolk admired his friendly nature and the undeniable bond he shared with his family. Wherever he went, the red collar shone like a beacon, a reminder of the unconditional love and loyalty he offered to those who embraced him.ate dinner at 6 pm like clockwork every day. Mrs. Johnson ensured that Rex's meals were always prepared with love and care, with a wholesome blend of nutritious ingredients. Rex would sit patiently, his tail wagging eagerly, until the clock struck six. As the aroma of his favorite meal wafted through the air, he couldn't contain his excitement. His red collar jingled with each step he took towards his feeding bowl, a sound that had become synonymous with the joyous anticipation of dinner time.Beyond the boundaries of the town, Rex's escapades expanded into a realm of wild exploration. He roamed through vast meadows and ventured into dense forests, his red collar contrasting against the vibrant hues of nature. On one such adventure, he met a pack of fellow canines, and together, they formed an inseparable bond. They navigated through the wilderness, encountering thrilling encounters with other animals, all while sharing tales of loyalty and bravery under the watchful stars.As time passed, Rex grew older, and the years began to leave their mark on his once vibrant fur. Though his steps may have slowed, his spirit remained unwavering. The red collar, now slightly faded, continued to adorn his neck, a symbol of the unforgettable memories he had woven into the fabric of his family's life. As he approached the twilight of his days, the Johnsons made sure to reciprocate the love and care he had bestowed upon them throughout the years. They cherished every moment, knowing that the time they spent together was a precious gift, and they were determined to make it memorable.

In a quaint little village nestled amidst picturesque landscapes, there lived a delightful canine named Luna, whose presence brought an unyielding sense of joy to all. Luna was a beautiful mix of Labrador and Border Collie, with a silky black coat that shimmered under the sun and a pair of striking amber eyes that sparkled with intelligence. The town's residents couldn't help but smile whenever they caught a glimpse of Luna's wagging tail and the exuberance in her every step. Her playful energy was infectious, drawing people from all walks of life to her side, eager to bask in the warmth of her affectionate nature. From dawn to dusk, the village was adorned with the was his owner's best friend. Every morning, Luna would race through the cobbled streets, her paws dancing in rhythm with her excitement for the day ahead. The local children adored her, and she became their loyal companion in every adventure they embarked upon. She would join them in the meadows, chasing butterflies, and rolling in the soft grass. Her presence at the village square became a delightful spectacle, as she delighted in meeting new faces and playfully nuzzling anyone willing to indulge her in games of fetch or tug-of-war. Luna's boundless energy was a testament to the sheer happiness that can be found in the simplest of moments.One of Luna's favorite pastimes was when she loved playing fetchwith the children. Every afternoon, she would eagerly await the school bell, knowing that her little friends would soon come running towards her. With a bright red ball gripped firmly in their hands, they would take turns throwing it far into the distance, and Luna would dash like the wind to retrieve it, her tail swishing back and forth in sheer delight. The children's laughter filled the air as they cheered her on, and the bond between Luna and her young playmates grew stronger with each game. Through their innocent games of fetch, they learned the value of camaraderie and the joy of giving and receiving unconditional love.

One of the world’s largest and most advanced LLMs is GPT-4, OpenAI’s latest artificial intelligence model which the company says exhibits “human-level performance” on several academic and professional benchmarks such as the US bar exam, advanced placement tests and the SAT school exams.

GPT-4 can generate and ingest large volumes of text: users can feed in up to 25,000 English words, which means it could handle detailed financial documentation, literary works or technical manuals.

The product has reshaped the tech industry, with the world’s biggest technology companies — including Google, Meta and Microsoft, who have backed OpenAI — racing to dominate the space, alongside smaller start-ups.

The LLMs they have released include Google’s PaLM model, which powers its chatbot Bard, Anthropic’s Claude model, Meta’s LLaMA and Cohere’s Command, among others.

While these models are already being adopted by an array of businesses, some of the companies behind them are facing legal battles around their use of copyrighted text, images and audio scraped from the web.

The reason for this is that current LLMs are trained on most of the English-language internet — a volume of information that makes them far more powerful than previous generations.

From this enormous corpus of words and images, the models learn how to recognise patterns and eventually predict the next best word.

But things don’t always go to plan. While the text may seem plausible and coherent, it isn’t always factually correct. LLMs are not search engines looking up facts; they are pattern-spotting engines that guess the next best option in a sequence.

Because of this inherent predictive nature, LLMs can also fabricate information in a process that researchers call “hallucination”. They can generate made-up numbers, names, dates, quotes — even web links or entire articles.

Users of LLMs have shared examples of links to non-existent news articles on the FT and Bloomberg, made-up references to research papers, the wrong authors for published books and biographies riddled with factual mistakes.

In one high-profile incident in New York, a lawyer used ChatGPT to create a brief for a case. When the defence interrogated the report, they discovered it was littered with made-up judicial opinions and legal citations. “I did not comprehend that ChatGPT could fabricate cases,” the lawyer later told a judge during his own court hearing.

Although researchers say hallucinations will never be completely erased, Google, OpenAI and others are working on limiting them through a process known as “grounding”. This involves cross-checking an LLM’s outputs against web search results and providing citations to users so they can verify.

Humans are also used to provide feedback and fill gaps in information — a process known as reinforcement learning by human feedback (RLHF) — which further improves the quality of the output. But it is still a big research challenge to understand which queries might trigger these hallucinations, as well as how they can be predicted and reduced.

Despite these limitations, the transformer has resulted in a host of cutting-edge AI applications. Apart from powering chatbots such as Bard and ChatGPT, it drives autocomplete on our mobile keyboards and speech recognition in our smart speakers.

Its real power, however, lies beyond language. Its inventors discovered that transformer models could recognise and predict any repeating motifs or patterns. From pixels in an image, using tools such as Dall-E, Midjourney and Stable Diffusion, to computer code using generators like GitHub CoPilot. It could even predict notes in music and DNA in proteins to help design drug molecules.

For decades, researchers built specialised models to summarise, translate, search and retrieve. The transformer unified all those actions into a single structure capable of performing a huge variety of tasks.

“Take this simple model that predicts the next word and it . . . can do anything,” says Aidan Gomez, chief executive of AI start-up Cohere, and a co-author of the transformer paper.

Now they have one type of model that is “trained on the entire internet and what falls out the other side does all of that and better than anything that came before”, he says.

“That is the magical part of the story.”

This story is free to read so you can share it with family and friends who don’t yet have an FT subscription.

Madhumita Murgia is the FT’s artificial intelligence editor.

Visual storytelling team: Dan Clark, Sam Learner, Irene de la Torre Arenas, Sam Joiner, Eade Hemingway and Oliver Hawkins.

With thanks to Slav Petrov, Jakob Uszkoreit, Aidan Gomez and Ashish Vaswani.

To generate the 50D word embeddings we used the GloVe 6B 50D pre-trained model and converted to Word2Vec format. To generate the 2D representation of word embeddings we used the BERT large language model and reduced dimensionality using UMAP. The self-attention values and the probability scores in the beam search section are conceptual.

We used the free version of ChatGPT-3.5 to generate some of the example sentences used in the visual part of the word embedding and self attention section.



Comentários

Mensagens populares deste blogue

Le Grand Raid des Pyrénées

Portugueses com 50 ou mais Maratonas e Ultras

Where The Schooling Went

The Completion

How Trust Becomes Access

How an AI Agent Works