Home / The A to Z of Generative AI / What generative AI is

The A to Z of Generative AI

What generative AI is, in plain words

The book starts where most leaders get stuck: what is this thing, and how is it different from the AI we already had? Here are the basics, in short words.

By Kieran Gilmurray and Olivier Gomez (OG). Page numbers (p.) are the book’s own, 2024 ebook.

Traditional AI does one task. Generative AI makes new content of many kinds from one prompt (p. 12).

Old AI does one job

Before ChatGPT, most business AI was narrow. A model did one task, for example a churn score that tells you which customers may leave. It did that task well, and nothing else.

Generative AI is broad. Give it an instruction in normal language and it makes new text, images, code, music, speech or video. The book calls it a subset of AI, not a replacement for it (p. 12).

“Traditional AI is narrow.”

From the book, p. 12

The book lists what the older tools do well: machine learning to predict, reading text, reading images, and robotic process automation for repetitive work. Generative AI adds writing, summarizing, making images and sound, and simulating many possible outcomes instead of one forecast (p. 21-22). You do not have to choose. The two work best together.

Why everyone noticed at once

The authors put the turning point in November 2022, when OpenAI released ChatGPT 3.5 (p. 14). Two things changed. You talk to it in plain language, so you do not need to be a data scientist. And the big pre-trained models are open through a screen or an API, so a company does not pay to build one from zero (p. 12, p. 42).

The book gives the speed: one million users in less than five days, and 100 million two months after launch (p. 12). A figure in the D chapter calls AI an 80-year-old overnight success: neural networks in the 1940s and 1950s, AI named as a science in 1956, the ELIZA chatbot in 1964, and the ChatGPT models in the last stretch, 2017 to 2024 (p. 43).

Its business example is a law firm with 50,000 documents to review. Before, that meant an army of paralegals. With generative AI the same review is faster and cheaper, and the economics of the firm change (p. 13).

The book also names who feels it first: people in translation, creative writing, marketing, law, technology, health and customer service, where large parts of the work can be automated sooner than expected (p. 43).

What a large language model is

A large language model, or LLM, is a program trained on a huge amount of text. It learns the patterns of language, then writes new text that follows them. It can translate, answer questions and summarize (p. 127).

The book uses a kitchen picture. The model is a chef with a very big library of recipes. More recipes give more varied dishes. But a chef with bad ingredients still serves a bad meal, so the model must be trained and watched with care (p. 127).

A foundation model is the big general model underneath. It is trained once, at scale, then adapted to many tasks. The book names GPT-4 and BERT as examples (p. 71).

Bigger is not the only way. The M chapter explains that you can also train smaller models for one industry or one job, the way a model airplane is a smaller, simpler version of the real one (p. 152).

How a model learns

Supervised, unsupervised and reinforcement learning, as the book explains them (p. 229-230).

The book explains three ways (p. 229-230). Supervised learning: the model studies examples with the right answer, like a student with a teacher. Unsupervised learning: the model gets data with no labels and finds the patterns itself. The book says this is the main method behind models like GPT-3 and GPT-4. Reinforcement learning: the model tries, gets a reward or a penalty, and adjusts.

In real systems they are mixed. Unsupervised learning lets a model read far more data, but with less human control, so bias and harmful content can get in. Supervised steps and human feedback are then used to clean the output (p. 230).

Why it sometimes makes things up

A hallucination is an answer that is false but sounds sure. The book links it mainly to weak or misleading training data (p. 192).

“Crud in, crud out applies more than ever with foundation models.”

From the book, p. 191

The fixes it lists are practical: clean and varied training data, clear and specific prompts, asking the model for its sources, linking it to external data to check answers, human feedback, and a lower temperature setting so the output is less random (p. 192).

Giving the model your facts

Retrieval augmented generation: the librarian finds your facts, the novelist writes the answer (p. 203-204).

A general model does not know your company. Retrieval augmented generation, or RAG, is the book’s answer. It first searches your own documents for the relevant facts, then asks the model to write the answer from them (p. 203-204).

“In simple terms, retrieval-based models are like librarians.”

From the book, p. 204

The generative model is the novelist. It writes well, but it can invent. Put the librarian first and the novelist second, and the answer is easy to read and based on your real data (p. 204).

Next: models that act

The L chapter introduces Large Action Models. An LLM can suggest a hotel room. An action model can suggest it and book it, by working through the same screens a person would use (p. 127-128). The book presents this as the next step, not as something finished.

Next: the A to Z map of the book.

Get the book

The full A to Z: 26 chapters, seven best practice guides and hundreds of example prompts. Kindle, paperback, hardcover and audiobook.

Tool names and prices in the book date from 2024 and change fast. Check them before you act.

Order on AmazonBack to the book page