What is AI? (AI vs ML vs Deep Learning vs Generative AI)

The vocabulary of the field: AI, machine learning, deep learning, generative AI, models, parameters, training and inference.

What is it?

Artificial intelligence (AI) is the broad goal of making computers do things that normally need human intelligence: understanding language, recognising images, making decisions, writing text. It is an umbrella term, not one specific technique.

For decades most AI was rule-based: a programmer wrote the rules by hand. 'If the email contains the word lottery and comes from an unknown sender, mark it as spam.' That works until the world changes faster than you can write rules. Spammers simply write l0ttery.

Machine learning (ML) flips this around. Instead of writing the rules, you show the computer many examples (emails labelled spam or not spam) and let an algorithm learn the rules from the data. The thing that gets learned is called a model.

A model is just a function with adjustable numbers inside it. Those adjustable numbers are called parameters (or weights). A model that predicts house prices might be price = w1 * size + w2 * bedrooms + b, where w1, w2 and b are parameters. Training means automatically nudging the parameters, example after example, until the model's outputs are close to the right answers. Using a trained model to get answers for new inputs is called inference.

Deep learning is a kind of machine learning that uses neural networks with many layers. A neural network is a big chain of simple maths operations (multiply by weights, add, apply a simple non-linear function), stacked in layers. 'Deep' just means many layers. Deep networks can have millions or billions of parameters, which lets them learn very rich patterns - from raw pixels to 'this is a cat', or from raw text to grammar, facts and style.

Generative AI is deep learning used to create new content - text, images, audio, code - rather than only to label or score things. A spam filter is discriminative (it picks a label). A model that writes an email is generative.

A Large Language Model (LLM) is a generative model for text, trained on enormous amounts of writing. Claude is an LLM. Everything in this subject - prompting, RAG, agents - is about building software around LLMs.

The nesting to remember:

  • AI - any technique that makes machines act intelligently (including hand-written rules).
  • Machine learning - AI that learns patterns from data instead of hand-written rules.
  • Deep learning - machine learning with many-layered neural networks.
  • Generative AI - deep learning models that produce new content.
  • LLMs - generative models specialised in language (and often images and code too).

Two other terms you will meet everywhere: training data (the examples a model learns from) and a dataset split into training data (used to adjust parameters) and test data (held back to check the model works on examples it has never seen). A model that memorises its training data but fails on new data is said to overfit.

As an engineer building AI applications you will almost never train an LLM yourself - that costs enormous compute. You will use a pretrained model through an API (a way for your program to send a request to another program over the network and get a response). Your job is to feed the model the right input, check its output, and connect it to your data and tools. That is what AI engineering means in this subject.

Explain like I'm 10

Rule-based AI is a cookbook: someone wrote down every step, and the cook follows it exactly. Machine learning is a trainee chef who tastes thousands of dishes, gets told 'good' or 'bad' each time, and slowly adjusts their instincts. Deep learning is that chef with an enormous memory for subtle patterns. Generative AI is the chef inventing a brand new dish in the style of everything they have tasted. An LLM is a chef who has 'tasted' a huge library of text and now writes new text in any style you ask for.

Examples

Rules vs learning: a tiny spam filter both ways

// 1) Rule-based: a human wrote the rule
function ruleBased(email) {
  return email.includes("lottery") ? "spam" : "ham";
}

// 2) Learned: count which words appear in labelled examples
const training = [
  ["win a free lottery prize now", "spam"],
  ["free prize claim now", "spam"],
  ["cheap pills free shipping", "spam"],
  ["meeting moved to monday", "ham"],
  ["lunch on monday with the team", "ham"],
  ["project update and meeting notes", "ham"],
];

const counts = { spam: {}, ham: {} };
for (const [text, label] of training) {
  for (const word of text.split(" ")) {
    counts[label][word] = (counts[label][word] || 0) + 1;
  }
}

// The "model": score = spam word hits minus ham word hits
function learned(email) {
  let score = 0;
  for (const word of email.split(" ")) {
    score += (counts.spam[word] || 0) - (counts.ham[word] || 0);
  }
  return score > 0 ? "spam" : "ham";
}

const tests = ["claim your free prize", "team meeting on monday", "l0ttery winner free now"];
for (const t of tests) {
  console.log(JSON.stringify(t), "rule:", ruleBased(t), "| learned:", learned(t));
}

The rule only knows the word 'lottery' and misses every other spam. The learned version picked up 'free', 'prize' and 'now' from data nobody wrote rules for. Real ML uses far better algorithms, but the idea is identical: the behaviour comes from data, not hand-written rules.

What 'training a parameter' really means

// Data secretly follows y = 3x + 1. The model is y = w*x + b.
// w and b are the PARAMETERS. Training = nudging them to reduce error.
const data = [[0, 1], [1, 4], [2, 7], [3, 10], [4, 13]];
let w = 0, b = 0;
const learningRate = 0.02;

for (let step = 0; step <= 2000; step++) {
  let gradW = 0, gradB = 0, loss = 0;
  for (const [x, y] of data) {
    const prediction = w * x + b;
    const error = prediction - y;
    loss += error * error;
    gradW += 2 * error * x;   // how loss changes if w grows
    gradB += 2 * error;       // how loss changes if b grows
  }
  w -= learningRate * gradW / data.length;  // step downhill
  b -= learningRate * gradB / data.length;
  if (step % 500 === 0) {
    console.log("step", step, "w =", w.toFixed(3), "b =", b.toFixed(3), "loss =", (loss / data.length).toFixed(4));
  }
}
console.log("prediction for x = 10:", (w * 10 + b).toFixed(2), "(true value 31)");

This is gradient descent, the algorithm used to train essentially every neural network, including LLMs. An LLM does exactly this, but with billions of parameters instead of two, and with 'predict the next token' as the task instead of a straight line.

Using a pretrained generative model (Python, Anthropic SDK)

# pip install anthropic   and   export ANTHROPIC_API_KEY=...
import anthropic

client = anthropic.Anthropic()  # reads ANTHROPIC_API_KEY from the environment

resp = client.messages.create(
    model="claude-opus-5-5",
    max_tokens=16000,
    messages=[{"role": "user", "content": "In two sentences, what is the difference between AI and machine learning?"}],
)

for block in resp.content:
    if block.type == "text":
        print(block.text)

You did not train anything: someone else trained the model, and you run inference on it through an API. This is the normal starting point for AI engineering.

How it works

Every machine learning system has the same three parts: data (examples), a model (a function with parameters), and a training procedure (an algorithm that adjusts the parameters to reduce a measure of error called the loss).

Training loop, in words: take an example, run the model to get a prediction, measure how wrong it was (loss), work out which direction to move each parameter to make it less wrong (the gradient), move each parameter a small step in that direction. Repeat millions of times. This is gradient descent.

Deep learning uses the same loop, but the model is a neural network with many layers. Each layer transforms its input a little; together the layers can represent complicated patterns. Nobody writes those patterns by hand - they emerge from training on data.

After training, the parameters are frozen and saved. Inference loads them and runs the function on new inputs. When you call an LLM API, the provider runs inference for you on their hardware and sends back the result.

+--------------------------------------------+
| AI: machines doing 'intelligent' tasks     |
|  (rules, search, planning, ML ...)         |
|  +--------------------------------------+  |
|  | Machine learning: learn from data    |  |
|  |  +--------------------------------+  |  |
|  |  | Deep learning: many-layer nets |  |  |
|  |  |  +--------------------------+  |  |  |
|  |  |  | Generative AI            |  |  |  |
|  |  |  |   +------+               |  |  |  |
|  |  |  |   | LLMs |               |  |  |  |
|  |  |  |   +------+               |  |  |  |
|  |  |  +--------------------------+  |  |  |
|  |  +--------------------------------+  |  |
|  +--------------------------------------+  |
+--------------------------------------------+

 data --> [training: adjust parameters] --> model
 new input --> [inference: run model] --> output

Why does it exist?

Many valuable tasks - understanding a customer's question, summarising a contract, spotting a tumour in a scan - have no clean set of rules a human could write down. Machine learning exists because data is easier to collect than rules are to write, and because learned models keep improving as data grows.

Generative AI and LLMs exist because a single model trained on broad text turns out to be useful for thousands of tasks without task-specific training: you just describe the task in words.

When to use it

Reach for ML when the task is fuzzy (language, images, judgement calls), when rules would be endless, or when you have lots of examples of the right answer. Reach for a generative LLM when the input or output is natural language, code or loosely structured content: drafting, summarising, extracting, classifying, answering questions, operating tools.

When not to use it

Do not use AI when a simple deterministic rule or formula does the job: tax calculations, sorting, validating an email format, looking up a record by id. Normal code is cheaper, faster, exact and testable. Do not use an LLM when every answer must be provably correct with no human or programmatic check - LLMs can be confidently wrong (see hallucinations-and-limitations).

Common mistakes

  • Using 'AI', 'ML' and 'LLM' as synonyms - LLMs are one small (if very visible) part of AI.

  • Thinking the model 'looks things up' in its training data at answer time; it only has the parameters it learned, not a copy of the data.

  • Believing you must train your own model to build an AI product; most applications use a pretrained model through an API.

  • Assuming a model that does well on its training examples will do well on new ones (overfitting).

  • Reaching for an LLM for problems that a few lines of ordinary code solve exactly.

  • Treating model output as fact instead of as a prediction that can be wrong.

Practice exercises

  1. Easy:

    For each task, say whether you would use plain code, classic ML, or an LLM, and why: (a) converting Celsius to Fahrenheit, (b) flagging fraudulent card transactions from history, (c) summarising support tickets, (d) checking a password is at least 12 characters.

  2. Easy:

    Explain in your own words the difference between a parameter, training and inference, using the house-price example.

  3. Medium:

    Extend the spam-filter demo: add 4 more training examples and 3 tricky test emails. Find an email the learned filter gets wrong and explain why from the counts.

  4. Medium:

    Change the learning rate in the gradient-descent demo to 0.2 and to 0.001. Describe what happens to the parameters and the loss, and explain why.

  5. Hard:

    Build a version of the gradient-descent demo that learns two inputs: y = 2a - 3b + 5. Generate 20 data points, train w1, w2 and b, and print them every 500 steps.

Interview questions

What is the difference between AI, machine learning and deep learning?

AI is the broad field of making machines act intelligently, including hand-written rules. Machine learning is the subset where behaviour is learned from data. Deep learning is the subset of ML using neural networks with many layers, which is what powers modern image, speech and language models.

What is a model parameter?

A number inside the model that is adjusted during training, such as a weight in a neural network. The set of all parameters is what the model 'knows'. Large language models have billions of them.

What is the difference between training and inference?

Training adjusts the parameters using data and a loss function, which is expensive and done rarely. Inference runs the trained model with fixed parameters on new input to get an output, which happens on every request.

What is generative AI?

Models that produce new content (text, images, audio, code) rather than only classifying or scoring inputs. LLMs are generative models for text.

What is overfitting?

When a model learns its training examples too specifically, including noise, so it scores well on training data but poorly on new data. You detect it by evaluating on held-out test data.

As an application engineer, why don't you usually train your own LLM?

Pretraining an LLM needs vast data and compute. Pretrained models are already broadly capable, so it is far cheaper to use one through an API and adapt it with prompting, retrieval (RAG) and tools; fine-tuning is reserved for specific needs.