AI Engineering: LLMs, RAG and Agents

From zero to production: how large language models work, prompting, embeddings, building complete RAG pipelines, and building AI agents from scratch.

Beginner

  1. What is AI? (AI vs ML vs Deep Learning vs Generative AI): The vocabulary of the field: AI, machine learning, deep learning, generative AI, models, parameters, training and inference.
  2. How Large Language Models Work: Next-token prediction, training data, the transformer and attention, pretraining vs instruction tuning vs RLHF, and why models can be wrong.
  3. Tokens and Context Windows: What tokens are, how tokenizers split text, why everything is counted in tokens, and how the context window limits what a model can see.
  4. Prompt Engineering Basics: How to write prompts that work: system vs user messages, clear instructions, context, examples (few-shot), XML-tag structure and asking for step-by-step reasoning.
  5. Calling an LLM API: How applications talk to a model: API keys, the messages format, roles, statelessness, streaming, usage and stop reasons, errors and retries.
  6. Sampling, Temperature and Model Settings: How the next token is chosen (greedy, temperature, top-k, top-p), and the settings you actually control: max tokens, stop sequences, effort and thinking.
  7. Hallucinations and Limitations: Why LLMs produce confident falsehoods, the other built-in limits (cutoff, no memory, arithmetic, bias, injection), and the engineering patterns that reduce the damage.
  8. Structured Output: JSON, Schemas and Validation: Getting machine-readable output from an LLM: JSON, JSON Schema, Pydantic models, the SDK's parse helper, validation and retry.
  9. Embeddings and Vector Similarity: Turning text into vectors of numbers that capture meaning, and comparing them with cosine similarity and dot product - the foundation of semantic search and RAG.
  10. What is RAG and Why It Exists: Retrieval-Augmented Generation: fetch relevant passages from your own data and give them to the model, so answers are current, private-data-aware and citable - compared with fine-tuning and long context.
  11. What is an AI Agent?: An LLM that decides its own next steps in a loop, using tools, until a goal is done - and when a simpler design is better.

Intermediate

  1. Loading and Parsing Documents: Turning PDFs, HTML, Markdown, tables and scanned pages into clean text plus metadata - the first and most underrated step of every RAG pipeline.
  2. Chunking Strategies: Splitting documents into retrievable pieces: fixed-size, overlap, recursive, structure-aware and semantic chunking, and how to choose a chunk size.
  3. Vector Databases and Indexes: Where embeddings live: exact (flat) search vs approximate nearest neighbour indexes (HNSW, IVF), Chroma and pgvector, and filtering by metadata.
  4. Retrieval: Semantic, Keyword (BM25) and Hybrid Search: How to find the right chunks: dense semantic search, sparse keyword search with BM25, and combining both with reciprocal rank fusion.
  5. Reranking: Two-stage retrieval: fetch many candidates cheaply, then reorder them with a slower, smarter cross-encoder or LLM so the best chunks reach the prompt.
  6. Building a Complete RAG Pipeline: End to end: load, clean, chunk, embed and store documents, then retrieve, rerank, build a grounded prompt and get a cited answer from Claude - as a full Python project and a runnable offline JavaScript version.
  7. Grounded Answers and Citations: Getting the model to answer only from retrieved context, say 'I don't know' when it should, cite its sources, and checking those citations in code.
  8. Conversational RAG: Making RAG work in a chat: follow-up questions, rewriting them into standalone search queries, managing chat history, and deciding when to retrieve at all.
  9. RAG in Production: Keeping a RAG system correct, fresh, secure, fast and observable: incremental re-indexing, deletions, permissions, caching, latency budgets and monitoring.
  10. Tool Use and Function Calling: How an LLM asks your code to run functions: tool definitions, JSON Schema inputs, tool_use and tool_result blocks.
  11. The Agent Loop: Perceive, think, act, observe - repeated: the loop at the heart of every agent, ReAct, and how and when it must stop.
  12. Build an Agent from Scratch: Step by step: a real command-line agent with file, calculator and notes tools, a guarded loop, error handling, logging and human approval.
  13. Agent Memory: How agents remember: the context window as working memory, summarization and compaction, and long-term stores with retrieval.
  14. Model Context Protocol (MCP): An open standard for plugging tools, data and prompts into any AI application through MCP servers.
  15. Agentic RAG: Retrieval as a tool: the model decides when to search, what to search for, and when it has enough to answer.
  16. Workflow Patterns: Five proven ways to compose LLM calls in code: prompt chaining, routing, parallelization, orchestrator-workers and evaluator-optimizer.

Advanced

  1. Multi-Agent Systems: Several LLM agents with separate contexts and roles working together - when that helps, how to structure it, and what it costs.
  2. Evaluating RAG: Measure retrieval with recall@k, precision@k, MRR and nDCG, and answers with faithfulness and relevance, using a golden dataset and LLM-as-judge.
  3. Evaluating Agents: Test agents on task success, trajectories, tool-call accuracy and step/cost budgets, with repeated trials and regression suites.
  4. Advanced RAG Techniques: Multi-query, HyDE, parent-child (small-to-big), contextual retrieval, metadata filtering, query routing and the GraphRAG idea.
  5. AI Security: Prompt Injection and Guardrails: Direct and indirect prompt injection, data exfiltration through tools and links, and layered defenses: least privilege, approval, separation and validation.
  6. Cost, Latency and Prompt Caching: Tokens drive cost and latency; cut both with prompt caching, batching, streaming, smaller prompts, smaller models and lower effort.
  7. Fine-tuning vs RAG vs Prompting: A decision guide: what prompting, retrieval and fine-tuning each change, what they cost, and which problem each actually solves.
  8. Observability and Debugging LLM Apps: Trace every model and tool call with inputs, outputs, tokens and latency; log structurally; classify failures and debug from traces.
  9. Capstone: Build a Documentation Assistant: A guided end-to-end project: ingest docs, hybrid retrieval, grounded answers with citations, an agent that searches and reads files, evals, guardrails and a deployment checklist.