AI Engineering: LLMs, RAG and Agents
From zero to production: how large language models work, prompting, embeddings, building complete RAG pipelines, and building AI agents from scratch.
Beginner
- What is AI? (AI vs ML vs Deep Learning vs Generative AI): The vocabulary of the field: AI, machine learning, deep learning, generative AI, models, parameters, training and inference.
- How Large Language Models Work: Next-token prediction, training data, the transformer and attention, pretraining vs instruction tuning vs RLHF, and why models can be wrong.
- Tokens and Context Windows: What tokens are, how tokenizers split text, why everything is counted in tokens, and how the context window limits what a model can see.
- Prompt Engineering Basics: How to write prompts that work: system vs user messages, clear instructions, context, examples (few-shot), XML-tag structure and asking for step-by-step reasoning.
- Calling an LLM API: How applications talk to a model: API keys, the messages format, roles, statelessness, streaming, usage and stop reasons, errors and retries.
- Sampling, Temperature and Model Settings: How the next token is chosen (greedy, temperature, top-k, top-p), and the settings you actually control: max tokens, stop sequences, effort and thinking.
- Hallucinations and Limitations: Why LLMs produce confident falsehoods, the other built-in limits (cutoff, no memory, arithmetic, bias, injection), and the engineering patterns that reduce the damage.
- Structured Output: JSON, Schemas and Validation: Getting machine-readable output from an LLM: JSON, JSON Schema, Pydantic models, the SDK's parse helper, validation and retry.
- Embeddings and Vector Similarity: Turning text into vectors of numbers that capture meaning, and comparing them with cosine similarity and dot product - the foundation of semantic search and RAG.
- What is RAG and Why It Exists: Retrieval-Augmented Generation: fetch relevant passages from your own data and give them to the model, so answers are current, private-data-aware and citable - compared with fine-tuning and long context.
- What is an AI Agent?: An LLM that decides its own next steps in a loop, using tools, until a goal is done - and when a simpler design is better.
Intermediate
- Loading and Parsing Documents: Turning PDFs, HTML, Markdown, tables and scanned pages into clean text plus metadata - the first and most underrated step of every RAG pipeline.
- Chunking Strategies: Splitting documents into retrievable pieces: fixed-size, overlap, recursive, structure-aware and semantic chunking, and how to choose a chunk size.
- Vector Databases and Indexes: Where embeddings live: exact (flat) search vs approximate nearest neighbour indexes (HNSW, IVF), Chroma and pgvector, and filtering by metadata.
- Retrieval: Semantic, Keyword (BM25) and Hybrid Search: How to find the right chunks: dense semantic search, sparse keyword search with BM25, and combining both with reciprocal rank fusion.
- Reranking: Two-stage retrieval: fetch many candidates cheaply, then reorder them with a slower, smarter cross-encoder or LLM so the best chunks reach the prompt.
- Building a Complete RAG Pipeline: End to end: load, clean, chunk, embed and store documents, then retrieve, rerank, build a grounded prompt and get a cited answer from Claude - as a full Python project and a runnable offline JavaScript version.
- Grounded Answers and Citations: Getting the model to answer only from retrieved context, say 'I don't know' when it should, cite its sources, and checking those citations in code.
- Conversational RAG: Making RAG work in a chat: follow-up questions, rewriting them into standalone search queries, managing chat history, and deciding when to retrieve at all.
- RAG in Production: Keeping a RAG system correct, fresh, secure, fast and observable: incremental re-indexing, deletions, permissions, caching, latency budgets and monitoring.
- Tool Use and Function Calling: How an LLM asks your code to run functions: tool definitions, JSON Schema inputs, tool_use and tool_result blocks.
- The Agent Loop: Perceive, think, act, observe - repeated: the loop at the heart of every agent, ReAct, and how and when it must stop.
- Build an Agent from Scratch: Step by step: a real command-line agent with file, calculator and notes tools, a guarded loop, error handling, logging and human approval.
- Agent Memory: How agents remember: the context window as working memory, summarization and compaction, and long-term stores with retrieval.
- Model Context Protocol (MCP): An open standard for plugging tools, data and prompts into any AI application through MCP servers.
- Agentic RAG: Retrieval as a tool: the model decides when to search, what to search for, and when it has enough to answer.
- Workflow Patterns: Five proven ways to compose LLM calls in code: prompt chaining, routing, parallelization, orchestrator-workers and evaluator-optimizer.
Advanced
- Multi-Agent Systems: Several LLM agents with separate contexts and roles working together - when that helps, how to structure it, and what it costs.
- Evaluating RAG: Measure retrieval with recall@k, precision@k, MRR and nDCG, and answers with faithfulness and relevance, using a golden dataset and LLM-as-judge.
- Evaluating Agents: Test agents on task success, trajectories, tool-call accuracy and step/cost budgets, with repeated trials and regression suites.
- Advanced RAG Techniques: Multi-query, HyDE, parent-child (small-to-big), contextual retrieval, metadata filtering, query routing and the GraphRAG idea.
- AI Security: Prompt Injection and Guardrails: Direct and indirect prompt injection, data exfiltration through tools and links, and layered defenses: least privilege, approval, separation and validation.
- Cost, Latency and Prompt Caching: Tokens drive cost and latency; cut both with prompt caching, batching, streaming, smaller prompts, smaller models and lower effort.
- Fine-tuning vs RAG vs Prompting: A decision guide: what prompting, retrieval and fine-tuning each change, what they cost, and which problem each actually solves.
- Observability and Debugging LLM Apps: Trace every model and tool call with inputs, outputs, tokens and latency; log structurally; classify failures and debug from traces.
- Capstone: Build a Documentation Assistant: A guided end-to-end project: ingest docs, hybrid retrieval, grounded answers with citations, an agent that searches and reads files, evals, guardrails and a deployment checklist.