Loading and Parsing Documents

Turning PDFs, HTML, Markdown, tables and scanned pages into clean text plus metadata - the first and most underrated step of every RAG pipeline.

What is it?

A RAG system can only retrieve what it has ingested, and it can only ingest what it can read. Document loading (also called parsing or extraction) is the step that turns your raw files - PDFs, web pages, Word documents, Markdown, wiki exports, spreadsheets - into plain text plus metadata that the rest of the pipeline can chunk, embed and search.

It sounds boring, and that is exactly why it is the most common hidden cause of bad RAG answers. If a PDF's two columns get interleaved line by line, if a table becomes a soup of numbers, or if every page starts with the same header and footer, the chunks you embed are garbage, and no clever retrieval later can fix that. The rule of thumb: garbage in, garbage retrieved.

What each format gives you:

  • Plain text and Markdown - easiest. The text is already there; Markdown also gives you structure (headings #, lists, code blocks) that you can keep for smarter chunking later.
  • HTML - the text is mixed with tags, menus, cookie banners, scripts and footers. You must extract the main content and drop the boilerplate (repeated navigation and chrome that appears on every page).
  • PDF - a PDF is a description of where to draw characters on a page, not a document with paragraphs. A text extractor (such as the Python library pypdf) reassembles the characters into lines, but reading order, columns, hyphenation and tables can come out scrambled.
  • Scanned PDFs and images - contain no text at all, only pictures of text. You need OCR (optical character recognition: software that recognises letters in an image) to get text out. If a PDF extractor returns an empty string for a page, it is probably scanned.
  • Office files (DOCX, PPTX, XLSX) - structured XML underneath; dedicated libraries read them reliably.

Metadata is data about each piece of text: which file it came from, the page number, the section heading, the author, the date, the URL, the department that owns it, who is allowed to read it. You attach it at load time and carry it all the way through the pipeline. Metadata is what lets you later cite sources ('handbook.pdf, page 12'), filter searches ('only 2025 policies'), enforce permissions ('only documents the HR team may see') and re-index a single file when it changes.

Cleaning is the step right after extraction. Typical cleaning operations:

  • Normalise whitespace (collapse runs of spaces, at most one blank line between paragraphs).
  • Re-join words hyphenated across line breaks (retriev-\nal becomes retrieval).
  • Remove repeated headers, footers and page numbers that appear on every page.
  • Remove boilerplate such as navigation menus, 'Accept cookies' text and share buttons from HTML.
  • Fix encoding problems (odd characters such as ’ where an apostrophe should be) and remove invisible characters such as soft hyphens.
  • Drop near-empty pages and exact duplicate documents.

Tables need special care. A table flattened into one line ('Plan Price Seats Basic 10 1 Pro 30 5') loses the link between a header and its value. A better approach is to convert each row into a self-describing sentence or Markdown row: 'Plan: Pro; Price: 30; Seats: 5'. Then any chunk containing that row still makes sense on its own.

The output of this step is a list of documents (or pages / sections), each with text and metadata. That list is the input to chunking.

Explain like I'm 10

Document loading is like a librarian receiving a delivery of books, photocopies, printed web pages and faxes. Before anything can be shelved, they must be unpacked, the coffee-stained cover pages thrown away, the faxes typed up (OCR), and a catalogue card written for each item: title, author, date, which shelf, who may borrow it. A library with no catalogue cards (metadata) and smudged pages (bad extraction) is useless, however clever the librarian is at finding things later.

Examples

Clean extracted text: whitespace, hyphenation, repeated headers

// Pretend this came out of a PDF text extractor, two pages
const pages = [
  "ACME Corp - Employee Handbook      Page 1\n\nAll employees receive 25 days of paid\nannual leave. Unused leave can be car-\nried over, up to 5 days.\n\n\n\n",
  "ACME Corp - Employee Handbook      Page 2\n\nRemote work   is allowed up to 3 days\nper week with manager approval.\u00ad",
];

function findRepeatedFirstLines(pages) {
  // A first line that appears on most pages (ignoring digits) is a header
  const counts = {};
  for (const p of pages) {
    const first = p.split("\n")[0].replace(/[0-9]+/g, "#").trim();
    counts[first] = (counts[first] || 0) + 1;
  }
  return Object.keys(counts).filter((line) => counts[line] >= Math.ceil(pages.length / 2));
}

function clean(text, headers) {
  const lines = text.split("\n").filter((line) => {
    const normalized = line.replace(/[0-9]+/g, "#").trim();
    return !headers.includes(normalized);
  });
  return lines.join("\n")
    .replace(/\u00ad/g, "")              // invisible soft hyphens
    .replace(/(\w)-\n(\w)/g, "$1$2")     // re-join car-\nried -> carried
    .replace(/[ \t]+/g, " ")              // collapse spaces
    .replace(/([^\n])\n([^\n])/g, "$1 $2") // single newlines inside a paragraph -> space
    .replace(/\n{3,}/g, "\n\n")           // at most one blank line
    .trim();
}

const headers = findRepeatedFirstLines(pages);
console.log("detected headers:", headers);
pages.forEach((p, i) => {
  console.log("--- page " + (i + 1) + " ---");
  console.log(clean(p, headers));
});

Real extractors produce exactly these problems. Detecting repeated lines across pages is a simple, effective way to drop headers and footers. Note how the hyphenated word 'car-ried' is restored, so a search for 'carried' can now match.

HTML to text, Markdown sections and tables to sentences (with metadata)

// 1) HTML: drop scripts, nav and footer, keep main content
const html = "<html><head><script>track()</script></head><body>" +
  "<nav>Home | Pricing | Login</nav>" +
  "<main><h1>Refund policy</h1><p>Refunds are issued within <b>14 days</b>.</p>" +
  "<p>Contact support&amp;billing for help.</p></main>" +
  "<footer>(c) ACME</footer></body></html>";

function htmlToText(h) {
  return h
    .replace(/<(script|style|nav|footer|header)[^>]*>[\s\S]*?<\/\1>/gi, "")
    .replace(/<\/(p|h[1-6]|li|div)>/gi, "\n")
    .replace(/<[^>]+>/g, "")
    .replace(/&amp;/g, "&").replace(/&lt;/g, "<").replace(/&gt;/g, ">")
    .split("\n").map((l) => l.trim()).filter(Boolean).join("\n");
}
console.log(htmlToText(html));

// 2) Markdown: split by headings and keep the heading path as metadata
const markdown = "# Handbook\nIntro text.\n## Leave\n25 days per year.\n## Remote work\nUp to 3 days a week.";
const sections = [];
let path = [];
for (const line of markdown.split("\n")) {
  const m = line.match(/^(#+) (.*)$/);
  if (m) {
    path = path.slice(0, m[1].length - 1).concat(m[2]);
    sections.push({ text: "", metadata: { source: "handbook.md", section: path.join(" > ") } });
  } else if (sections.length) {
    sections[sections.length - 1].text += line;
  }
}
console.log(JSON.stringify(sections, null, 1));

// 3) Table rows -> self-describing sentences
const table = [["Plan", "Price", "Seats"], ["Basic", "10", "1"], ["Pro", "30", "5"]];
const [header, ...rows] = table;
for (const row of rows) {
  console.log(header.map((h, i) => h + ": " + row[i]).join("; "));
}

Regex HTML stripping is fine for a demo; in production use a real HTML parser (BeautifulSoup in Python) or a main-content extractor. The important ideas are: remove boilerplate, keep structure (headings) as metadata, and turn tables into rows that make sense on their own.

Real loaders in Python: PDF pages with pypdf, HTML with BeautifulSoup

# pip install pypdf beautifulsoup4
from pathlib import Path
from pypdf import PdfReader
from bs4 import BeautifulSoup

def load_pdf(path: Path):
    reader = PdfReader(str(path))
    for page_number, page in enumerate(reader.pages, start=1):
        text = page.extract_text() or ""
        if not text.strip():
            print(f"warning: {path.name} page {page_number} has no text layer (scanned? needs OCR)")
            continue
        yield {"text": text, "metadata": {"source": path.name, "page": page_number, "type": "pdf"}}

def load_html(path: Path):
    soup = BeautifulSoup(path.read_text(encoding="utf-8"), "html.parser")
    for tag in soup(["script", "style", "nav", "footer", "header", "aside"]):
        tag.decompose()                      # drop boilerplate
    main = soup.find("main") or soup.body or soup
    title = soup.title.get_text(strip=True) if soup.title else path.stem
    yield {"text": main.get_text("\n", strip=True),
           "metadata": {"source": path.name, "title": title, "type": "html"}}

def load_text(path: Path):
    yield {"text": path.read_text(encoding="utf-8", errors="replace"),
           "metadata": {"source": path.name, "type": path.suffix.lstrip(".")}}

LOADERS = {".pdf": load_pdf, ".html": load_html, ".htm": load_html, ".md": load_text, ".txt": load_text}

def load_folder(folder: str):
    for path in sorted(Path(folder).rglob("*")):
        loader = LOADERS.get(path.suffix.lower())
        if loader:
            yield from loader(path)

if __name__ == "__main__":
    docs = list(load_folder("docs"))
    print(len(docs), "documents/pages loaded")
    print(docs[0]["metadata"], docs[0]["text"][:200])

One loader per file type, each producing the same shape: {text, metadata}. Loading PDFs page by page means you keep the page number for citations. Pages with no text are reported instead of silently ignored, because they usually need OCR (for example with the Tesseract OCR engine) or a document-parsing service.

How it works

A loading pipeline has four stages: discover (list files from a folder, bucket, wiki or database), extract (format-specific parser turns bytes into text), clean (normalise and remove noise) and annotate (attach metadata).

PDF extraction works by reading the drawing instructions in the file ('put glyph A at x=72, y=700') and grouping characters into words and lines by their positions. Simple extractors follow the order in the file, which is usually but not always the reading order. Multi-column layouts, sidebars and footnotes are where it goes wrong. Layout-aware parsers (and some vision-model-based parsers) look at positions on the page to recover columns and tables; they are slower but much more accurate on complex documents.

OCR runs a recognition model over an image of the page and outputs text, often with a confidence score per word. Low-confidence pages are worth flagging for human review. Some modern LLMs can also read page images directly, which is useful for diagrams and messy scans, but it costs more per page than plain extraction.

Deduplication is usually done with a hash: a short fingerprint computed from the text (for example SHA-256). Two documents with the same hash are identical, so you keep one. Storing the hash in metadata also lets you skip unchanged files when you re-ingest (see rag-in-production).

Whatever the source, normalise everything into one shape - { text, metadata } - so that chunking, embedding and storage code never needs to know whether something came from a PDF or a web page.

  files / URLs / wiki / DB
            |
            v
 +---------------------+
 | discover            |  list sources
 +---------------------+
            v
 +---------------------+   pdf -> pypdf (+OCR)
 | extract (per type)  |   html -> parser
 +---------------------+   md/txt -> read
            v
 +---------------------+   whitespace, hyphens,
 | clean               |   headers, boilerplate
 +---------------------+
            v
 +---------------------+   source, page, section,
 | annotate metadata   |   date, owner, access
 +---------------------+
            v
   [{ text, metadata }, ...]  --> chunking

Why does it exist?

Your knowledge does not live in neat strings: it lives in PDFs, wikis, tickets and web pages built for human eyes. An LLM and an embedding model need clean text. Loading exists to bridge that gap, and metadata exists because answers are only trustworthy when you can say where they came from and who is allowed to see them.

When to use it

Every RAG pipeline needs this step. Invest in it heavily when your corpus has PDFs with columns or tables, scanned documents, or HTML pages with lots of navigation. Always look at a sample of extracted text by eye before building anything else - it is the cheapest debugging you will ever do.

When not to use it

If your data is already structured (rows in a database, JSON from an API), do not flatten it into a document and parse it back: query it directly with SQL or a tool call (see tool-use). If the whole corpus is small and stable, you may be able to paste it into the prompt and skip retrieval entirely (see what-is-rag).

Common mistakes

  • Never looking at the extracted text, then blaming the embedding model for bad answers.

  • Dropping metadata (file name, page, section, date) at load time, which makes citations and filtering impossible later.

  • Silently skipping scanned pages that returned empty text instead of logging them for OCR.

  • Leaving repeated headers, footers and navigation in every chunk, so every chunk looks similar to every query.

  • Flattening tables into one long line so row values lose their column headers.

  • Indexing the same document several times (copies, versions), which fills the top results with duplicates.

  • Treating loader output as trusted: documents can contain hidden text and instructions aimed at the model (see ai-security).

Practice exercises

  1. Easy:

    Run the cleaning demo and add a footer line 'Confidential - do not distribute' to both pages. Extend findRepeatedFirstLines to also detect repeated LAST lines and remove them.

  2. Easy:

    Pick three real files you own (a PDF, a web page saved as HTML, a Markdown file). Extract their text with the Python loaders and read the first 1,000 characters of each. List every quality problem you see.

  3. Medium:

    Extend the Markdown section parser to keep code blocks (lines between ``` fences) intact, even if they contain lines that start with #.

  4. Medium:

    Write a function that converts a CSV string (first row headers) into one self-describing sentence per row, and attach metadata {source, row} to each.

  5. Hard:

    Build a loader that computes a SHA-256 hash of each document's cleaned text, skips exact duplicates, and prints a report: files loaded, pages with no text, duplicates skipped.

Interview questions

Why is document parsing often the biggest quality lever in RAG?

Every later step works on the extracted text. Scrambled reading order, merged columns, lost table structure or repeated boilerplate produce chunks that embed poorly and read badly, so retrieval returns the wrong things and the model gets confusing context. Fixing extraction often improves answers more than changing the embedding model or the LLM.

How do you handle scanned PDFs?

Detect them (the text extractor returns little or no text per page), then run OCR or a document-parsing service that renders the page as an image and recognises the text. Log OCR confidence and route low-confidence pages to review. Some multimodal LLMs can read page images directly, which helps for diagrams, at a higher per-page cost.

What metadata would you attach to each document and why?

Source identifier and path or URL (for citations and re-indexing), page or section (for precise citations), title, dates such as last-modified (for freshness and filters), owner or department and access-control info (for permission filtering), document type, and a content hash (for deduplication and change detection).

How would you represent a table for retrieval?

Keep each row together with its column headers, for example as 'Plan: Pro; Price: 30; Seats: 5' or a Markdown table row with the header repeated in the chunk. Optionally add the table caption. Large tables are better stored in a database and queried with SQL, with only a description of the table indexed for retrieval.

How do you remove boilerplate from web pages?

Parse the HTML, drop script, style, nav, header, footer and aside elements, prefer the main or article element, and remove lines that repeat across many pages of the same site. Main-content extraction libraries automate this.

Why normalise all loaders to one output shape?

So that chunking, embedding, storage and citation code is format-agnostic. Each new source type only needs a new loader that returns {text, metadata}, and the rest of the pipeline stays unchanged and testable.