Skip to content
Logo
←Applied AI Engineering
Chapter 3 of 23 · Level 1 · Foundations: How LLMs Work

What Models Take In

Multimodal input: how an image becomes patches then tokens, why images cost money, and when OCR still beats native vision.

  1. 3.1

    What Models Take In

    A model does not see your image. It receives patches turned into tokens, in the same stream as your text. Once you know that, image cost, resolution limits and the choice between OCR and native vision stop being mysteries.

    6 min read
    →
← Previous chapterHow Models Generate TextNext chapter →Reasoning and Test-Time Compute
© 2026 Said Mustafa Said
LinkedInGitHubEmail
Logo