Skip to main content
AI Engineering Multimodal AI
Level: Advanced Updated: September 2026

Multimodal AI

Not just text: building apps that understand images, documents and audio. One of the most useful areas for real business automation.

What Multimodal AI is

A multimodal model understands more than one medium — not just text, but also images, documents, audio and even video. The current flagship models from the major providers are multimodal by design, which unlocks a set of business uses that used to need specialist software: reading a photographed invoice, analyzing a screenshot, transcribing a support call and summarizing it.

This turns manual, tedious processes (typing data from documents, sorting images) into automatic ones — which is why it's one of the highest-ROI areas.

Why it's powerful for businesses

Most data in the world isn't tidy text — it's documents, images and recordings. Multimodal AI "unlocks" that data for automation.

Images (Vision)

You send an image to the model and ask about it in natural language. Uses: analyzing screenshots, quality control, content recognition, accessibility. Example with OpenAI:

from openai import OpenAI
client = OpenAI()

# Set this from your provider's current model list — ids change often.
MODEL = "your-provider-vision-model-id"

resp = client.chat.completions.create(
    model=MODEL,
    messages=[{"role": "user", "content": [
        {"type": "text", "text": "What is shown in the image? Return a short description."},
        {"type": "image_url", "image_url": {"url": "https://.../photo.jpg"}}
        # or base64: {"url": "data:image/jpeg;base64,...."}
    ]}],
)
print(resp.choices[0].message.content)

Documents & PDF — Document AI

One of the most in-demand uses: extracting structured data from documents (invoices, receipts, forms, contracts). Instead of traditional OCR + brittle rules, you send the document to a multimodal model and ask for JSON.

resp = client.messages.create(
    model=MODEL, max_tokens=1024,
    messages=[{"role":"user","content":[
        {"type":"text","text":"Extract from the invoice: vendor, total, VAT, date. Return valid JSON only."},
        {"type":"image","source":{"type":"base64","media_type":"image/jpeg","data": img_b64}}
    ]}],
)
# -> {"vendor":"...","total":1234,"vat":210,"date":"2026-08-01"}

Audio & transcription

For audio you use a transcription model (like Whisper) that converts speech to text, then feed the text to an LLM for summary/analysis. A common chain: recording → transcription → summary + action items.

# Step 1: transcription
audio = open("call.mp3", "rb")
transcript = client.audio.transcriptions.create(model="whisper-1", file=audio)

# Step 2: analysis with an LLM
summary = client.chat.completions.create(
    model=MODEL,
    messages=[{"role":"user","content": f"Summarize the call and give action items:\n{transcript.text}"}],
)

Uses: meeting summaries (see the AI Meeting Summary template), support-call analysis, generating transcripts for accessibility.

A vision model is not OCR, and the difference is the whole story

Document automation used to mean OCR plus rules: recognise the characters, then find the total because it sits in a box at a known position, or because it follows the word "Total". That approach is deterministic and brittle — it either works or visibly fails, and a supplier changing their invoice template breaks it on a Tuesday.

A multimodal model reads the document the way a person does. It does not care where the total is, it handles a layout it has never seen, and it copes with a phone photograph taken at an angle. That flexibility is the reason this technology replaced a category of specialist software in about two years.

What you give up is the part nobody mentions: guarantees. The old pipeline told you when it could not read something. The new one produces a confident answer either way — the same property that makes chat models hallucinate, applied to your accounts payable.

So the architecture that works is not "send the document, store the result". It is extract, then verify, then decide whether a human looks at it. The next three sections are that.

Validation is not optional, and it is cheap

Because the model will not tell you it was unsure, you build the check yourself. Most of it is arithmetic you already know.

Then route by confidence you computed: clean extractions go straight through, anything that failed a check goes to a person. That threshold is the product decision, and it is where this stops being a demo.

Dates, numbers and the errors you will not notice

The dangerous extraction errors are not the unreadable ones. They are the ones that produce a valid-looking value that is wrong.

The general defence: ask for the raw text alongside the interpreted value. It costs a few tokens, it makes every error traceable, and it lets you fix a systematic misreading in your own code instead of re-prompting.

Prompting for extraction

Extraction prompts have their own rules, and three instructions do most of the work.

Give it permission to fail. "If a field is not clearly readable, return null and do not guess" changes behaviour substantially, and it is what converts a silent wrong answer into a routable exception. Without it, the model does what it always does — produces the most plausible value.

Ask where it found each field. Requesting a short quote of the surrounding text alongside each value gives you something to check against, and it discourages invention because the model has to point at something.

Define the fields precisely. "Total" is ambiguous on a document that has a subtotal, a total before tax, a total due and a previous balance. Spell out which one you mean, in the words the documents use.

And keep the schema stable. A pipeline whose output shape changes between runs is one nobody downstream can rely on, which is the practical argument for enforced structured output rather than asking nicely for JSON.

Images: resolution, cost and what to send

Images are converted into tokens, and the count scales with size — so resolution is a cost dial as well as a quality one.

The instinct is to send the highest resolution available. Usually wrong. Providers process large images in tiles, so an oversized scan costs several times what a right-sized one does and frequently extracts no better. The useful discipline is to find, on a sample of your real documents, the smallest size at which accuracy stops improving — and then downscale everything to it.

Where resolution genuinely matters: small print, dense tables, handwriting, and stamps or signatures. Where it does not: a clear typed invoice, a screenshot, a photograph of a whiteboard. For a mixed pipeline, consider deciding per document rather than applying one setting to all of them.

Two practical notes. Rotate before sending — a sideways scan costs accuracy for no reason, and detecting orientation is cheap. And crop to the region you need when you know where it is; a cropped field extracts more reliably and costs less than the full page.

What still breaks

Measuring whether it works

"It seemed accurate in testing" is how document pipelines reach production and then quietly cost somebody a quarter of manual corrections.

Build a small labelled set — fifty to a hundred real documents, spanning your actual suppliers and formats, with the correct values typed out by a person. It is a tedious afternoon and it is the only thing that lets you answer any later question.

Then measure per field, not overall. "97% accurate" is meaningless if the 3% is concentrated in the total. You will typically find that dates and supplier names are the weak fields while amounts are near-perfect, and that tells you exactly where the human review belongs.

Keep the set and re-run it whenever you change the prompt, the model or the image pipeline. Providers update models, and behaviour shifts — the set is how you find out it shifted before your finance team does.

Video, and why it is mostly frames

Video support sounds like a separate capability and is usually the same one applied repeatedly. Most systems sample frames — a few per second, or at detected scene changes — and reason over those images plus the audio transcript. Knowing that tells you how to use it and what it will miss.

It works well for anything whose meaning survives sampling: what is on screen, when a topic changes, whether a person appears, what a slide says, finding the moment something is mentioned. It works badly for anything between the frames — fast motion, a gesture, the precise instant of an event, counting things that move.

The cost follows directly from the sampling rate, and this is where video bills surprise people: an hour of footage at a few frames per second is thousands of images. Two habits keep it sane. Sample on scene changes rather than on a timer, which captures what matters and skips the static stretches. And use the transcript as the index — find the relevant minutes from the words, then look only at frames from those minutes. Most video questions are really transcript questions with a visual confirmation step.

These documents contain other people's data

An invoice carries names, addresses and bank details; a support recording carries whatever the customer said. Sending them to a third-party API is a processing decision with obligations attached wherever your customers are.

Getting the audio chain right

The recording to transcript to summary chain is straightforward and has two failure points worth knowing.

Names and jargon are what transcription gets wrong, and the summary then reasons confidently about the wrong word. Most serious transcription services accept a vocabulary list — load your product names, your people and your industry terms before processing, not after reading a mangled transcript.

Who said what. Plain transcription produces a wall of text; speaker separation matters enormously for a sales call or a multi-person meeting, because "the customer said" and "we said" are different facts. Check whether your provider offers it before building on the assumption.

And as with documents, errors compound downstream: fix the transcript first, because every summary, action item and analysis below it inherits whatever was wrong.

When the old approach is still better

Worth saying, because "use a vision model" is now the reflex answer.

The strongest pipelines are usually hybrids: cheap deterministic extraction where the format is known, a vision model for the long tail of formats that would never justify a template, and a human on whatever failed a check.

Tips & cost

Next step

Build real document automation with a ready-made template, or strengthen your structured output.