Multimodal AI
Not just text: building apps that understand images, documents and audio. One of the most useful areas for real business automation.
What Multimodal AI is
A multimodal model understands more than one medium — not just text, but also images, documents, audio and even video. The current flagship models from the major providers are multimodal by design, which unlocks a set of business uses that used to need specialist software: reading a photographed invoice, analyzing a screenshot, transcribing a support call and summarizing it.
This turns manual, tedious processes (typing data from documents, sorting images) into automatic ones — which is why it's one of the highest-ROI areas.
Most data in the world isn't tidy text — it's documents, images and recordings. Multimodal AI "unlocks" that data for automation.
Images (Vision)
You send an image to the model and ask about it in natural language. Uses: analyzing screenshots, quality control, content recognition, accessibility. Example with OpenAI:
from openai import OpenAI
client = OpenAI()
# Set this from your provider's current model list — ids change often.
MODEL = "your-provider-vision-model-id"
resp = client.chat.completions.create(
model=MODEL,
messages=[{"role": "user", "content": [
{"type": "text", "text": "What is shown in the image? Return a short description."},
{"type": "image_url", "image_url": {"url": "https://.../photo.jpg"}}
# or base64: {"url": "data:image/jpeg;base64,...."}
]}],
)
print(resp.choices[0].message.content)
- Combine with structured output to get JSON (e.g. a list of detected items) instead of free text.
- High resolution costs more tokens — match the size to the need.
Documents & PDF — Document AI
One of the most in-demand uses: extracting structured data from documents (invoices, receipts, forms, contracts). Instead of traditional OCR + brittle rules, you send the document to a multimodal model and ask for JSON.
resp = client.messages.create(
model=MODEL, max_tokens=1024,
messages=[{"role":"user","content":[
{"type":"text","text":"Extract from the invoice: vendor, total, VAT, date. Return valid JSON only."},
{"type":"image","source":{"type":"base64","media_type":"image/jpeg","data": img_b64}}
]}],
)
# -> {"vendor":"...","total":1234,"vat":210,"date":"2026-08-01"}
- Require a schema and add validation — a blurry document can cause errors.
- For long documents — split into pages or use the provider's native PDF capability.
- Ready-made template: see Invoice & Finance Bot on the templates page.
Audio & transcription
For audio you use a transcription model (like Whisper) that converts speech to text, then feed the text to an LLM for summary/analysis. A common chain: recording → transcription → summary + action items.
# Step 1: transcription
audio = open("call.mp3", "rb")
transcript = client.audio.transcriptions.create(model="whisper-1", file=audio)
# Step 2: analysis with an LLM
summary = client.chat.completions.create(
model=MODEL,
messages=[{"role":"user","content": f"Summarize the call and give action items:\n{transcript.text}"}],
)
Uses: meeting summaries (see the AI Meeting Summary template), support-call analysis, generating transcripts for accessibility.
A vision model is not OCR, and the difference is the whole story
Document automation used to mean OCR plus rules: recognise the characters, then find the total because it sits in a box at a known position, or because it follows the word "Total". That approach is deterministic and brittle — it either works or visibly fails, and a supplier changing their invoice template breaks it on a Tuesday.
A multimodal model reads the document the way a person does. It does not care where the total is, it handles a layout it has never seen, and it copes with a phone photograph taken at an angle. That flexibility is the reason this technology replaced a category of specialist software in about two years.
What you give up is the part nobody mentions: guarantees. The old pipeline told you when it could not read something. The new one produces a confident answer either way — the same property that makes chat models hallucinate, applied to your accounts payable.
So the architecture that works is not "send the document, store the result". It is extract, then verify, then decide whether a human looks at it. The next three sections are that.
Validation is not optional, and it is cheap
Because the model will not tell you it was unsure, you build the check yourself. Most of it is arithmetic you already know.
- Cross-field consistency. Line items sum to the subtotal; subtotal plus tax equals the total; the tax rate is one of the rates that exist in your jurisdiction. This single check catches a large share of extraction errors, because a misread digit almost never survives the arithmetic.
- Schema validation. Require a strict shape and reject anything that does not parse — see structured outputs. A model asked for JSON will occasionally return prose around it.
- Range and sanity rules. A date in the future on a receipt, an invoice for a thousand times your normal amount, a supplier you have never used. These are business rules, not AI problems, and they are the ones that catch a genuinely wrong extraction that happens to be internally consistent.
- Known-value lookups. Match the supplier against your existing list rather than trusting the extracted name. Company names are exactly the sort of thing that gets nearly right.
- Two-pass agreement for high-value documents: extract twice, at temperature zero if available, and flag any disagreement between the runs. Doubles the cost of the cheapest step in the pipeline and finds the unstable fields.
Then route by confidence you computed: clean extractions go straight through, anything that failed a check goes to a person. That threshold is the product decision, and it is where this stops being a demo.
Dates, numbers and the errors you will not notice
The dangerous extraction errors are not the unreadable ones. They are the ones that produce a valid-looking value that is wrong.
- Dates. 01/02/2026 is two different days depending on where the document came from, and the model guesses from context that may not be there. Ask for the raw string as it appears and the parsed date, then resolve the format from the supplier's country rather than from the document.
- Decimal separators. 1.234 is one number in one convention and a thousand times that in another. Ask for the currency and the raw string, and normalise in your own code.
- Currency. A bare symbol is ambiguous across several currencies. Extract the symbol and the country, and never assume.
- Negative numbers and credits, which appear as parentheses, a minus, or a word — and a credit note booked as an invoice is an error that survives every arithmetic check.
- Similar fields. Invoice date against due date, order number against invoice number, subtotal against total. When two fields are adjacent and similar, ask for both explicitly rather than one.
The general defence: ask for the raw text alongside the interpreted value. It costs a few tokens, it makes every error traceable, and it lets you fix a systematic misreading in your own code instead of re-prompting.
Prompting for extraction
Extraction prompts have their own rules, and three instructions do most of the work.
Give it permission to fail. "If a field is not clearly readable, return null and do not guess" changes behaviour substantially, and it is what converts a silent wrong answer into a routable exception. Without it, the model does what it always does — produces the most plausible value.
Ask where it found each field. Requesting a short quote of the surrounding text alongside each value gives you something to check against, and it discourages invention because the model has to point at something.
Define the fields precisely. "Total" is ambiguous on a document that has a subtotal, a total before tax, a total due and a previous balance. Spell out which one you mean, in the words the documents use.
And keep the schema stable. A pipeline whose output shape changes between runs is one nobody downstream can rely on, which is the practical argument for enforced structured output rather than asking nicely for JSON.
Images: resolution, cost and what to send
Images are converted into tokens, and the count scales with size — so resolution is a cost dial as well as a quality one.
The instinct is to send the highest resolution available. Usually wrong. Providers process large images in tiles, so an oversized scan costs several times what a right-sized one does and frequently extracts no better. The useful discipline is to find, on a sample of your real documents, the smallest size at which accuracy stops improving — and then downscale everything to it.
Where resolution genuinely matters: small print, dense tables, handwriting, and stamps or signatures. Where it does not: a clear typed invoice, a screenshot, a photograph of a whiteboard. For a mixed pipeline, consider deciding per document rather than applying one setting to all of them.
Two practical notes. Rotate before sending — a sideways scan costs accuracy for no reason, and detecting orientation is cheap. And crop to the region you need when you know where it is; a cropped field extracts more reliably and costs less than the full page.
What still breaks
- Handwriting. Much better than it was, still the least reliable input, and worse for numerals than for words — which is the wrong way round for finance.
- Tables that span pages. Send pages separately and the model loses the header row; send them together and you may exceed what it attends to reliably. Extract the header once and pass it as context with each page.
- Multi-column layouts, where reading order is ambiguous and text from two columns can interleave.
- Poor scans — skew, shadow, glare from a phone camera. A quick quality check before the expensive call saves money and produces better exceptions.
- Documents in scripts the model saw less of. Test on your own samples rather than trusting a multilingual claim; the drop is real and uneven.
- Long documents. Detail in the middle of a very long input gets less reliable attention than material at either end. Chunk by page and extract per page.
Measuring whether it works
"It seemed accurate in testing" is how document pipelines reach production and then quietly cost somebody a quarter of manual corrections.
Build a small labelled set — fifty to a hundred real documents, spanning your actual suppliers and formats, with the correct values typed out by a person. It is a tedious afternoon and it is the only thing that lets you answer any later question.
Then measure per field, not overall. "97% accurate" is meaningless if the 3% is concentrated in the total. You will typically find that dates and supplier names are the weak fields while amounts are near-perfect, and that tells you exactly where the human review belongs.
Keep the set and re-run it whenever you change the prompt, the model or the image pipeline. Providers update models, and behaviour shifts — the set is how you find out it shifted before your finance team does.
Video, and why it is mostly frames
Video support sounds like a separate capability and is usually the same one applied repeatedly. Most systems sample frames — a few per second, or at detected scene changes — and reason over those images plus the audio transcript. Knowing that tells you how to use it and what it will miss.
It works well for anything whose meaning survives sampling: what is on screen, when a topic changes, whether a person appears, what a slide says, finding the moment something is mentioned. It works badly for anything between the frames — fast motion, a gesture, the precise instant of an event, counting things that move.
The cost follows directly from the sampling rate, and this is where video bills surprise people: an hour of footage at a few frames per second is thousands of images. Two habits keep it sane. Sample on scene changes rather than on a timer, which captures what matters and skips the static stretches. And use the transcript as the index — find the relevant minutes from the words, then look only at frames from those minutes. Most video questions are really transcript questions with a visual confirmation step.
These documents contain other people's data
An invoice carries names, addresses and bank details; a support recording carries whatever the customer said. Sending them to a third-party API is a processing decision with obligations attached wherever your customers are.
- Check retention and training. Whether the provider stores inputs and for how long, and whether business tiers exclude them from training. This is usually a setting or a plan, and it is worth confirming rather than assuming.
- Know where it is processed if you have data-residency commitments.
- Redact what you do not need. If the job is extracting the total, the bank details do not need to be in the image.
- Consider local models for genuinely sensitive material. Smaller open vision models now handle clean typed documents adequately, and nothing leaves your network.
- Log what you sent, so that answering "what did you process and when" is a query rather than an investigation.
Getting the audio chain right
The recording to transcript to summary chain is straightforward and has two failure points worth knowing.
Names and jargon are what transcription gets wrong, and the summary then reasons confidently about the wrong word. Most serious transcription services accept a vocabulary list — load your product names, your people and your industry terms before processing, not after reading a mangled transcript.
Who said what. Plain transcription produces a wall of text; speaker separation matters enormously for a sales call or a multi-person meeting, because "the customer said" and "we said" are different facts. Check whether your provider offers it before building on the assumption.
And as with documents, errors compound downstream: fix the transcript first, because every summary, action item and analysis below it inherits whatever was wrong.
When the old approach is still better
Worth saying, because "use a vision model" is now the reflex answer.
- Very high volume of one fixed format. A template-based extractor is far cheaper per document and deterministic. If you process fifty thousand identical forms, the flexibility is not worth paying for.
- When you need the same answer every time. Some regulated processes require reproducibility that a probabilistic system cannot promise.
- Machine-readable documents. If there is a structured PDF layer, an API or a data feed, use it. Photographing a screen to read it with AI is a fashionable way to add errors to data you already had cleanly.
- Barcodes, QR codes and MRZ. Purpose-built decoders with error correction, essentially free, and exact.
The strongest pipelines are usually hybrids: cheap deterministic extraction where the format is known, a vision model for the long tail of formats that would never justify a template, and a human on whatever failed a check.
Tips & cost
- Images cost tokens. High resolution = more expensive. Downscale images that don't need full resolution.
- Always validate. Extraction from a document can get a field wrong — verify (e.g. subtotal+vat=total) before feeding the system.
- Non-English documents. Models read most languages reasonably, but handwriting or a poor scan make it hard — test on a real sample.
- Combine with hallucination prevention: instruct the model to say "unreadable" instead of guessing a field.
- Gemini is especially strong at multimodal and long context — consider it for long documents/video (guide).
Next step
Build real document automation with a ready-made template, or strengthen your structured output.