Anyone who has run support for appliances, cars or electronics knows the conversation. “There’s a light blinking.” Which light? “The one on the left, near the bottom.” What colour? Three messages later someone asks for a photo, and the photo is blurred. Customers have been sending videos into chat for years because it is the obvious thing to do. Until recently, the AI on the other end could not watch them.
This week we shipped video understanding in chat. A customer sends a clip of up to a minute; the system watches it, reads any display or error code, listens to what is said, and the assistant answers with that in hand. Building it took less time than deciding where it should live, and that decision is the part worth writing about.
Two ways to give a conversation eyes
The first option is the elegant one. Move the whole conversation onto a model that takes video natively. The answering model then sees the actual footage alongside the history and the retrieved knowledge, and can look at it again when the customer asks a follow-up. Technically it is the stronger design.
The second option is plainer. Keep the conversation exactly where it is, and add a separate step in front of it: a multimodal model watches the video and writes a description, and that description is handed to the answering model as context. The model holding the conversation never sees a single frame.
We measured what each would cost against our own traffic — about 1.3 million chat turns a month — rather than against a pricing page.
Measured input, with prompt caching applied. The jump in January is an introductory price ending.
Moving every conversation to the video-native model would have raised the chat bill from roughly $1,100 a month to roughly $5,900 — about $4,800 more today, and closer to $10,700 more once the introductory price ends. The video feature itself, at ten thousand one-minute clips a month, costs about $64. We would have been paying seventy-five times the cost of a feature to change the engine under every conversation that does not use it.
So video is a side path. It is off by default, switched on per assistant, and metered per tenant. Nothing about the conversations that never see a video has changed. It is the same shape of decision as keeping expensive work away from turns that do not need it, which we wrote about in where the AI bill actually goes.
The price of the cheaper design
The trade-off deserves to be said out loud, because it is real. If the answering model only ever sees a description, then the description is the video as far as the conversation is concerned. Anything the describer leaves out never existed. Anything it gets wrong cannot be checked downstream, because nothing downstream has seen the footage. And a follow-up like “what does the light on the right mean?” is answered from the stored description unless the clip is sent to be described again — so we keep the video and allow exactly that.
That makes one property of the describing model matter more than any other, and it is not the one we expected.
The model that made up an error code
We tested candidate describers on six real appliance-repair videos — a washing machine, a gas range, a leaking sink, a circuit board, an oven display, a toilet — and asked each to identify the appliance, the fault, and any readable text.
The model we chose read the manufacturers’ logos, and caught an oven clock ticking over from 11:26 to 11:27. On one clip it said plainly that the display was too pixelated to read. The cheapest alternative — attractive because it needed no new integration work — looked at that same display and reported a programme code, a sixty-minute timer and a sixty-degree temperature. None of it was there. A third, a small open-weights model, described the clips as if they were still photographs: it saw objects, not events.
In a design where the describer is the only witness, a model that invents an error code is not usable at any price. The customer would be told, with confidence, how to fix a fault their machine never reported. It is also the condition that makes the side-path design safe at all: an honest describer feeding a strong conversational model is acceptable; the same pipeline with a describer that fills gaps is not.
Only one thing moves the bill
The second finding saved us from building something useless. The natural instinct with video is to shrink it before sending — lower the resolution, compress the file. We tested that directly: the same clip at 360p, 720p and 1080p produced identical token counts, and so did real 4K footage. File sizes from 1 MB to 30 MB made no difference. Nine container formats made no difference. The audio track costs nothing extra; it is already inside the per-second rate.
Duration is the only lever. The model normalises everything internally, so a downsampling step would have cost engineering time and infrastructure and reduced the bill by exactly zero — while degrading the image twice, which is the last thing you want when the whole point is reading a small digit off a small screen. The model’s own high-detail mode was no better a deal: over three times the tokens per second, and in testing it did not read a single character the low-detail mode missed.
Two levers we deliberately left alone. Sampling fewer frames per second cuts the video tokens on a thirty-second clip by more than half, but a fault that shows for a moment — a single flash of a warning light — can fall between frames, and we have not validated it on real footage. And the model’s internal reasoning cannot be switched off: every setting we tried still produced roughly 300 to 550 reasoning tokens per call, billed at the output rate. Our cost estimate assumes output rather than measuring it, and we say so in the budget.
What generalises
Three things carry beyond video. When a capability is needed by a small share of traffic, bolt it on beside the main path rather than moving the main path to reach it; the premium otherwise lands on every conversation. Before optimising a cost, find out empirically which input actually moves it — the obvious lever is often priced at zero. And when a component is the only one that sees the evidence, its most important property is honesty about what it could not see. Accuracy on a benchmark does not measure that. Putting an unreadable screen in front of it does.
