What is a vision-language model? VLMs for video analytics

Written by Alphan Eker, Co-FounderUpdated

A vision-language model (VLM) is an AI model that reads images or video together with language: it can describe what is happening in a scene, answer questions about it and explain its reasoning in plain words. Where classic computer vision detects objects and stops at "a person, a forklift", a VLM can read the action and its context — who is doing what, and whether it is normal. In video analytics, that means findings written as sentences with an evidence clip, instead of boxes and labels someone still has to interpret.

What a vision-language model (VLM) is

A vision-language model (VLM) is a model that takes visual input — an image, a sequence of frames, a video clip — and works with it through language. It can be asked a question about the scene ("Is the guard on this machine open?"), told what to look for ("a solder bridge between two pins"), and it answers in words: what it sees, and why it reaches its conclusion.

The useful part is the combination. The visual side gives the model the scene; the language side gives it a way to carry knowledge about that scene — a safety rule, an acceptance criterion, a store procedure — and to report what it found in a form a person can read and check. That is why VLMs are a natural fit for problems an experienced person solves by looking: a safety professional on the floor, a loss-prevention auditor at the register, an inspector at a review station.

Where the term comes from

The name says what the model joins: vision (images and video) and language (text). Large language models work with text only; a VLM adds the ability to read visual input, so the same kind of reasoning can be applied to what a camera sees.

How a VLM differs from classic computer vision

Classic computer vision — object detection, classification, segmentation — answers a fixed question for each frame: which objects are present, and where? A detector trained on people and forklifts reports "person" and "forklift" with a box around each. Whatever happens next — whether the two are about to meet, whether the person is a visitor or a driver — is left to rules written on top of those boxes, or to a person watching the screen.

  • Objects versus actions. A detector says what is in the frame. A VLM can read what is being done: an operator reaching into a machine, a cashier handing over an order, a forklift reversing toward someone behind it.
  • Labels versus explanations. A detector outputs a class and a score. A VLM can output a sentence: what happened, where, and why it matters.
  • Fixed classes versus described targets. A detector knows the classes it was trained on. A VLM can be told, in language, what to look for and what the rules of the site are — which is how domain knowledge is put on top of the model.
  • Single frames versus context. A detector judges each frame on its own. A VLM can read the scene in context: the zone, the task, what happened a few seconds earlier.

Classic computer vision still has its place: it is fast, light and predictable for well-defined questions. For a neutral side-by-side, see VLM vs rule-based video analytics.

Classic computer vision answers "what is in the frame?" A vision-language model can answer "what is happening here, and is it normal?" — and say why.

How VLMs read video

Video understanding is more than reading frames one by one. Most of what matters on a floor or at a register is an action that unfolds over a few seconds, so a VLM-based system reads video in windows of time rather than single frames.

  1. 01

    Windows, not frames

    The stream is split into short windows, so the before and after of an action are read together. That is what makes direction, sequence and intent visible: a sale paid and handed over before it is voided, a person walking toward an intersection a forklift is about to enter.

  2. 02

    A prefilter picks the moments

    Most of any video stream is uneventful. A lighter first stage watches every stream and passes the windows that deserve attention to the deeper analysis, so the expensive reading is spent where it matters.

  3. 03

    Domain knowledge on top

    The model reads the window with the knowledge it was configured with: the site's safety rules, the board's acceptance criteria, the store's procedures. The same core can act as a safety expert in one place and an inspector in another.

  4. 04

    A finding in language, with evidence

    The result is a sentence a person can read, attached to the time-stamped clip it was read from, and kept in a memory that can be asked about later.

What VLMs make possible on the floor

Because a VLM reads actions and context rather than objects alone, problems that used to need a person watching the footage can be read on every camera, every hour. Some examples, each covered in its own guide:

Workplace safety

A near miss between a forklift and a pedestrian is a few seconds of converging paths that end without harm — and usually without a record. A VLM can read the route and the direction together and record the moment; see near miss detection with AI video. The same reading in context is what lets PPE be judged against the zone and the task, so a visitor and a welder are not held to the same rule; see PPE compliance with AI video.

The register

A POS log says what was entered; the camera shows what happened. A VLM can read the action at the register — a scan, a skipped item, a refund with no one at the counter — and set it against the POS line from the same second. See how to catch employees stealing from the cash register.

Visual quality control

On an electronics line, a VLM can take a second look at the locations an AOI machine flags and read each one the way an experienced inspector would; see AOI false call reduction. The same approach applies to incoming quality control and to packing and dock checks, where a skipped step or a damaged pallet is a visual question.

Limits, and how to work with them

A VLM is a strong reader of scenes, not a guarantee. Knowing its limits is part of using it well.

  • It reads only what the image shows. A blind corner with no camera on it, a joint hidden under a component, an action blocked by a shelf: none of these can be read. Coverage is a camera-placement question first.
  • It has to be proven on your own footage. Lighting, camera angles and the way work is done differ from site to site. A pilot on your own cameras or images, judged against criteria agreed before it starts, is the only fair test.
  • Findings are for people to review. Each finding should come with its evidence clip or image, so a person can check it before acting. A single event is a reason to look, not a verdict.
  • It costs more compute than a simple detector. Reading windows of video with a large model takes more processing than counting boxes. Systems handle this with a prefilter, so deep analysis runs only on the moments that matter.
  • Footage is sensitive. Video of people at work raises privacy questions. On-premise deployment, anonymization and analysis focused on the relevant zone are design choices to ask about, not add-ons.

What to look for in a VLM-based solution

Whichever vendor you talk to, these are fair questions to ask about a video analytics system built on vision-language models:

  • Works on your existing cameras. Can it connect to the IP cameras you already have, or does it need new hardware?
  • Explains every finding. Does each result come as a readable explanation with an evidence clip, or as a bare label and score?
  • Takes your domain knowledge. Can your rules, procedures and acceptance criteria be built into how it reads the scene — and how much work is it to change them?
  • Handles false alarms openly. Ask how findings are reviewed by a person, and how single events are separated from repeating patterns.
  • Can be piloted on your own footage. Your cameras, your floor, your criteria — not a demo video.
  • Keeps your data with you. Is there an on-premise option where footage never leaves the facility, and does it meet GDPR and other data-protection requirements?
  • Remembers. Can past findings be searched and asked about, so repeats by zone, shift or line can be seen?
  • Gets results to the right role. Instant alerts, periodic digests and dashboards, each for the person who acts on them.

How Primarch uses VLMs

This section is about our product. Primarch builds on and enhances vision-language models. Primarch Vision is the core: it reads the scene the way a person would and writes what it finds in plain language. Every Primarch expert — from workplace safety to the register — is this core in a different configuration, and Custom VLM Solutions configure it for problems that can be seen but are not in a catalog.

  • Holistic scene reading. The stream is split into real-time windows; a prefilter hands the moments that matter to the deep analysis layer. The output is a reading at the level of the action, not a label.
  • Experts defined by configuration. An expert is defined by domain documentation, regulations and thresholds on top of the same core. We don't train models from scratch; what makes the expert specific is the knowledge.
  • Natural-language findings with evidence clips. Every finding is a human-readable interpretation of the scene, with a time-stamped video segment attached.
  • Queryable event memory. Findings are written into a facility memory that can be asked about in plain language.
  • Alerts, digests and dashboards. Instant notifications, periodic digests and role-based dashboards.
  • Any visual input. Besides RTSP cameras, photo archives, periodic captures, drone footage and microscope output can feed an expert.
  • Real-time, multi-camera. Dozens of streams watched in parallel, with the windows that deserve attention prioritized automatically.

Putting it to work

  1. 01

    Discovery

    We define the problem and the success criteria together: which scene, which decision, which action.

  2. 02

    Configuration

    An expert profile is built from a domain knowledge base, rule templates and thresholds; targeted fine-tuning if needed.

  3. 03

    Pilot

    Validation on real footage from your own site, tested together against measurable criteria. Depending on scope, discovery to pilot typically takes a few weeks.

  4. 04

    Deployment

    On-premise or cloud rollout, team training and a continuous improvement loop.

  • Your existing cameras. Primarch connects to RTSP-capable IP cameras you already own; no new cameras or special hardware.
  • Your data stays with you. With on-premise deployment, footage never leaves the facility, and compliant anonymization is part of the design.

Frequently asked questions

What is a vision language model used for?

To read images and video the way a person would and report in plain language. In video analytics that means recognizing actions in context — a near miss on a warehouse floor, a refund entered with no one at the register, a defect on a circuit board — and writing each finding as a sentence with an evidence clip.

VLM vs computer vision: what's the difference?

Classic computer vision detects and classifies objects frame by frame: "a person, a forklift". A VLM reads the scene with language: it can describe the action and its context, be told what to look for, and explain its conclusion. Classic detectors remain useful for simple, well-defined questions; VLMs open up problems that need judgment.

How is a VLM different from an LLM?

A large language model (LLM) works with text. A vision-language model also takes images or video as input, so it can apply the same kind of language-based reasoning to what a camera sees.

Does a VLM work with existing cameras?

It can. Primarch connects to existing RTSP-capable IP cameras with no new hardware. What a VLM can read still depends on what the camera sees, so coverage is checked during the pilot.

Does our footage have to leave the facility?

No. With on-premise deployment, footage is processed inside your facility and never leaves it.

How long does it take to put a VLM to work on a new problem?

Because Primarch experts are defined by configuration rather than trained from scratch, a new use case typically goes live in days. A custom solution typically takes a few weeks from discovery to pilot, depending on scope.

Related scenarios

READY?

Ready to make your cameras think?

A 15-minute live demo — with a scenario tailored to your industry.