VLM vs rule-based video analytics: a neutral comparison

Written by Alphan Eker, Co-FounderUpdated

VLM vs rule-based video analytics is a choice between two ways of turning camera footage into findings. Rule-based analytics detects objects and fires when a fixed rule is met — a line crossed, a zone entered, a count exceeded — which makes it predictable, light and easy to audit, but blind to context. VLM-based analytics uses vision-language models that read a window of video together with domain knowledge, so it can recognize the action and the situation and write the finding in plain language, at the cost of more compute and a need to be validated on your own footage. Many sites are best served by using each where it fits, or by combining them.

What rule-based video analytics is

Rule-based video analytics is the long-established way of getting alerts out of cameras. A detector finds objects in each frame — a person, a vehicle, a forklift, a hard hat — and draws a box around them. On top of those boxes, someone writes rules for each camera: an alarm when a person crosses this line, when a vehicle enters this zone, when someone stays in an area longer than a set time, when more than a set number of people are counted. The rule is the logic; the detector supplies the inputs.

  • Zones and line crossing — an alarm when an object of a given class enters an area or crosses a virtual line.
  • Object classes and thresholds — an alarm when a class is present or absent above a confidence threshold, such as a person without a detected helmet.
  • Dwell time — an alarm when an object stays in an area longer than a set time.
  • Counting — people or vehicles counted across a line, or an alarm when a count goes over a limit.

What VLM-based video analytics is

A vision-language model (VLM) reads images or video together with language. Instead of only labelling objects frame by frame, it reads a window of time and can describe what is happening: who is doing what, and whether that fits the situation. Domain knowledge — the rules a safety professional, an inspector or a loss-prevention auditor works to — is put on top of the model as configuration, so the reading follows that domain. The output is a finding written in plain language, with the clip it was read from. For a longer introduction, see what a vision-language model is.

Primarch builds on and enhances vision-language models, so the descriptions of VLMs in this guide are our own experience. The comparison below is meant to be fair to both approaches: each has a place.

VLM vs rule-based video analytics, side by side

  1. 01

    What it detects

    Rule-based: objects, positions, counts and times — what a rule can be written for. VLM-based: actions and situations in context, such as an operator reaching into a machine that is still energized, or a refund entered with no customer at the counter.

  2. 02

    Setup

    Rule-based: zones, lines and thresholds are drawn and tuned camera by camera. VLM-based: an expert profile is configured with domain knowledge, documentation and thresholds; site-specific tuning still happens, but there is less to draw per camera.

  3. 03

    False alarms

    Rule-based: a rule fires whenever its condition is met, whatever the context — a parked forklift and a turning one look the same to a zone rule, so normal work can produce a stream of alarms. VLM-based: context is designed to separate normal work from risk, but readings can still be wrong in other ways, so a review policy matters.

  4. 04

    What it misses

    Rule-based: anything nobody wrote a rule for. VLM-based: anything the camera can't see clearly — an occluded hand, a dark corner, a detail too small for the resolution.

  5. 05

    Explanations

    Rule-based: the explanation is the rule itself ("zone 3 entered"), which is easy to trace but says little about what happened. VLM-based: a plain-language sentence about what happened and why it matters, with an evidence clip.

  6. 06

    Compute and cost

    Rule-based: light; can often run on the camera or a small edge device, with very low latency. VLM-based: heavier; needs more compute, and a reading of a window of video takes longer than checking a single rule.

  7. 07

    Predictability and audit

    Rule-based: deterministic — the same input gives the same alarm, and a rule is easy to audit. VLM-based: less deterministic, so thresholds, validation on your own footage and human review of findings carry more weight.

  8. 08

    When things change

    Rule-based: a moved camera, a new layout or a new task means redrawing and retuning rules. VLM-based: the reading follows the scene, and new situations are handled by updating the configuration — but changes still need to be checked on real footage.

When rules are enough, and when context is needed

Rules are a good fit when the event is fully described by position, presence, count or time, and the context doesn't change its meaning:

  • Perimeter intrusion out of hours — anyone crossing the fence line at night is worth an alarm.
  • People and vehicle counting — footfall, queue length, gate throughput.
  • Simple line crossing — a vehicle passing a stop line, a person entering an area that no one should ever enter.
  • Low-power sites — where only a small edge device is available and a fast, simple alert is what's needed.

Context is needed when the same picture can mean normal work or a real problem, depending on what is happening:

  • PPE in context. Whether missing equipment matters depends on the zone and the task — a visitor on a walkway and a welder at work are not the same thing. See PPE compliance with AI video.
  • Forklift and pedestrian near misses. Distance alone doesn't separate a picker working beside a parked forklift from one stepping into the path of a moving one; direction and movement over time do. See near miss detection with AI video.
  • A void after payment at the register. The POS log shows a void; only the video shows whether the sale was paid and handed over first. See POS exception reporting vs AI video.
  • AOI false calls. A machine measuring a joint against fixed limits flags acceptable joints along with real defects; a second look that reads the call in context is designed to separate the two. See AOI false call reduction.

A useful test: if a person looking at the clip would need to know what is happening before deciding whether it's a problem, a rule alone will struggle. If position, presence, count or time decides it, a rule may be all you need.

Combining the two

The choice doesn't have to be either-or. A light first stage can watch every stream and pick out the moments that deserve attention, and a deeper stage can then read those moments in context. This keeps compute focused where it is needed and keeps the speed of a simple check while adding the reading a rule can't make. It is how the Primarch platform is built: the stream is split into real-time windows, and a prefilter hands the moments that matter to the platform's deep analysis layer.

Existing rule-based alarms can also stay where they work well — a perimeter fence at night, a people counter at a gate — while context-heavy questions go to a VLM-based layer.

What to look for in a solution

Whichever approach or vendor you consider, these are fair questions to ask about video analytics software:

  • Works on your existing cameras. Can it connect to the IP cameras you already have, and what compute does it need on site?
  • Fits the question you are asking. Is the event you care about described by position, count or time — or does it depend on the action and the context?
  • Handles false alarms openly. Ask how alarms are reviewed by a person, how normal work is kept from triggering them, and how single events are separated from repeating patterns.
  • Explains every finding. Does a finding say what happened, with a clip, or only which rule fired?
  • Can be piloted on your own footage. Your cameras, your layout, your shifts — judged against criteria agreed before the pilot starts, not a demo video.
  • Adapts to change. What happens when a camera moves, a layout changes or a new task starts — redraw every rule, or update a configuration?
  • Respects privacy. Is there an on-premise option where footage never leaves the site, is anonymization applied, and does it meet GDPR and other data-protection requirements?
  • Gets results to the right role. Instant alerts, periodic digests and dashboards — each person sees what they need to act on.

How Primarch Vision and Custom VLM Solutions help

This section is about our product. Primarch Vision is the platform core every Primarch expert runs on. It reads the scene holistically: it interprets actions and context, writes its findings in natural language and stores them in a queryable memory. Custom VLM Solutions put your domain knowledge on top of that core for problems that can be seen but are not in a catalog.

  • Prefilter plus deep analysis. The stream is split into real-time windows; a prefilter hands the moments that matter to the deep analysis layer, and dozens of streams can be watched in parallel.
  • Configuration, not rule drawing. An expert is defined by configuration — domain documentation, regulations and thresholds. The core stays the same, which is why a new use case is configured rather than trained from scratch.
  • Natural-language findings with evidence clips. Every finding is a human-readable reading of the scene with a time-stamped clip, so everyone sees the same thing.
  • Queryable event memory. Findings are written into a facility memory you can ask in plain language.
  • Alerts, digests and dashboards. Instant notifications, periodic digests and role-based dashboards.
  • Existing cameras, your data. It connects to the RTSP cameras you already own, and with on-premise deployment footage never leaves your facility, with compliant anonymization as part of the design.

Putting it to work

  1. 01

    Discovery

    We define the problem and the success criteria together: which scene, which decision, which action.

  2. 02

    Configuration

    An expert profile is built from a domain knowledge base, rule templates and thresholds; targeted fine-tuning if needed.

  3. 03

    Pilot

    Validation on real footage from your site, tested together against measurable criteria.

  4. 04

    Deployment

    On-premise or cloud rollout, team training and a continuous improvement loop.

Frequently asked questions

What is rule-based video analytics?

It is video analytics that detects objects in each frame and raises an alarm when a fixed rule is met: a line crossed, a zone entered, an object present or absent, a dwell time or a count exceeded. Rules are written and tuned for each camera.

What is the difference between VLM and rule-based video analytics?

Rule-based analytics checks objects against fixed conditions; it is fast, light and predictable, but doesn't know what is happening. VLM-based analytics reads a window of video with domain knowledge, so it can recognize the action and the situation and explain the finding in plain language; it needs more compute and validation on your own footage.

Will VLMs replace rule-based systems?

Not everywhere. For events fully described by position, count or time — a fence crossed at night, people counted at a gate — rules remain a good, light choice. VLMs add the most where context decides whether something is a problem. Many setups combine the two: a light first stage and a deeper reading of the moments that matter.

Which one produces fewer false alarms?

It depends on the question. For simple events, a well-tuned rule can be very reliable. Where the same picture can mean normal work or a real risk, rules tend to fire without context, and reading the scene in context is designed to reduce that noise. In either case, ask how findings are reviewed and test on your own footage.

Does VLM-based video analytics work with existing cameras?

Primarch connects to your existing RTSP-capable IP cameras; no new cameras or special hardware are required. What it can read still depends on what each camera sees.

Related scenarios

READY?

Ready to make your cameras think?

A 15-minute live demo — with a scenario tailored to your industry.