StreamVLM™: Write a Sentence, Get a Detector

Conventional video analytics needs months of data collection, labeling and training for each new use case. StreamVLM™ uses a Vision-Language Model that already understands images and language, so a detector you describe in a sentence starts working immediately.

Get in Touch

The Long Tail

The Conditions No Vendor Ships out of the Box

Every city has a list of things it needs to see that no analytics vendor sells a module for: a person lying on the ground, a damaged fence, illegal dumping behind a depot. Individually none of them justifies a training program. Together they are most of what a municipality actually worries about and the image-led library below brings 24 such conditions together with the camera view and the page that explains each one.

StreamVLM™ closes that gap. An operator writes the condition in plain English, the engine evaluates live camera frames against the prompt and matches raise real-time alerts like any other event. Each camera channel supports multiple simultaneous prompt-based detectors, each with its own confidence threshold, alert cooldown and event type.

StreamVLM™ is included in the standard IREX price: there is no separate license, module, or subscription fee. The only additional cost is a local GPU node for VLM inference where a deployment requires one.

Example

Person Down in a Metro Station

A Detector Nobody Trained

"Alert when a person is lying on the ground." No dataset of people collapsing on platforms exists to train on and collecting one would be neither practical nor decent. A Vision-Language Model does not need one: it already understands the scene, so the prompt is the specification.

A man lies on his side in a sleeping bag against the wall of a metro connecting passage, with a suitcase and carrier bag beside him.
One prompt, one event class. The platform reports a person on the floor; what it turns out to be is for the operator to decide.

Use-Case Library

24 StreamVLM Use Cases and Related Analytics

Explore 24 prompt-defined conditions, each with an illustrative camera image and a dedicated page. A related trained two-wheeler workflow rounds out the visual library.

How It Differs

One Engine Instead of a Module per Problem

No Training Pipeline

No data collection, no labeling, no per-use-case model. A new detector is a sentence and it works from the moment you save it.

One Engine, Unlimited Scenarios

Designed to replace and unify several legacy modules, including fire detection, video-quality monitoring and tamper detection.

Tunable per Detector

Confidence threshold, alert cooldown and event type are set individually, so a noisy condition does not flood the operator.

Logged like Everything Else

Every prompt, analyzed frame, alert and agent action is logged with timestamp, operator identity and Case ID, under role-based access.

Feeds Search and Dashboards

StreamVLM events flow into the same three modes as classic modules: real-time alerts, investigations over the archive and big-data export.

Bias Safeguards Apply

A prompt-defined detector sits inside the same ethics framework as the trained modules, including the narrow-constraints rule.

Tell Us What You Need to See

Bring the conditions your city actually worries about. If a prompt can describe it, we can usually show you a detector for it in the pilot.

Book a Demo

Tell us roughly how many cameras you operate, where the data has to stay and the three or four things you most need to detect. We will scope a walkthrough on your own footage.

If the form is unavailable, email [email protected].

Common Questions

Is StreamVLM available on our instance?

StreamVLM runs on a GPU node in the customer instance. Whether a specific deployment is in scope is confirmed with IREX engineering rather than promised in advance.

Does it cost extra?

No. StreamVLM is included in the standard IREX price with no separate license, module, or subscription fee. Where a deployment needs a local GPU node for VLM inference, that node is sized and priced with IREX engineering.

How many detectors can one camera run?

Several. A camera channel carries multiple prompt-defined detectors at once, each with its own confidence threshold, alert cooldown and event type. IREX does not yet publish a per-camera ceiling: how the density scales depends on the customer’s hardware and it is measured on your own cameras during the pilot.

Can it replace our existing analytics modules?

It is designed to unify several of them, including fire detection, video-quality monitoring and tamper detection. The trained high-frame-rate detectors remain the right tool for tracking people and vehicles at speed.

Which model is behind StreamVLM and what can it reach?

An open-weight vision-language model published by a third party and used as published: IREX does not fine-tune it or alter its weights, pins it by version and checksum and serves it only on designated endpoints inside your instance, on the instance’s own GPU node. No frame, prompt or verdict leaves the instance and no third-party cloud inference is used on any instance in any region. The model makes no tool calls, keeps no memory between frames, has no access to the archive, the event database, watchlists, other cameras or the network and has no biometric function: it judges scenes and conditions and cannot identify a person. IREX does not publish which model it is; the identity, version and checksum are disclosed to a customer under NDA, because a named model is an attack surface and a fact that goes stale. Part B of the public AI Model Governance Policy sets all of this out.