StreamVLM™: Write a Sentence, Get a Detector

Conventional video analytics needs months of data collection, labeling, and training for each new use case. StreamVLM™ uses a Vision-Language Model that already understands images and language, so a detector you describe in a sentence starts working immediately.

The Long Tail

The Conditions No Vendor Ships out of the Box

Every city has a list of things it needs to see that no analytics vendor sells a module for: a person lying on the ground, a damaged fence, illegal dumping behind a depot. Individually none of them justifies a training program. Together they are most of what a municipality actually worries about, and eight more are further down this page, each one written the way an operator would type it.

StreamVLM™ closes that gap. An operator writes the condition in plain English, the engine evaluates live camera frames against the prompt, and matches raise real-time alerts like any other event. Each camera channel supports multiple simultaneous prompt-based detectors, each with its own confidence threshold, alert cooldown, and event type.

StreamVLM™ is shipping in beta version on selected instances. It is included in the standard IREX price: there is no separate license, module, or subscription fee. The only additional cost is a local GPU node for VLM inference where a deployment requires one.

Example

Person down in a Metro Station

A Detector Nobody Trained

"Alert when a person is lying on the ground." No dataset of people collapsing on platforms exists to train on, and collecting one would be neither practical nor decent. A Vision-Language Model does not need one: it already understands the scene, so the prompt is the specification.

An elderly woman in a quilted coat lies on her side at the foot of a bench on a Stockholm metro platform, her rollator tipped over and shopping spilled beside her, with nobody nearby.
One prompt, one event class. The platform reports a person on the floor; what it turns out to be is for the operator to decide.

On the Map

The Same Sentence, Anywhere You Run It

One location at a time: the sentence an operator there would type, and the three rows the platform writes back for it. Some are prompt-defined StreamVLM™ detectors, the rest the trained modules and archive search they feed. Switch auto play off to read one, or open a match to see the frame at full resolution.

San Diego, United States
SEARCH
Case ID #IRX-2026-0103 · provided Searching…
MATCHES
Searched 3,000+ cameras in 0.4 seconds

How It Differs

One Engine Instead of a Module per Problem

No Training Pipeline

No data collection, no labeling, no per-use-case model. A new detector is a sentence, and it works from the moment you save it.

One Engine, Unlimited Scenarios

Designed to replace and unify several legacy modules, including fire detection, video-quality monitoring, and tamper detection.

Tunable per Detector

Confidence threshold, alert cooldown, and event type are set individually, so a noisy condition does not flood the operator.

Logged like Everything Else

Every prompt, analyzed frame, alert, and agent action is logged with timestamp, operator identity, and Case ID, under role-based access.

Feeds Search and Dashboards

StreamVLM events flow into the same three modes as classic modules: real-time alerts, investigations over the archive, and big-data export.

Bias Safeguards Apply

A prompt-defined detector sits inside the same ethics framework as the trained modules, including the narrow-constraints rule.

FAQ

Is StreamVLM available on our instance?

It is shipping in beta on selected instances. Whether a specific deployment is in scope is confirmed with IREX engineering rather than promised in advance.

Does it cost extra?

No. StreamVLM is included in the standard IREX price with no separate license, module, or subscription fee. Where a deployment needs a local GPU node for VLM inference, that node is sized and priced with IREX engineering.

How many detectors can one camera run?

Several. A camera channel carries multiple prompt-defined detectors at once, each with its own confidence threshold, alert cooldown, and event type. IREX does not yet publish a per-camera ceiling: how the density scales depends on the customer’s hardware, and it is measured on your own cameras during the pilot.

Can it replace our existing analytics modules?

It is designed to unify several of them, including fire detection, video-quality monitoring, and tamper detection. The trained high-frame-rate detectors remain the right tool for tracking people and vehicles at speed.

Which model is behind StreamVLM, and what can it reach?

An open-weight vision-language model published by a third party and used as published: IREX does not fine-tune it or alter its weights, pins it by version and checksum, and serves it only on designated endpoints inside your instance, on the instance’s own GPU node. No frame, prompt or verdict leaves the instance, and no third-party cloud inference is used on any instance in any region. The model makes no tool calls, keeps no memory between frames, has no access to the archive, the event database, watchlists, other cameras or the network, and has no biometric function: it judges scenes and conditions and cannot identify a person. IREX does not publish which model it is; the identity, version and checksum are disclosed to a customer under NDA, because a named model is an attack surface and a fact that goes stale. Part B of the public AI Model Governance Policy sets all of this out.

Tell Us What You Need to See

Bring the conditions your city actually worries about. If a prompt can describe it, we can usually show you a detector for it in the pilot.

Book a Demo