Audio Analytics on the Cameras You Already Have

A camera only sees what is in frame. Its microphone hears the whole room. AudioTrack turns the audio stream a camera already carries into four event classes, alerted and searched like any visual detection, with no acoustic sensor, no array and no extra cabling.

The Module

AudioTrack, a Detector on the Camera Channel

AudioTrack classifies the sound arriving on a camera’s audio stream into four event types: gunshot, glass breaking, scream, and loud sound. Each is switched on separately, carries its own priority, and lands on the same event list, alarm monitor and archive as every visual detection on that camera. It is a detector rather than a module in its own right, so it layers onto whatever else the channel is already running.

The hardware is the camera. Any camera with a microphone can run it: the audio arrives on the stream as AAC, G.711 mu-law or PCM, audio streaming is optional per channel, and there is no acoustic sensor to procure, mount, power or maintain. A control room that already owns cameras with microphones already owns an acoustic sensor network. This page is about switching it on where sound is worth listening to.

Sound covers what sight misses. A gunshot behind a pillar, glass breaking in a stockroom the camera faces away from, a scream in a stairwell at the edge of the field of view: none is in frame, all are heard. That is why the gunshot class also sits on the weapon detection page, alongside GunTrack, as the second of three independent ways to raise the same alarm.

Four Classes

What It Listens For

Gunshot

The event class that pages a response. It runs alongside visual weapon detection on the same channel, so a weapon fired out of frame is still a sound that was heard.

See Weapon & Gunshot Detection

Glass Breaking

A shopfront, a stockroom window, a display case, a vehicle window on a forecourt. The sound arrives before anyone is in frame, which is the point.

Scream

A person in distress outside the camera’s view, in a corridor, a stairwell or a platform end. The frame, the camera and the time reach an operator to verify.

Loud Sound

The catch-all. Anything loud the module cannot identify as one of the other three is reported rather than dropped: a door, a dropped pallet, or something worth a call, for a person to decide.

Setup

Three Settings per Camera

Everything an operator sets lives on the camera’s Detectors tab, and nothing is trained or tuned site-wide.

  1. 01

    Turn On the Classes You Want

    Each of the four event types is ticked separately and assigned its own priority. A retail stockroom might run glass breaking and loud sound only; a campus corridor runs all four.

  2. 02

    Set the Sensitivity

    A single sensitivity slider sets how loud an event has to be relative to the background noise on that camera. It is the setting the pilot spends its time on, because the right value in a quiet lobby is the wrong one on a station concourse.

  3. 03

    Route the Event

    The resulting events are scoped by camera and type onto the alarm monitors of the roles that own them, and pushed to real-time crime centers, dispatch workflows and the Sover secure messenger with the frame and the location attached. Afterwards they are searchable under Additional events like any other detection.

Fit

Where It Fits, and Where It Does Not

Runs On
Any connected camera that has a microphone, with audio enabled on the stream. AAC is recommended; G.711 mu-law and PCM are also supported. No separate sensor hardware.
Compute
The guide rates the module’s computational complexity Low, so it adds little to a server already running visual analytics on the same channel.
Combines With
Several modules and detectors run on one camera at the same time. GunTrack plus AudioTrack on one channel is the normal weapon configuration; AudioTrack beside intrusion detection or a StreamVLM™ prompt is just as valid.
Not Recommended
Noisy outdoor locations such as underground stations and railway terminals. Constant ambient noise is what the sensitivity threshold has to work against, and we would rather say so here than discover it together in a pilot.

What We Do Not Claim

Read This Before the Pilot

  • IREX publishes no range, accuracy or latency figure for acoustic detection. The commitment is the one every module carries: measured on your own cameras during the pilot, against criteria agreed in writing, with the measured numbers going into the contract.
  • IREX does not claim to locate a sound. There is no triangulation across microphones; the event carries the camera that heard it, so coverage follows your camera plan rather than a sensor grid.
  • It is not a speech system. The four classes are the whole vocabulary, and everything else is either Loud sound or silence.
  • Detections are signals for a person to verify. Nothing responds autonomously, and IREX products are expressly not life-safety systems.

FAQ

Do we need gunshot sensors on poles?

No. AudioTrack runs on any camera that already has a microphone, so there is no acoustic array to procure, mount, power, or maintain, and no separate sensor network to run alongside the cameras. The audio arrives on the camera stream as AAC, G.711 mu-law, or PCM, and it is optional per channel, so you turn it on where sound is worth listening to and leave the rest of the estate as it is.

How far away can it hear, and how often is it right?

IREX publishes no range, accuracy, or latency figure for acoustic detection, and we are not going to invent one for a website. The commitment is the same one the visual detectors carry: we measure it on your own cameras during the pilot, against criteria agreed in writing beforehand, and the measured numbers go into the contract. Detection works from the microphone in the camera, so coverage follows your camera plan rather than a sensor grid.

Will it work on a busy street or a station?

It is not recommended in noisy outdoor locations such as underground stations and railway terminals, and we say so on the product page rather than in a pilot report. The sensitivity setting is relative to the background noise on that camera, so a channel with constant ambient noise leaves little room between the background and an event. Indoor and quieter outdoor cameras are where it belongs.

Does it record or understand conversations?

No. The module classifies sound into four event types and nothing else. It does not transcribe speech and it does not turn audio into a searchable text record. Anything loud that it cannot classify as a gunshot, breaking glass, or a scream is reported simply as Loud sound.

Can one camera run audio and video analytics together?

Yes, and that is the normal configuration. Several video analytics modules and detectors run on the same camera at the same time, and AudioTrack is a detector, so it sits beside GunTrack, a perimeter rule, or a StreamVLM™ prompt on one channel. The two then cover each other: an incident out of frame is still a sound that was heard, and a silent one is still an object that was seen.

Does this run on its own or does someone have to watch it?

It runs on its own and alerts a person. Detections are signals for a human to verify: no response runs autonomously, and the verification and decision are logged against the same Case ID as the detection.

Turn On the Microphones You Already Own

Pick the cameras that face away from the trouble. Those are the ones a microphone changes.