Gunshot
The event class that pages a response. It runs alongside visual weapon detection on the same channel, so a weapon fired out of frame is still a sound that was heard.
See Weapon & Gunshot DetectionA camera only sees what is in frame. Its microphone hears the whole room. AudioTrack turns the audio stream a camera already carries into four event classes, alerted and searched like any visual detection, with no acoustic sensor, no array and no extra cabling.
The Module
AudioTrack classifies the sound arriving on a camera’s audio stream into four event types: gunshot, glass breaking, scream, and loud sound. Each is switched on separately, carries its own priority, and lands on the same event list, alarm monitor and archive as every visual detection on that camera. It is a detector rather than a module in its own right, so it layers onto whatever else the channel is already running.
The hardware is the camera. Any camera with a microphone can run it: the audio arrives on the stream as AAC, G.711 mu-law or PCM, audio streaming is optional per channel, and there is no acoustic sensor to procure, mount, power or maintain. A control room that already owns cameras with microphones already owns an acoustic sensor network. This page is about switching it on where sound is worth listening to.
Sound covers what sight misses. A gunshot behind a pillar, glass breaking in a stockroom the camera faces away from, a scream in a stairwell at the edge of the field of view: none is in frame, all are heard. That is why the gunshot class also sits on the weapon detection page, alongside GunTrack, as the second of three independent ways to raise the same alarm.
Four Classes
The event class that pages a response. It runs alongside visual weapon detection on the same channel, so a weapon fired out of frame is still a sound that was heard.
See Weapon & Gunshot Detection →A shopfront, a stockroom window, a display case, a vehicle window on a forecourt. The sound arrives before anyone is in frame, which is the point.
A person in distress outside the camera’s view, in a corridor, a stairwell or a platform end. The frame, the camera and the time reach an operator to verify.
The catch-all. Anything loud the module cannot identify as one of the other three is reported rather than dropped: a door, a dropped pallet, or something worth a call, for a person to decide.
Setup
Everything an operator sets lives on the camera’s Detectors tab, and nothing is trained or tuned site-wide.
Each of the four event types is ticked separately and assigned its own priority. A retail stockroom might run glass breaking and loud sound only; a campus corridor runs all four.
A single sensitivity slider sets how loud an event has to be relative to the background noise on that camera. It is the setting the pilot spends its time on, because the right value in a quiet lobby is the wrong one on a station concourse.
The resulting events are scoped by camera and type onto the alarm monitors of the roles that own them, and pushed to real-time crime centers, dispatch workflows and the Sover secure messenger with the frame and the location attached. Afterwards they are searchable under Additional events like any other detection.
Fit
What We Do Not Claim
Where It Runs
The same four classes, a different reason on each estate. These are the verticals where a camera microphone earns its keep.
A gunshot or a scream in a corridor no camera faces, on the cameras a school already owns, alongside the visual detector.
Read More →Shots fired out of frame in a public space, reaching officers with the camera and the time attached.
Read More →One more event class on the wall, verified with the frame rather than waiting for a call.
Read More →Breaking glass in a stockroom or a shopfront after hours, across a dispersed estate nobody is standing in.
Read More →A scream or a report inside a concourse at full load, when the visual detector is looking at a crowd.
Read More →What happens between rounds, heard on the cameras that were already there.
Read More →No. AudioTrack runs on any camera that already has a microphone, so there is no acoustic array to procure, mount, power, or maintain, and no separate sensor network to run alongside the cameras. The audio arrives on the camera stream as AAC, G.711 mu-law, or PCM, and it is optional per channel, so you turn it on where sound is worth listening to and leave the rest of the estate as it is.
IREX publishes no range, accuracy, or latency figure for acoustic detection, and we are not going to invent one for a website. The commitment is the same one the visual detectors carry: we measure it on your own cameras during the pilot, against criteria agreed in writing beforehand, and the measured numbers go into the contract. Detection works from the microphone in the camera, so coverage follows your camera plan rather than a sensor grid.
It is not recommended in noisy outdoor locations such as underground stations and railway terminals, and we say so on the product page rather than in a pilot report. The sensitivity setting is relative to the background noise on that camera, so a channel with constant ambient noise leaves little room between the background and an event. Indoor and quieter outdoor cameras are where it belongs.
No. The module classifies sound into four event types and nothing else. It does not transcribe speech and it does not turn audio into a searchable text record. Anything loud that it cannot classify as a gunshot, breaking glass, or a scream is reported simply as Loud sound.
Yes, and that is the normal configuration. Several video analytics modules and detectors run on the same camera at the same time, and AudioTrack is a detector, so it sits beside GunTrack, a perimeter rule, or a StreamVLM™ prompt on one channel. The two then cover each other: an incident out of frame is still a sound that was heard, and a silent one is still an object that was seen.
It runs on its own and alerts a person. Detections are signals for a human to verify: no response runs autonomously, and the verification and decision are logged against the same Case ID as the detection.
Keep Reading
Pick the cameras that face away from the trouble. Those are the ones a microphone changes.