What on-premises video management systems, as a category, are not built to do, written for IT and security teams deciding whether to extend an existing recorder estate or layer something over it. Every constraint carries its evidence, its operational impact, and its workaround.

This page describes a category, not a company, so there is no single vendor's documentation to cite and none is invented here. Every constraint below describes behavior that on-premises video management products publish about themselves in their own supported device lists, license terms, lifecycle notices and product documentation, reviewed on August 17, 2026. Where the category documents nothing, this page says so rather than guessing, and every line has to be checked against the release notes and license terms for your own product and version, because these vary more between products than the category description suggests. Spot AI sells a competing platform, so nothing here rests on an anonymous source or an aggregated user rating. Spot AI has published its own page about this category, and this review found nothing in it that the category's own documentation contradicts, so this banner carries no correction.
Every box describes behavior common to on-premises video management products rather than one vendor's documentation. Capture and retention are the category's strength. What happens after capture is where the gap sits.
A constraint list is only useful next to an honest account of the category.
Capture every camera, every hour, and still have the footage weeks later when somebody finally asks. Recording servers write full-resolution video to local storage sized to the retention policy, capacity is a purchasing decision rather than a plan tier, and no window shrinks because a subscription changed. Estates that have been burned by retention moving under them tend to defend this architecture hardest, and they are right to.
Established products maintain supported hardware lists covering thousands of camera models across dozens of manufacturers, updated per software version, so a procurement team can check a real inventory against a real document before committing. Analog channels reach the system through encoders or a hybrid recorder. On a fleet assembled over ten years from four vendors, that list is the single most useful artifact in an evaluation.
The system runs on infrastructure you own, video never has to leave the building, and an internal team can audit the storage, the access log and the network path without a vendor in the room. A recording server on the local network does not care about an outage at the carrier, so a site with a poor uplink loses nothing. For a hospital, a school district, a utility or a defense supplier that is often a condition of purchase rather than a preference.
Each one is a consequence of how the category is designed rather than a defect in any product. What matters is whether it collides with your starting point.
Product documentation across the category describes playback, timeline scrubbing, bookmarking, export and motion alerting as the operator workflow, and motion detection as the analytic that ships. What is not described anywhere in a base product is a set of named detections that raise the events worth seeing without somebody going to look for them.
The system answers what happened once somebody knows roughly where to look and has time to scrub, which is why investigations in these estates are routinely measured in hours. Most security teams can name the last three that cost half a day each, and none of them can name the incidents nobody thought to search for, which is the larger number.
Time your last three investigations and put the hours into the business case rather than arguing the point in the abstract. Tighten motion zones and bookmarking discipline for the short term, then evaluate a detection layer that reads the same cameras rather than assuming the recorder has to be replaced to get one.
Motion alerting is close to universal, though what it covers varies by product. Deeper detection generally arrives as an add-on module from the same vendor or as a partner product, each carrying its own licensing, its own server requirement and its own integration work, and each published on its own terms rather than inside the recorder's license.
The platform decision turns out to be the easy half. A 200-camera estate that wants protective equipment checks, forklift proximity and dock dwell is looking at a second procurement, a second vendor relationship and an integration project, and the quality of the result is then set outside the recorder vendor's release cycle.
Write down the five things you need spotted, then ask which ship in the version you already own and which are a second purchase. Price the analytic, the server it needs and the integration together over three years, and confirm the analytic supports the camera models you actually have before anything is signed.
The documented answer to an event is an operator watching a monitor or reading a motion alert, with rules that can trigger recording, a bookmark, an email or an output relay. What happens on the ground, a talk-down, a strobe or a horn sequence, is not part of what these products describe themselves as doing. The documentation covers what an alert does inside the system, not what happens at the fence.
For a site whose real requirement is that an intruder leaves rather than that the footage exists, this is the difference between a record and a response. It is also why so many estates run a separate alarm or monitoring contract alongside the recorder, and pay twice for two halves of the same event.
Name the deterrence requirement as its own line rather than assuming the recorder covers it. Where sites already have a public address system or standard speakers, ask each option on the shortlist which of them drives that hardware directly, and put the alarm contract renewal date next to the video decision so the two are argued together.
The architecture is per site by design: recording servers or a recorder in each building, storage sized locally, and administration handled by whoever owns that building. Whether a single multi-site view exists at all depends on the specific product and edition, and federation is published as a feature of the higher editions rather than as a category-wide assumption.
One building with one recording server is manageable. Twenty buildings means twenty upgrade windows, twenty storage calculations and a user list per site, and every version move becomes a coordinated change across servers and clients rather than a maintenance window. That work rarely appears in the original business case and never goes away.
Count the upgrade windows per year and the hours per window, then put that number in front of whoever owns the budget. Check whether your edition supports federation before assuming a central view is possible, standardize versions across sites before adding more, and decide deliberately which sites need a local server at all.
Products in this category publish lifecycle and end-of-support dates for their versions and their server platforms, and license terms that pair per-channel licenses with an annual maintenance or support agreement. Storage sizing is published as a calculation against resolution, frame rate and retention rather than as a fixed figure, because it is one.
Recording servers reach end of support and need a capital refresh, storage grows every time resolution or retention does, and the maintenance agreement renews on a date somebody in finance already has in a spreadsheet. A five-year total that only counts the original licenses will be wrong by the size of a refresh cycle.
Build a five-year model with the refresh year explicitly in it, and ask your integrator for the published end-of-support date of the exact server platform and software version you run. Recalculate storage at the resolution you will actually run in year three rather than the one you run today.
The same five constraints in one view, sized to paste into an evaluation document.
Swipe the table sideways to see every column.
This page describes a category rather than one vendor. Every line reflects behavior on-premises video management products publish about themselves, reviewed on August 17, 2026, and must be verified against your own product and version.
If the requirement is that every camera records reliably, that retention meets a policy somebody will audit and that video never leaves the building, this category answers it and has answered it for years. Estates with a data-residency condition, a poor uplink or an internal team that wants to inspect the storage themselves have a genuinely strong reason to stay exactly where they are, and a finance function that plans on five-year capital cycles often prefers a per-channel license and an integrator relationship to a subscription that renews forever.
Spot AI is built for the half of the problem this category leaves open, and it is designed to sit on the cameras that are already recording rather than to replace them. 15+ pre-trained Video AI Agents cover vehicle break-in, fire, cash register theft, after-hours intrusion, personal protective equipment, forklift near-miss, falls and crowding in hazard zones, and Iris builds anything else in natural conversation in about eight minutes, so the detections arrive with the platform instead of as a second procurement with its own server.
The architecture keeps the part of the category that was never the problem. Full-resolution video stays on the Intelligent Video Recorder in the building and only event metadata crosses the network, so residency and uplink answers stay as strong as they were, while one dashboard covers every site instead of one console per building. Any ONVIF or RTSP IP camera works at full functionality and legacy analog comes in through the IVR, and an event at 03:00 is answered on site with talk down, strobes and horns through standard speakers.
The useful question is not whether on-premises recording was a mistake, but how many separate purchases sit between the recorder and something actually firing.
None of that makes Spot AI the right answer for every estate. It is a subscription platform rather than a perpetual per-channel license, so an organization whose finance model is capital with a known five-year refresh is being asked to change shape, not just vendors. Spot AI states camera compatibility by protocol rather than publishing a supported hardware list running to thousands of models, which is genuinely less specific than what an established video management product gives a procurement team. The third-party module ecosystems around the larger products also reach requirements no single vendor's roadmap will. A constraint list earns its keep by lining each option's shape up against your starting point, and here that starting point is a recorder estate that mostly works.
A live pilot on your existing cameras answers in a week what a spec sheet cannot.
Customer-reported outcomes from named Spot AI customers.
Silver Bay Seafoods replaced fragmented legacy camera systems across 22 locations, including remote Alaska facilities, and lifted operational efficiency 15%.
Bridge33 Capital standardized more than 25 assets on one platform, cut footage search from hours to minutes and dropped the weekly manual camera audit.
Liberty-Perry School District resolves an incident in about five minutes, after evaluating ten systems before choosing Spot AI.
"Spot AI has replaced all of our legacy systems and enables us to view and review all of our sites from one central location...It was an easy choice to go with Spot AI."
Five, and they are category-wide rather than one vendor's. The system waits to be asked, because playback and motion alerting are the published workflow and no named detection set ships in a base product. Deeper analytics are a separate purchase with their own server and integration. Response at the moment of an incident is not part of the category. Administration multiplies by building, since servers, storage and users sit per site. And the cost is periodic, because servers reach end of support and storage grows with resolution and retention.
Yes, and better than almost anything else in this market. Established products publish supported hardware lists covering thousands of camera models across dozens of manufacturers, updated per software version, and analog channels reach the system through encoders or a hybrid recorder. If anything, camera compatibility is the category's strongest published claim, which is why a recorder estate is usually worth keeping rather than replacing.
Motion detection ships across the category and is genuinely useful for zones and simple triggers. Named detections such as protective equipment compliance, forklift proximity or dock dwell generally arrive as an add-on module or a partner product, each with its own license, its own server requirement and its own integration project. The practical test is to ask which of your five required detections ship in the exact version you already own, and price the rest properly.
Often, and the honest comparison has to include the periodic costs on both sides. On the existing estate that means the server refresh year, storage growth as resolution and retention rise, per-channel licenses and the annual maintenance agreement, plus any analytics module and its server. A layer that reads the cameras already recording is usually cheaper than a rip and replace, so the real choice is rarely keep versus replace, it is keep versus keep and add.
It depends on which half of the job is failing. If recording and retention work and the problem is that nobody has time to watch, the answer is a detection layer over the same cameras rather than a new recorder. A platform such as Spot AI fits that shape: any ONVIF or RTSP camera connects as it is, legacy analog comes in through the Intelligent Video Recorder, full-resolution video stays in the building, and 15+ pre-trained Video AI Agents run across every site from one dashboard.