What Is Video Analytics? An Introduction to the Technology
What is video analytics and how does it work? The fundamentals of AI-powered camera analysis technology and where it is used across industry and retail.
Translated from the Turkish original · Türkçe aslı
Video analytics, also called image analytics, is the technology that turns camera footage from a recording to be watched into a data source that produces decisions. Deep learning models detect objects in every frame, track them over time and generate events according to defined rules: a violation alert, counting data, a quality record. Instead of an operator constantly watching a screen, the system notifies you only when something meaningful happens.
How does video analytics work?
The chain consists of the same four links in every project. The first is the image source: the IP camera on site delivers footage as an RTSP stream, usually compressed with H.264 or H.265. The second is the processing unit: the edge server inside the facility decodes the stream and feeds the frames to the model running on a GPU; we describe this architecture in detail in the RTSP streams and edge GPU article. The third is the model: it draws a box around classes such as person, vehicle, forklift, hard hat, vest and smoke in every frame; the family most used in the field is single-stage detectors (object detection with YOLO). The fourth is the rule engine: it makes what the model sees meaningful in the facility's own terms.
The real work happens in the fourth link. "There is a person in the frame" means nothing on its own; "a person without a hard hat has been standing in a restricted zone for a set time while a forklift is operating" is an event. The rule engine holds the zone drawings, time thresholds, shift calendar and counting lines. When an event occurs, a time-stamped clip, a dashboard metric and an API message are generated; the operator sees the alert together with its clip and makes the decision.
How it differs from classic motion detection
Motion detection in recorders looks at pixel changes: a tarpaulin flapping in the wind, vehicle headlights or a cloud's shadow also trigger alarms. Video analytics, by contrast, classifies what is moving; it tells a person from a shadow and a forklift from a hand cart. In addition, an event is not generated from a single frame; the condition must persist for a certain number of frames and a certain time. This time window is the core mechanism that keeps false alarms from wearing out the operator. False positives are reviewed daily throughout the pilot and thresholds are tuned to the site.
Why now?
Three curves crossed at the same time: GPU costs fell, open model architectures matured, and millions of cameras were already installed in facilities. Today even a 10-year-old IP camera becomes a real-time AI eye with an edge server added behind it, without any hardware replacement.
A fourth factor needs to be added: training models on field data is now something small teams can do. The training loop that only large research groups could run a few years ago now runs on images collected in the field, careful labelling and a single GPU server. This means site-specific problems that off-the-shelf products do not cover have become solvable.
Use cases
In occupational safety, PPE and zone monitoring; in security, fire, perimeter and licence plate recognition; in manufacturing, OEE, quality control and counting; in retail, people counting, heatmaps and queue management. At CX Teknoloji we deploy field-proven solutions across all of these areas.
The areas may look different, but the output types fall into three categories. The first is the alert: an instant notification when a rule is broken, such as entry without a hard hat, entering a restricted zone, or smoke. The second is measurement: continuously turning a process into numbers, such as machine downtime, waiting time at the checkout, or the number of people entering a store. The third is evidence: a clip showing when and how an event happened, for a shipment dispute, a workplace accident investigation or an audit. The same camera can serve more than one output; for example, a camera at a warehouse entrance can both raise alerts on forklift–pedestrian proximity and store evidence clips of the shipping area. The first question when designing a project is which of these three the need is; everything from camera angle to reporting frequency changes accordingly.
Off-the-shelf solution or custom model?
For common needs such as PPE, fire and people counting, pre-trained models are calibrated to the site. But some problems exist only in that facility: a defect type seen only on that line, a movement that is meaningful only on that site. In that case custom video analytics comes in. The process starts with a feasibility study: what can be measured with the existing footage, what accuracy is realistic and what the labelling cost will be are set out clearly. A typical custom project goes live within 12–14 weeks, and the biggest item in that time is not the model but the data; capturing rare cases determines the quality of the project.
How is accuracy measured?
The accuracy figure in a brochure tells you nothing about your site; light, camera angle, distance and workflow differ in every facility. Accuracy is a field value, not a lab value: throughout the pilot, real events are labelled manually and compared with the system's detections. What counts as a "correct detection", how many false alarms are acceptable and on which dataset the test is run are defined in writing at the outset. With custom models the curve rises during the pilot; the example curve on our page follows the path 71% → 94%, because false detections from the field are added to the training data every week.
Data privacy and KVKK
Because camera footage can contain personal data, it is assessed under Law No. 6698 on the Protection of Personal Data (KVKK, Türkiye's Personal Data Protection Law). In our installations the default architecture is that footage is processed on the edge server inside the facility; no raw video is sent to the cloud. Face recognition is not used, analytics outputs are anonymous, and event clips are automatically deleted at the end of a defined retention period. Matters such as the privacy notice, retention period and access rights are decisions for the organisation's own legal team; we have gathered the general framework in the KVKK and image processing article. This information is not a substitute for legal advice.
What do you need to get started?
The first step is not buying new cameras but reviewing your existing camera inventory. Which cameras can be used as they are, which need a changed angle, where a new camera is needed: we carry out this assessment free of charge. Then the metric to be measured is defined, zones are drawn together with the site team and the pilot begins. We walk through the order of the steps and what is decided at each one in the installation process article.
Frequently asked questions
Can video analytics be used with our existing cameras?
In most cases, yes. IP cameras that provide an RTSP stream can be analysed with an edge server added behind them; even 10-year-old cameras can fall into this group. What matters is not the brand but whether the object to be measured is large and sharp enough in the image, and whether the angle and light are suitable. Before installation we assess your camera inventory free of charge and report the suitable points and, if needed, any camera that should be added.
Is footage sent to the cloud, and what is the position under KVKK?
In the default installation, footage is processed on the edge server inside the facility and no raw video is sent to the cloud. Face recognition is not used; analytics outputs are anonymous and event clips are automatically deleted at the end of a defined retention period. Decisions such as the privacy notice and retention period are made by the organisation's legal team; we configure the technical side according to those decisions. This information is not legal advice.
How accurate will the system be?
Accuracy depends on the site and is measured during the pilot, on your cameras. Real events are labelled manually, compared with the system's detections and the result is reported. What counts as a correct detection and the acceptable level of false alarms are defined in writing at the start of the pilot. Thresholds are calibrated to the site throughout the pilot; with custom models, accuracy rises as false detections from the field are added to training.
Does it integrate with our existing VMS, MES or alarm system?
Yes. Events are published via REST API and MQTT, so they can connect to your existing MES, ERP, VMS and alarm infrastructure. Because a standard RTSP stream is used on the camera side, your existing recording system keeps working as it is; the analytics only reads a copy of the stream. On the production side, MES, ERP/SAP and Power BI integration is also possible.
How long does installation take?
For off-the-shelf solutions, installation including camera assessment and zone drawing is typically completed within a week; the 30-day pilot that follows delivers measurable results with existing cameras. The first week of the pilot is usually set aside for calibration; sources of false alarms are flagged and thresholds are tuned. For projects requiring a site-specific model, data collection and labelling are added, so the typical deployment time is 12–14 weeks.