NAZR combines computer vision, person tracking, identity recognition, pose analysis and visual understanding to transform CCTV video into structured events, alerts and operational intelligence.
Detection
Tracking
Pose
Context
The processing architecture moves progressively from raw camera footage toward structured intelligence. Each stage contributes context before an event is confirmed and surfaced to the dashboard.
Incoming CCTV or camera video provides the visual input to the system.
Detected people are tracked across sequential frames.
Body keypoints, posture and movement provide additional activity context.
Signals are combined to determine whether an observed situation represents a meaningful event.
Events and operational information become visible through the NAZR dashboard.
People are identified within individual video frames.
When enabled, enrolled users can be recognised when image quality permits.
Broader scene context is added to help interpret what is occurring.
Confirmed events become structured records or alerts for operational use.
Multiple models contribute different pieces of visual intelligence.
This layer exposes the current model stack for technical evaluators, with each model assigned a specific role within the processing pipeline.
Used for person detection, identifying people within incoming camera frames as the first computer-vision stage.
Used for identity recognition for enrolled users. Recognition works best when the face is visible and image quality is sufficient. The technical team currently recommends roughly 20–40 clear enrollment images from different angles and lighting conditions.
Adds broader visual context to the scene, supporting the system's ability to interpret what is happening beyond individual detections
Tracks people across sequential video frames, allowing the system to maintain continuity as individuals move through the scene.
Provides body keypoints and posture information that can contribute to activity interpretation, including standing, sitting, lying and hand movement.
The current tested cloud architecture uses a Raspberry Pi as the edge gateway, with secure tunnelling into GPU infrastructure for AI processing and downstream event handling.
Camera input layer
Gateway layer
AI processing layer
Event delivery layer
Ethernet CCTV / USB Camera → NVIDIA GPU → AI Processing → Event Record → Dashboard
Computer vision turns continuous visual information into structured operational signals.
The architecture is intentionally modular. New detection and activity capabilities can be added without changing the fundamental flow from camera input through visual analysis and event confirmation.
Whether your deployment requires cloud GPU processing, local NVIDIA infrastructure or a combination of both, the NAZR architecture can be evaluated around your existing camera environment and operational requirements.
A modular computer-vision pipeline connecting camera feeds with detection, tracking, identity, pose, visual understanding and structured event records.