Technology / Computer Vision / Edge AI

From Camera Feed to Actionable Intelligence.

NAZR combines computer vision, person tracking, identity recognition, pose analysis and visual understanding to transform CCTV video into structured events, alerts and operational intelligence.

Computer Vision
Edge AI
Person Tracking
Pose Analysis
Visual Understanding

Live Computer Vision Pipeline

PROCESSING

01

Detection

02

Tracking

03

Pose

04

Context

01 / Architecture overview

A video pipeline designed to turn pixels into events.

The processing architecture moves progressively from raw camera footage toward structured intelligence. Each stage contributes context before an event is confirmed and surfaced to the dashboard.

01

Camera Feed

Incoming CCTV or camera video provides the visual input to the system.

03

Person Tracking

Detected people are tracked across sequential frames.

05

Pose & Activity Analysis

Body keypoints, posture and movement provide additional activity context.

07

AI Event Confirmation

Signals are combined to determine whether an observed situation represents a meaningful event.

09

NAZR Dashboard

Events and operational information become visible through the NAZR dashboard.

02

Person Detection

People are identified within individual video frames.

04

Identity Recognition

When enabled, enrolled users can be recognised when image quality permits.

06

Visual Understanding

Broader scene context is added to help interpret what is occurring.

08

Alert / Activity Record

Confirmed events become structured records or alerts for operational use.

Pipeline principle: Rather than treating a single frame as the complete answer, NAZR builds context through detection, tracking, recognition, pose and visual understanding before an event is confirmed.

Computer vision layer

Multiple models contribute different pieces of visual intelligence.

Model-driven intelligence

Different models. Different jobs. One processing pipeline.

NAZR technology stack separates the core computer-vision problems into specialised processing stages. Detection identifies people, tracking maintains continuity, pose analysis interprets body position and visual understanding adds broader context.

02 / AI model stack

The models behind the intelligence.

This layer exposes the current model stack for technical evaluators, with each model assigned a specific role within the processing pipeline.

01

YOLO26L

Used for person detection, identifying people within incoming camera frames as the first computer-vision stage.

Person detection

03

InsightFace

Used for identity recognition for enrolled users. Recognition works best when the face is visible and image quality is sufficient. The technical team currently recommends roughly 20–40 clear enrollment images from different angles and lighting conditions.

Identity recognition

05

Moondream2

Adds broader visual context to the scene, supporting the system's ability to interpret what is happening beyond individual detections

Visual understanding

02

ByteTrack

Tracks people across sequential video frames, allowing the system to maintain continuity as individuals move through the scene.

Person tracking

04

YOLO26L-Pose

Provides body keypoints and posture information that can contribute to activity interpretation, including standing, sitting, lying and hand movement.

Pose & activity

03 / Cloud architecture

Edge-connected processing without putting the entire workload on the camera.

The current tested cloud architecture uses a Raspberry Pi as the edge gateway, with secure tunnelling into GPU infrastructure for AI processing and downstream event handling.

Current Tested Cloud Architecture

EDGE → CLOUD → DASHBOARD

01
CCTV Camera
Video source
02
DHCP Reservations
Network access
03
Raspberry Pi
Edge gateway
04
Cloudflare Tunnel
Secure connection
05
Runpod GPU
AI compute
06
AI Processing
Model stack
07
Dashboard
Events & alerts

01

Camera input layer

EDGE

Gateway layer

GPU

AI processing layer

API

Event delivery layer

04 / Local architecture

Keep core processing local when the environment requires it.

The local architecture removes the cloud GPU layer and performs the intended core processing directly on an NVIDIA GPU PC. The technical architecture is designed for environments where Internet access is not required for the core local processing workflow.

Local processing

Ethernet CCTV / USB Camera → NVIDIA GPU → AI Processing → Event Record → Dashboard

Local Processing Architecture

LOCAL-FIRST WORKFLOW

01
Camera
Ethernet / USB
02
Local compute
Local compute
03
AI Processing
Vision models
04
Event Record
Structured output
05
Dashboard
Operational view

Architecture distinction: the cloud configuration and local configuration are separate deployment patterns. The local architecture is intended to support core processing without requiring Internet access.

05 / Detection capabilities

What the system currently detects.

NAZR current tested capabilities are separated from expandable roadmap capabilities so technical and operational teams can clearly distinguish what is available today from future development areas.
Current tested capabilities

Detection and activity intelligence available today.

Fall detection

Chest-holding / chest-pain gesture detection

Taking medicine

Eating

Drinking

Patient identification

Person tracking

Doctor / nurse visits

Selected staff care activities

Expandable / roadmap

Additional capabilities identified for future expansion.

Sleep monitoring

Bed occupancy

Abnormal behavior

Respiratory distress

Nurse visit tracking

Emergency-call integration

Mobile alerts

Hospital-management-system integrations

AI interpretation layer

Computer vision turns continuous visual information into structured operational signals.

Designed for expansion

A technology stack that can grow with the detection library.

The architecture is intentionally modular. New detection and activity capabilities can be added without changing the fundamental flow from camera input through visual analysis and event confirmation.

Technical evaluation

See how the architecture fits your environment.

Whether your deployment requires cloud GPU processing, local NVIDIA infrastructure or a combination of both, the NAZR architecture can be evaluated around your existing camera environment and operational requirements.

Architecture summary

Camera → Vision → Context → Event → Dashboard

A modular computer-vision pipeline connecting camera feeds with detection, tracking, identity, pose, visual understanding and structured event records.