Counting and identifying food plant workers on the servers the plant already owns
How Gezora.ai designed an on premises computer vision system for Unilever Pakistan Foods, and brought its cost down by more than 90% without dropping a single function.
Sufi Inam Ul Hassan
AI Engineer15 minute read

Summary
Unilever Pakistan Foods asked Gezora.ai for a camera system with two jobs at its food plant. The first was to count workers: those standing in marked zones inside three assembly halls, and those passing through three entry and exit points. The second was to say who each of those people is. Results had to reach administrators on a web dashboard and on their phones, and the plant set an accuracy target of 92% to 98%.
Our first complete design met the brief on paper and cost about PKR 73 million in its first year. More than a third of that went on two new GPU servers. Over four rounds of revision we rebuilt the design around equipment the plant already owned. We replaced premium cameras with NDAA compliant models at about a quarter of the price, chose open source models with commercial licences wherever accuracy allowed, and reduced model training to a single rented GPU used once. The final proposal costs PKR 5.8 million with cameras and installation, or PKR 1.9 million if the plant buys and installs the cameras itself. Both versions keep every function the brief asked for.
| 15 | 0 | 92% to 97% | 24 weeks |
|---|---|---|---|
| cameras: 6 for counting, 9 for faces | new servers bought | lower cost than the first design | planned from contract to handover |
The client and the brief
Unilever Pakistan Foods runs a food plant with three assembly halls and three entry and exit points. Each hall has a spot marked on the floor with a yellow box, and the plant wanted to know how many people were standing inside that box at any moment, ignoring everyone outside it. At each entry and exit point it wanted counts of people coming in and going out, per hour and per day, plus a running figure for how many people were inside the building.
Counting was only half of the request. For every person the cameras saw, the system had to try to match them against the plant's enrolled workers and record a name, department, location and time. Anyone who did not match had to be flagged as unknown, and repeat sightings of the same unknown person had to be grouped under one temporary identity so that a visitor walking past three cameras would not be counted as three visitors. Unilever would supply photos and videos of every worker for enrollment.
Two lines in the brief shaped most of what followed. The plant runs its own wifi, its own servers and its own databases, and the system had to run on that infrastructure. And the accuracy target of 92% to 98% had been written as a single number for a system that does several very different things.
| Item | Detail |
|---|---|
| Client | Unilever Pakistan Foods, food plant |
| Industry | Food manufacturing |
| Scope | 3 assembly hall zones, 3 entry and exit points, 15 cameras |
| Gezora.ai role | Research, solution architecture, camera layout, costing and delivery plan |
| Engagement period | September to October 2026 |
| Delivery plan | 24 weeks from contract to handover, starting with a pilot at one entry point |
Why the problem was harder than it looked
The brief read like one problem. It is really two. A person detector such as YOLO can count bodies in a zone with high accuracy, but it has no idea who anyone is. Identification needs a separate chain of models: one that finds faces, one that judges whether a face image is clear enough to use, one that turns a face into a numeric template, and a matcher that decides whether that template belongs to an enrolled worker or to nobody on the list. The two chains fail for different reasons and have to be measured separately.
Camera height was the second issue. The plant's cameras were to be mounted high on the walls, which suits counting well because a high camera sees the whole yellow box with little overlap between people. The same camera mostly sees the tops of heads. Face recognition vendors ask for a camera that looks at the face from no more than about 15 degrees above eye level. A camera 5 metres up a wall, looking at someone a few metres away, sees them from 30 to 60 degrees above. No single camera position could do both jobs well.
Food plants add their own difficulties. Workers wear hairnets, often beard nets and masks, and the same uniform. NIST's 2020 study of masked faces found that even the best algorithms made errors on 2.4% to 5% of masked faces, several times their unmasked error rate. We could find no published study at all on hairnets and caps. Identical uniforms also ruled out a common shortcut: following a person from camera to camera by the colour and pattern of their clothes, which is how body re identification models mostly work.
Licensing narrowed the options further. The most popular free face recognition weights, including the InsightFace model packs, are released for non commercial research only. Ultralytics YOLO26, the newest YOLO release, is licensed under AGPL 3.0, and closed commercial use requires an Enterprise licence whose price Ultralytics does not publish. Most crowd detection datasets, CrowdHuman among them, also forbid commercial use. A system built for a paying client could use none of these without a licence.
Privacy was the last constraint and, for a UK headquartered group, a serious one. Face templates are biometric data. Pakistan's Personal Data Protection Bill was still a draft in September 2026, but the draft treats biometric data as sensitive and requires explicit consent. In the UK, the Information Commissioner's Office ordered Serco Leisure in 2024 to stop using facial recognition to record staff attendance. That case set the standard we expected Unilever's own reviewers to apply.
How we approached it
We began with research rather than a preferred architecture. In parallel, we looked at detection and tracking models, face recognition engines, cameras available through Pakistani and international suppliers, training datasets and their licences, GPU rental prices, on premises hardware and the state of Pakistani data protection law. Every model version, specification and price went into the notes with its source and the date we checked it, and every figure we could not confirm was marked as an estimate.
The first output that mattered was a definition of accuracy. We split the client's 92% to 98% into separate targets for door counts, yellow box counts, worker identification, wrong identifications and unknown flagging. Each target names where it applies and how it will be measured: a two week acceptance test in which two people count the same footage independently and the system is scored against their agreed result. This changed the conversation with the client. Door counting at 95% or better is well supported by published results for overhead counters. Identifying people inside the hall boxes, who face their work rather than any camera, is not something we could honestly promise at 92%, and the targets say so.
The second decision was to separate the two jobs physically. Counting cameras stay high, where they see the zones clearly. Face cameras go at head height, 2.4 metres up, at controlled points where people walk towards them: the approach aisle to each hall zone, and both directions at each door. A person leaving the building has their back to the camera that saw them come in, so every door needs two face cameras.
The solution
The finished design uses 15 cameras, all from Hanwha Vision's Wisenet A range: six ANV-L6082R domes for counting and nine ANO-L6082R bullets for faces. Both are 1080p cameras with 120 dB wide dynamic range, a 3.3 to 10.3 mm varifocal lens, H.265 video and NDAA compliance. In the halls, the counting dome sits 4 to 6 metres up, looking down at 40 to 60 degrees with the yellow box in the lower middle of the picture. At the doors it looks straight down from 3 to 4 metres at a zone just inside the entrance, where two virtual lines a metre apart record direction. Each face camera covers a walking lane from 3.0 to 4.8 metres in front of it, where a face fills at least 750 pixels per metre, and a small LED light keeps that lane bright on night shifts.

Figure 1. System architecture and technology stack. All processing runs on Unilever premises.
The cameras connect over wired PoE+ cable to switches on a separate camera network with no internet access. We kept video off the plant wifi on purpose. A busy wireless network drops video frames, a dropped frame can break a person's track in the middle of a doorway, and each camera needs power at its mount anyway. The wifi carries only what people use: the dashboard and the mobile app.
Everything else runs on one of Unilever's existing servers, fitted with an NVIDIA GPU. NVIDIA DeepStream 9.1 decodes all fifteen streams on the GPU. On the counting cameras, an RF-DETR Medium detector finds each person, ByteTrack gives each person an ID that holds from frame to frame, and Gezora.ai's counting logic records a count when a person stands inside a yellow box or crosses both door lines in order. We chose RF-DETR over YOLO26 because it scores 54.7 mAP on the COCO benchmark against 53.1 for YOLO26m, runs slightly faster on the same GPU, and carries the Apache 2.0 licence, which allows closed commercial use at no cost.
On the face cameras, Neurotechnology's SentiVeillance engine finds faces, picks the clearest frames of each person and builds a template that is matched against the enrolled workers stored in PostgreSQL with the pgvector extension. Neurotechnology was one of the few vendors we found that publishes its prices: EUR 790 for the development kit and EUR 200 per camera, as a perpetual licence. Faces that match nobody are grouped into temporary identities with HDBSCAN clustering, and a supervisor can tag a temporary identity as a visitor or an outsourced worker. Events travel over a NATS message bus to a FastAPI backend, which serves an Angular dashboard and a Flutter app for Android and iOS. Users sign in with their normal Unilever accounts through Keycloak.
Model training happens once. We record the cameras for a few days across all shifts, pick 3,000 to 5,000 frames, and label every person in them with CVAT running on the plant's server. Faces are blurred before any frame leaves the building. The person detector is then trained on a single NVIDIA A100 rented on RunPod at USD 1.59 an hour, for about 200 hours, which comes to roughly PKR 88,000. The finished model is installed on the plant's server and the rented GPU is released. The face engine needs no training. Workers are enrolled from the photos and videos Unilever supplies, and the match threshold is set on site.
The full stack is listed below. Every open source component carries a licence that allows closed commercial use.
| Layer | Technology |
|---|---|
| Cameras | Hanwha Wisenet A ANV-L6082R for counting, ANO-L6082R for faces |
| Network | Cat6 cable with PoE+ on a separate camera VLAN, RTSP video in H.265 |
| Server | Existing Unilever server with an NVIDIA GPU, Ubuntu Linux, Docker, CUDA 13.2 |
| Video processing | NVIDIA DeepStream 9.1 and TensorRT 10.16 |
| Detection and tracking | RF-DETR Medium and ByteTrack |
| Face recognition | Neurotechnology SentiVeillance, with HDBSCAN for grouping unknown people |
| Data and backend | NATS JetStream, Python with FastAPI, PostgreSQL 18 with pgvector 0.8 |
| Apps and sign in | Angular 22 dashboard, Flutter 3.47 mobile app, Keycloak 26.7 linked to Unilever sign in |
| Operations | Prometheus and Grafana for monitoring, MLflow for model versions, CVAT for labelling |
Bringing the cost down
The design did not arrive at that shape in one step. Our first full proposal priced the system at PKR 73.0 million for the first year. It included two Dell PowerEdge servers with GPUs in a failover pair (PKR 26.1 million on their own), premium Axis and Hanwha cameras at USD 800 to 840 each, a training workstation, storage, UPS units, a full development team, a year of support and a contingency sum. It was a sound design for a plant starting from nothing. This plant was not starting from nothing.

Figure 2. Gezora.ai proposal price by design round.
The second round started from what the plant already had. Unilever runs its own servers, databases, storage, server room and power backup, so we removed every one of those from the price and replaced the training workstation with rented GPU time. That brought the figure to PKR 27.0 million, most of it software development. In the third round the camera choice changed. The Hanwha Wisenet A models cost USD 180 to 240 each and keep the specifications the job needs, so the fifteen cameras fell from about PKR 4.6 million to PKR 1.2 million. The network switches and the GPU card moved below the line, to be added only if the plant cannot provide free switch ports or a GPU of its own. Software development was also taken out of this quotation to be priced separately, so not all of the drop in this round is a saving in the strict sense. The total came to PKR 5.8 million. The fourth round produced a second version of the quote for the case where the plant buys and installs the cameras through its own suppliers. That leaves Gezora.ai's price at PKR 1.9 million, covering the face recognition licence and the one time training.
None of these changes removed a function. The camera count, the two separate chains, the accuracy targets and the privacy controls are the same in the last version as in the first.
Privacy designed in from the start
We treated privacy as part of the architecture, not as a policy written after it. Counting is anonymous by design: the counting cameras produce track numbers and totals, never names. Only the face cameras identify anyone, and only people who have given written consent are enrolled. A worker who declines can use a badge or a supervisor sign in instead, with no penalty.
Retention limits are built into the database. Templates of unknown people are deleted after 24 hours unless a supervisor flags an incident, face snapshots attached to sightings are kept for 30 days, and a worker's template is deleted within 14 days of withdrawing consent. All data is encrypted, access is controlled by role, and every enrollment, search and export is written to an audit log. Signs at every camera area explain what is recorded and why. No video, face image or template leaves the plant at any point, including during training, when only face blurred frames go out. The plan requires Unilever to sign off a data protection impact assessment before the system goes live.
Delivery plan
The plan runs for 24 weeks. The first three weeks cover kick off, a survey of all six locations with test footage, and the consent forms and privacy assessment. Cameras are bought in weeks 3 to 9, while we set up the software on the plant's server. The pilot runs in weeks 9 to 13 at the first entry and exit point and the three hall zones, and the footage it records is labelled and used for the single training run between weeks 11 and 16. The remaining cameras go in from week 14 to week 18. Worker enrollment follows in weeks 16 to 19, and the two week acceptance test runs in weeks 19 to 22. The last two weeks are handover and training for the plant's administrators.
How success will be measured
The engagement is at the design stage, so its results are still to come. They will be judged against the targets agreed in the proposal, measured during a two week acceptance test against counts made independently by two people.
| What is measured | Where | Target |
|---|---|---|
| People entering and leaving | 3 entry and exit points | At least 95% per hour, at least 97% per day |
| People inside the yellow box | 3 hall zones | 92% to 95% |
| Enrolled workers correctly identified | Face cameras at the doors | At least 92%, aiming for 98% |
| Wrong names given to people not enrolled | All face cameras | Less than 1% |
| People inside the hall zones identified | 3 hall zones | Best effort, reported separately |
The pilot answers the questions no published study could, chief among them how much hairnets and caps affect recognition on this plant's own workers. It also produces the footage the person detector is trained on. The full rollout goes ahead only after the pilot results have been reviewed.
What we learned
The biggest saving came from a question we should have asked before pricing any hardware: what does the client already own? Our first design assumed a plant with no computing capacity. Once the existing servers, storage and network were counted in, more than PKR 26 million of hardware disappeared from the quote without any loss of function.
A single accuracy figure in a brief is an invitation to agree on what it means. Breaking 92% to 98% into separate targets for counting and identification, each tied to a location and a test method, gave the client a realistic picture and gave us commitments we can defend in an acceptance test.
Licences belong at the start of model selection, not at the end. Several of the strongest freely available models could not legally be shipped in a commercial system, and finding that out after building on them would have cost weeks.
For face recognition, camera placement matters more than the choice of model. No engine we reviewed can identify a worker seen from 40 degrees above with a hairnet pulled low. A camera at head height, in good light, facing people as they walk towards it, gives any competent engine a fair chance.
About Gezora.ai
Gezora.ai builds AI automation products and computer vision systems for businesses in Pakistan, including CRM, procurement and HR automation. For this engagement our team handled the research, solution architecture, camera layout, licensing review, privacy design, costing and delivery plan.
Working on a similar problem?
If your plant, warehouse or site needs people counted or identified on its own infrastructure, Gezora.ai can review your brief and camera layout with you. Visit gezora.ai.
Topics
- Computer vision case study
- Worker counting and identification
- Facial recognition
- On premises AI
- Manufacturing automation
- Unilever Pakistan Foods
Get started
Stop paying people to do what an agent can
Tell us what you want to automate. We will map the workflow, deploy the right agents, and train your team to run them.
- Every agent is trained on your own workflows, never a generic template
- Most deployments are live within two to four weeks
- SOC 2 compliant, with a complete audit trail on every deployment
