← Back to News
October 9, 2026

Require MAEpp and X Accuracy Before You Buy People Counting Systems

Verify people counting accuracy by demanding MAEpp, X Accuracy, and directional-bias reports. Validate on-site with at least 100 events per direction...

Require MAEpp and X Accuracy Before You Buy People Counting Systems

Require MAEpp and X Accuracy Before You Buy People Counting Systems

Technician observing people pass beneath a counting sensor

People counting accuracy is the measured agreement between an automated count and a manually verified ground truth, recorded under stated conditions like mounting height, entrance width, and traffic density. A single percentage means little on its own. Ask vendors for MAE per person (MAEpp), X-Accuracy, and separate in-count versus out-count bias, validated across both busy and quiet periods.


TL;DR:

  • Request MAE per person, X Accuracy within a stated tolerance, and separate entry and exit bias; a blended accuracy percentage can hide directional errors.
  • Require at least 100 manually verified events in each direction across busy and quiet periods, then document mounting angle, entrance width, and software versions.
  • A two camera overhead fisheye setup in a 2,000 square foot space reduced MAE per person from 0.373 to 0.097, with 92% within five people.
  • Beam and PIR counters suit controlled single file doorways; wide entrances or dense crowds need multiple sensors and duplicate removal when undercounts affect safety.

Table of Contents

What accuracy means in practice and the metrics professionals must use

Most vendor sheets quote a single accuracy percentage, and that number usually comes from exact-match testing: did the system's count equal the human-verified count, event by event. Exact-match accuracy is intuitive, but it penalizes a system that is off by one person in a crowd of 40 just as harshly as one that is off by one person in a doorway of two. That is not a useful way to judge performance at scale.

Mean Absolute Error (MAE) and Root Mean Square Error (RMSE) fix part of that problem by measuring the average size of the error rather than whether it hit exactly zero. RMSE penalizes large misses more heavily than small ones, which is useful when occasional big errors (a missed group, a duplicate count) matter more than small drift. But neither MAE nor RMSE tells you whether a given error is proportionally large or small relative to how many people actually passed through.

That gap is why the COSSY research on overhead fisheye counting proposes MAE per person (MAEpp), which scales the error to average true occupancy, and X-Accuracy, which reports the share of counting events that land within a stated tolerance of X people. Context decides which metric matters most.

Illustrated comparison of counting errors and tolerance

Directional bias deserves its own line item. An entrance sensor can overcount entries and undercount exits (or the reverse) and still produce a plausible-looking net total. Axis's own technical documentation shows how reported accuracy swings between roughly 99.46% and 100.55% depending on which direction you measure, which is why in-count and out-count errors should always be reported separately, rather than folded into one blended figure.

Sample size is the last trap. Practitioner guidance from Axis and similar manufacturers recommends validating against at least 100 events per direction, spread across both quiet and busy periods, before accepting any accuracy claim.

  • Exact-match accuracy tells you how often the count was perfectly right, which matters less as occupancy rises.
  • MAEpp scales error to average occupancy, making it comparable across sites of different sizes.
  • X-Accuracy reports how often counts fall within an operationally meaningful tolerance.
  • Directional bias (in-count vs out-count error) exposes asymmetries a blended number hides.

A multi-camera overhead fisheye system tested in a 2,000 square foot space cut MAEpp from 0.373 to 0.097 and raised X-Accuracy to 92% within five people and 98% within ten. That gap between a single-camera and two-camera setup is the clearest evidence that method and geometry, not just the sensor itself, drive real-world accuracy.

How different sensor methods typically perform and their trade-offs

Sensor choice sets the ceiling on achievable accuracy before installation decisions ever come into play. Each method handles density, lighting, and entrance geometry differently, and none dominates every scenario.

Camera-based systems using AI object detection tend to deliver the strongest localization and the highest accuracy in moderate-density settings, because they can track individual silhouettes rather than inferring presence from a beam break or a heat signature. Their weakness is occlusion: when people overlap or move in tight groups, detection models lose track of individuals, and accuracy degrades as density climbs. Multi-camera overhead fisheye configurations address this directly, with the COSSY study's two-camera setup nearly quadrupling accuracy over a single-camera baseline by covering blind spots and resolving overlapping bodies from different angles.

Depth sensing and time-of-flight (ToF) or stereo cameras handle glare and shadow far better than standard RGB cameras, since they measure distance rather than relying on contrast or color. Their limitation is range: most depth sensors lose reliable resolution beyond a few meters, which makes them a poor fit for wide entrances or atriums where a single unit cannot cover the full width.

Thermal and infrared sensors, along with radar, are resilient in low light and inherently privacy-friendly since they capture heat or reflected radio signatures rather than identifiable images. Research evaluating supervision levels for infrared-based counting found that weaker, image-level annotation can approach detector-level performance on some datasets, offering a real cost-versus-accuracy trade-off for teams with limited labeling budgets. The trade-off is spatial resolution: thermal and radar systems struggle to separate closely packed individuals in dense crowds, which caps their usefulness for exact counts where every person matters.

Beam and PIR tripwires, along with turnstiles, remain the simplest and often most dependable option for a single, controlled entry point. They are also the most brittle: any grouped or parallel passage, two people walking shoulder to shoulder through a doorway built for one, produces an undercount that the sensor has no way to correct.

  • Camera plus AI: high accuracy in moderate density, weak against occlusion unless multiple cameras cover the space.
  • Depth and ToF: strong against glare and shadow, limited by effective range on wide entrances.
  • Thermal, IR, and radar: resilient in low light and privacy-friendly, lower resolution in dense crowds.
  • Beam, PIR, and turnstiles: reliable for single-file entries, fail when people pass through side by side.

Published accuracy figures across these methods rarely use the same metric or test conditions, so a "97% accurate" thermal sensor and a "97% accurate" camera system are not necessarily comparable unless both report MAEpp and X-Accuracy under similar density and geometry.

Key real-world factors that drive errors in the field

Accuracy claims made in a lab rarely survive contact with a real entrance. Several recurring stressors explain most of the gap between a spec sheet and field performance.

  1. Occlusion and group movement: when people cluster or walk in tight formation, sensors lose the ability to distinguish individuals, which typically produces undercounts rather than overcounts.
  2. Wide entrances with overlapping camera coverage: multiple units covering the same zone without proper calibration can double-count the same person as they cross from one field of view into another.
  3. Lighting variation, glare, and shadow: sudden changes in brightness, reflections off glass or flooring, and long shadows can trigger spurious detections in camera-based systems, inflating counts.
  4. Mounting height and angle: a sensor mounted too low or at too shallow an angle sees people at an oblique perspective, which increases the chance of missed or merged detections.
  5. Crowd density at the extremes: as occupancy rises, exact-match accuracy collapses even in well-engineered systems, which is precisely why MAEpp and X-Accuracy give a more honest operational picture than a single percentage.

Axis's own guidance on 3D people counting hardware stresses that geometry is not a cosmetic detail: overhead or 3D sensing reduces the errors shadows and glare introduce, and wide entrances generally require multiple devices with explicit duplicate-removal logic rather than a single wide-angle unit stretched past its effective coverage. Skipping that step is one of the most common reasons a system that tested well in a pilot underperforms once installed at scale.

How to validate and test people-counting accuracy at your site

A credible accuracy number comes from a validation plan you can reproduce, not a vendor's marketing sheet. The process starts with ground truth: a manually verified count against which the automated system is compared.

  • Manual tally: a trained observer counts in real time, which is cheap but prone to fatigue errors during long or busy shifts.
  • Synchronized video review: reviewers count from recorded footage after the fact, trading speed for higher accuracy on the ground-truth side itself.
  • Turnstile or access-control logs: useful where a physical barrier already exists, though they cannot capture tailgating or grouped entries.

Sample size matters as much as method. A short demo run of a few dozen events can look perfect by chance and prove nothing about real performance. Practitioner guidance recommends validating with at least 100 events per direction, spread across both a busy period and a quiet one, since accuracy often shifts meaningfully between the two.

The results worth reporting are MAE, MAEpp, X-Accuracy at a tolerance that matches your operational needs, and directional bias for in-counts versus out-counts separately. Document the test conditions alongside the numbers: mounting height and angle, entrance width, detection algorithm version, and firmware release, since any of these changing later can shift accuracy without anyone noticing.

Translating those numbers into an acceptance threshold is the final step. A retail analytics deployment tracking general footfall trends can tolerate a wider MAEpp than a security application counting occupancy against a fire-code limit, where even small undercounts carry real consequences.

Pro Tip: Ask vendors to run their validation test live at your actual site, during your actual peak hour, not at a reference installation chosen to flatter the numbers.

Design and installation choices that materially improve accuracy

Sensor selection sets a ceiling on accuracy, but installation choices decide how close to that ceiling a deployment actually gets. Several design decisions consistently move the needle.

Multi-camera setups with duplicate-removal algorithms are the clearest lever available for wide or irregular entrances. The COSSY study's two-camera overhead fisheye configuration demonstrated that adding a second overlapping camera, paired with logic to resolve the same person seen twice, nearly quadrupled MAEpp performance over a single unit. Overhead or near-overhead mounting generally outperforms oblique, eye-level angles because it minimizes the occlusion that comes from people walking in front of one another.

Edge processing and sensor fusion, combining depth, visual, and sometimes radar data at the device itself, reduce the false positives that glare and shadow otherwise introduce, while also keeping raw identifiable footage off the broader network. Routine calibration and periodic re-validation matter just as much as the initial install: a system validated once at deployment and never checked again will drift as furniture, signage, or seasonal foot traffic patterns change the scene geometry it was tuned for. Firmware and algorithm updates should be tracked explicitly, since a silent update can shift accuracy in either direction without a corresponding change in the reported numbers.

Operational design choices round out the picture. Favoring anonymized counting outputs over stored identifiable video reduces both privacy exposure and storage overhead, and it is worth deciding deliberately how long any retained data needs to live rather than defaulting to indefinite storage. Latency matters too: a system feeding a live occupancy dashboard needs near-real-time output, while one feeding a weekly footfall report can tolerate batch processing with no accuracy penalty.

  • Multi-camera fusion with duplicate-removal logic is the strongest lever for wide entrances.
  • Overhead or near-overhead mounting reduces occlusion compared with oblique angles.
  • Edge processing cuts false positives from glare and shadow while limiting raw video exposure.
  • Scheduled re-validation catches drift that a one-time install check will miss.

Field-proven proof points behind accurate people counting

Our engineering work on sensor-based security systems gives us a direct view into which of these principles hold up outside a lab. A two-week pilot for loitering detection analytics showed how a short, structured validation window can surface accuracy issues before a full rollout. In a queue-detection deployment for a Singapore procurement application, our analytics reached 93% accuracy once edge fusion was applied to combine multiple sensor inputs rather than relying on a single feed.

That same edge-fusion approach, combining depth, visual, and radar data at the device level, is what reduces the duplicate counts and false positives that plague single-sensor setups in crowded or glare-prone spaces. A related deployment pushed license plate recognition accuracy from 67% to 99.4% after the same fusion and validation discipline was applied, which is the kind of gap a short pilot is built to catch.

  • Edge fusion combining depth, visual, and radar inputs reduces duplicate counts from overlapping coverage.
  • Short structured pilots (one to two weeks) surface accuracy gaps before full deployment commitment.
  • Sub-10 ms tracking latency supports real-time occupancy dashboards without batch delay.

These are our own field results, offered as illustrations of the principles above rather than a substitute for independent validation. Any pilot should still be measured against the MAEpp, X-Accuracy, and sample-size standards covered earlier, on your own site, under your own conditions.

A procurement checklist worth running before you sign anything

Every accuracy claim I have reviewed gets more honest once you ask for MAEpp, X-Accuracy, and directional bias instead of a single headline number, validated over at least 100 events per direction across both busy and quiet periods. Put that requirement in the RFP itself, not as a follow-up question after the contract is signed.

Privacy-by-design deserves equal weight in the evaluation. PDPC guidance in Singapore treats identifiable CCTV footage as personal data, so anonymized counting outputs are the safer default unless your use case genuinely requires identifiable video for security purposes.

Lower-cost beam or PIR sensors are fine for a single controlled doorway with predictable, single-file traffic. Reserve multi-sensor, fusion-based systems for wide entrances, dense crowds, or any application where an undercount carries real operational or safety consequences.

— Eumir

How we support pilots, validation, and integration

Getting from a vendor's spec sheet to a number you can trust at your own site usually means running a real pilot, not reading another brochure — learn more about how package sales GMV now shown in exports can integrate into your operational analytics. Our Solution Integration team works directly with system integrators and facility operators to plan a site survey, set up an acceptance test built around MAEpp and X-Accuracy rather than a single blended percentage, and run a short pilot before any full commitment.

Beyondsensor

For teams already working with hardware vendors or regional distributors, Ecosystem Matchmaking helps connect the right sensor configuration, camera, depth, thermal, or fused, to the entrance geometry and density profile that actually matches your site. If you are evaluating how a counting deployment fits into a broader security or operations rollout, our Solution Integration page is the place to start a site survey and scope a validation run on your own terms.

FAQ

What are some accurate sensors for counting people?

Camera-based AI detection, depth or time-of-flight sensors, thermal and infrared units, radar, and beam or PIR tripwires are the main sensor categories used for people counting. Each handles density, lighting, and entrance width differently, so accuracy depends more on matching the method to the site than on any single sensor type being universally best.

How much does a people counter sensor cost?

Pricing varies widely by sensor type, coverage area, and whether the deployment needs multi-camera fusion or a single unit, and no published figure applies across the market. The most reliable way to get a real number is a site survey that scopes your entrance geometry and density needs before quoting hardware.

What is a people counting system?

A people counting system combines one or more sensors (camera, depth, thermal, radar, or beam-based) with processing logic that converts raw detections into directional in-count and out-count figures. Accuracy depends on the sensor method, installation geometry, and how counts are validated against a manually verified ground truth.

What is a people counter called?

People counters are also referred to as footfall counters, occupancy sensors, or visitor counting systems depending on the industry. The underlying technology and accuracy considerations, metrics like MAE per person and X-Accuracy, remain the same regardless of the label used.

Sources

Recommended

Share this article:
Get In Touch

Let's Build YourSecurity Ecosystem.

Whether you're a System Integrator, Solution Provider, or an End-User looking for trusted advisory, our team is ready to help you navigate the BeyondSensor landscape.

Direct Advisory

Connect with our regional experts for tailored solutioning.