Object Tracking: from 4.7 to 64 FPS on the same CPU
A real-time multi-object tracker in Python. The first version ran at about 4.7 FPS on my laptop. The current one runs at about 64 FPS on the same CPU, with no GPU.
The problem
The textbook stack is YOLOv8 for detection and DeepSORT for identity. It is accurate, and on a CPU it is slow: about 210 ms of PyTorch inference and about 30 ms of re-identification per frame. At roughly 240 ms a frame, 4.7 FPS is what you get.
My approach
Measure where the time goes, then ask what each expensive part is buying. I was tracking cars and people in ordinary video. Objects do not teleport between frames, so their boxes overlap heavily from one frame to the next. For that use case, box overlap is enough to keep identities.
Architecture
capture thread (queue of 2, drop old frames)
|
v
frame skip (every Nth frame)
|
v
ONNX detector at 320x320 -> NMS
| boxes
v
IoU tracker (greedy match, highest overlap first)
|
v
draw boxes, IDs and FPS
The same pipeline sits behind three interfaces: a PySide6 desktop app, a command-line tracker, and a FastAPI server that streams MJPEG to a web dashboard with JWT login.
What I built
| Version | Detector | Tracker | FPS |
|---|---|---|---|
| v1 | YOLOv8 (PyTorch) | DeepSORT | 4.7 |
| v2 | YOLOv8 (ONNX Runtime) | DeepSORT | 12 |
| v3 | YOLOv8 (ONNX Runtime) | IoU tracker | 64 |
Engineering decisions
- ONNX Runtime instead of PyTorch. Same model, same weights. Detection went from 210 ms at 640 pixels to 36 ms at 320.
- A 40-line IoU tracker instead of DeepSORT. Tracker time went from about 30 ms to about 0.01 ms, with no model to load.
- Frame skipping. Detection runs every third frame and the tracker coasts on the last boxes in between.
- Threaded capture. Frames are read in the background into a small queue that drops old frames, so camera I/O never blocks inference.
What failed
The first version, which is the one every tutorial builds. It worked and it crawled. The mistake was paying for a CNN re-identification model and a Kalman filter to solve a problem this footage did not have.
Current limitations
The IoU tracker has no memory of what an object looks like. If two objects fully overlap and swap positions, or one is hidden for longer than the track's maximum age, it can swap or drop identities where DeepSORT's appearance model would survive. DeepSORT is still in the repo as an option for heavy occlusion.
Evidence
Measured on a CPU-only laptop:
| Metric | v1 | v3 |
|---|---|---|
| FPS | 4.7 | 64.1 |
| Detection latency | 210 ms | 36 ms |
| Tracker latency | about 30 ms | 0.01 ms |
| Runtime dependencies | 6 | 3 |
| Install size | about 2.1 GB | about 85 MB |