Many deep-learning computer vision solutions still fall short of the performance and stability needed to run in real time on low-end devices. I am travelling at the moment, and have used train journeys and breaks to improve my eye-tracking system for audiovisual performances. The initial version did its job, but with a few “quirks”.
For my current projects, I need a robust and efficient system. I explored AI implementations such as YOLO v3, v4, v8 and their “tiny” versions, as well as MobileNet-SSD v1 and v2. ChatGPT is useful here for quickly implementing and translating models between platforms; otherwise, it would have taken weeks to compare benchmarks, choose a model and implement it. I trained each model using data from a more accurate but slower version of my initial system, based on OpenCV’s k-means.
At this point, the results jittered and produced different data between almost identical frames. Reducing capabilities further to increase speed was not viable.
I returned to “classical” computer vision methods (soon we will call them retro), using good old reliable OpenCV. I used adaptive thresholding and median filters on a reduced image to find the region of interest (ROI), then detected contours and applied a least-squares circle fit on the original image. This achieved 40 fps on 1620×1080 images.
This solves a very specific case in a controlled context, but shows that traditional algorithms remain a very powerful alternative.
I have not ruled out improving efficiency further with AI algorithms. Along the way, ideas emerged that might work better by drawing on more complex image features in controlled environments.
One of the custom instruments.