All projects
Study Project2026-03-25

Real-time Spotify Control with Hand Gestures

A six-gesture Spotify controller that reached 0.87623 mAP50-95 and runs directly from a webcam on macOS.

PythonYOLOv8OpenCVPyTorchAppleScriptHaGRID
Real-time Spotify Control with Hand Gestures

Overview

This project delivers a real-time hand-gesture system for controlling Spotify and system media on macOS. A YOLOv8n model recognizes seven classes: six gestures mapped to play, pause, next track, previous track, volume up, and volume down, plus a no_gesture class representing the absence of a command.

Problem

An accurate prediction on an individual frame is not enough for reliable control. A webcam can capture multiple objects, transitional gestures, or fluctuating predictions, any of which may trigger unintended media commands. The project therefore had to address both gesture detection and the safe conversion of predictions into actions.

Role

I served as team lead, built the primary model, and supported paper writing. My scope covered data preparation, YOLOv8 training and evaluation, and integration of the trained model into the macOS media-control demo.

Solution

  • Prepared 10,608 images in YOLO format across seven gesture classes.
  • Fine-tuned YOLOv8n for 100 epochs and retained the best checkpoint for inference.
  • Built an OpenCV webcam pipeline that selects the highest-confidence prediction in each frame.
  • Mapped six gestures to Spotify and macOS media commands through AppleScript.
  • Added multi-frame stability checks and a cooldown between actions.

Technical decisions

YOLOv8n was selected to balance model size with real-time inference needs. Instead of triggering every detection, the pipeline considers only the highest-confidence gesture, requires it to persist for four consecutive frames, and applies a default two-second cooldown. This trades a small amount of responsiveness for fewer unintended commands. The no_gesture class remains in the training set as a negative signal for moments when the user is not issuing a command.

Results

The best model, recorded at epoch 85, achieved 0.99096 precision, 0.98995 recall, 0.99344 mAP50, and 0.87623 mAP50-95 on the validation set. From the first epoch to the best epoch, mAP50-95 increased from 0.68412 to 0.87623, a relative improvement of approximately 28.08%. The checkpoint was integrated into a webcam demo supporting six media controls on macOS.

Lessons learned

  • A real-time vision product depends on both model quality and post-inference signal stabilization.
  • A no-command class and debounce logic help reduce unintended actions.
  • Tracking metrics by epoch makes it possible to select the best checkpoint rather than defaulting to the final one.

Keep exploring

Residual SAC for Retail Replenishment Under a 25 m² Limit