The challenge
Explore recognition with a limited local vocabulary and limited data, while making uncertainty and the distinction between generic gestures and LSU explicit.
My contribution
- Built a desktop interface for webcam capture, landmarks and stable predictions.
- Added local data collection and classifier training workflows.
- Implemented phrase/gloss construction, copying and local text-to-speech.
Engineering decisions
- MediaPipe provides landmarks and a generic gesture demo. The training pipeline compares local classifiers with scikit-learn.
- Training and evaluation split samples by recording session to reduce leakage between nearly identical frames.
- Confidence thresholds and temporal stabilization reduce repeated or flickering predictions.
Outcome
An experiment with a limited vocabulary, not a complete sign-language translator. Generic MediaPipe gestures are not presented as validated Uruguayan Sign Language (LSU). Local LSU labels require expert/source review.