← All projects
Brno University of Technology

Sign Language Interpretation

Recognizing American Sign Language from hand and face pose and re-signing it in Czech Sign Language with a 3D avatar.

Rigged 3D hand model of the signing avatar

My master’s thesis at Brno University of Technology was about translating between sign languages. The system watches someone signing in American Sign Language, figures out which words they are signing, and a 3D avatar then signs the same words in Czech Sign Language.

Most research at the time worked with the video directly, often using 3D convolutional networks. I went another way and looked only at the pose of the hands and face. Hands are hands in every sign language, so everything that tracks them can stay the same, and only the final classifier needs to learn a new language. Leaving out the rest of the body turned out to be fine too, since the position of the hands relative to the face already says roughly what the arms are doing.

How it works

That choice meant building a chain of models, each handling one part of the problem. The first one finds hands and faces in every frame. I trained YOLOv5 and EfficientDet detectors for it and quickly ran into a problem: there was no public dataset with both hands and faces labeled. So I trained one detector on a hand dataset and another on a face dataset, let each of them fill in the missing labels in the other’s data, and trained the final detector on the merged result.

Next, a pose estimation network finds 98 keypoints on the face and 21 on each hand. Hands were the hard part. Signing is full of hands crossing, touching and hiding each other’s fingers, while most hand datasets show a single hand on its own. I ended up combining four datasets and giving the hand model a separate branch for interacting hands.

Single frames are noisy, so a tracker based on Deep SORT follows the hands and face through the video and filters out mistakes. Instead of adding yet another network to tell the left hand from the right, I matched them by how close their keypoints are from one frame to the next, which costs almost nothing in speed.

At the end of the chain, recurrent networks read the keypoint sequences and classify them into words. I trained them on the WLASL dataset with vocabularies from 100 up to 2000 signs, using the rest of the pipeline to annotate all of its videos automatically. The avatar is a 3D model animated in Blender and rendered in real time with Panda3D.

How it turned out

Recognition accuracy ended up close to the pose-based results published by the WLASL authors, and on the 1000-word vocabulary it was a few percent better in top-5 accuracy. Still, telling apart words from such a large vocabulary remained the weakest link, and it limits how well the whole system works in practice.

I presented the detection, pose estimation and tracking part at the EEICT 2022 student conference, and it won a sponsor award and the best paper award. The thesis itself also received the Dean’s award for an outstanding thesis.

Gallery