Automated sign language digitisation: Leveraging computer vision and archived broadcast footage for accessible avatar generation
Abstract
Sign language serves as the primary linguistic modality for the deaf and hard-of-hearing community. However, the vast majority of sign language content contained within historical video archives remains inaccessible — unindexed, unsearchable and effectively invisible to computational analysis. This paper proposes a novel framework for the automated detection, tracking and semantic annotation of signlanguage motions using advanced computer vision and archived broadcast footage. Building upon initial trials conducted by NHK Enterprises and utilising 2026-era hardware capabilities, the paper evaluates the shift from sensor-based motion capture to videobased pose estimation. Furthermore, it discusses the economic and ethical implications of creating large-scale 3D motion databases, proposing an international consortium model to govern the ‘physical artificial intelligence’ data marketplace. This article is also included in The Business & Management Collection which can be accessed at http://hstalks.com/business.
The full article is available to subscribers to the journal.
Author's Biography
Takashi Koyano has more than 20 years of experience in the digital domain at NHK. He led major digital initiatives for NHK’s coverage of the Tokyo Olympic and Paralympic Games, as well as the PyeongChang Games, with a particular focus on accessibility technologies such as automated sign-language generation. He is currently engaged in advancing NHK Enterprises’s automated sign-language computer graphics generation system, extending its applications beyond broadcasting to broader public information platforms, including transportation and aviation.