
[Abstract]
We propose a delay-free audio source separation method that estimates the source signals of human voices and background sounds by learning background sound information from the input audio signal in real time. By applying this method to digital video devices, we aim to offer users new functions while watching video content, such as making human voices easier to hear, reducing background sound, enhancing the sense of presence in sports, and supporting singing practice. We confirmed that the method requires little processing, operates without delay on a range of devices, and lets users experience its effects immediately while watching. We also confirmed that learning voices and background sounds separately enables accurate estimation of both across a variety of content.
[Publications (Japanese) ]
- 広畑 誠, 小野 利幸, 西山 正志,
多様な映像コンテンツに対応した遅延なし音源分離技術,
日本音響学会 春季研究発表会, pp. 1-4, March, 2013. - 西山 正志, 広畑 誠, 小野 利幸,
声と背景音のボリュームバランス調整に向けた音源分離, 情処研報 CVIM187-46, pp. 1-5, May, 2013.