The BenAV dataset contains a lexicon of 50 words from 128 speakers (107 male and 21 female) with 26,300 utterances. The average number of speakers for each word is 18 (max 20, min 12, and standard deviation 1.826). The total duration of the dataset is 7.3 hours. This is the first Bengali audio-visual dataset that can be used for various research, including acoustic speech recognition and audio-visual speech recognition.
Date made available | 10 Mar 2021 |
---|
Publisher | Springer |
---|
Date of data production | 2021 - |
---|