Abstract
Humans learn abstract representations of meaning from perception of audio-visual stimuli. Our work is focused on perception of audio-visual information in videos by the computer as a step towards learning se- mantic understanding and representation. We give and approach for the computer to recognize the different speakers (actors) in a video by distin- guishing their voice and tagging the speaker’s face for each different voice. We first segment speech and non-speech in the audio stream. From the speech, we next go for diarization, i.e. segmentation with respect to dif- ferent speakers. For each speaker, we try to assign a face from the visual stream.