Generating Text Descriptions From Images / Video

Anant Raj and Rohun Tripati

{anantraj, rohunt}@iitk.ac.in

Abstract

In recent years a lot of work has been done in understanding the connection betweent the two modalities computer vision and natural language processing. Due to exponentially increasing web scale images and their available captions new opportunities have been evolved for integrative models bridging natural language processing and computer vision. Generating the text description of images or video has been an important part of this whole ongoing research. This project is motivated from the same research problem. The goal of this project is to genetate syntactically correct sentence [1] or sentences for a given image or video by generating syntactic trees which semantically describes the content of images or video.

colosseum_teaser

Review of The Literature

As mentioned earlier new ways for multimodal fusion of computer vision and natural language processing algorithm have been emerged in last couple of years. Various methods have been proposed to genetrate descriptive summaries of the given images [2] [3]. Very recently work has also been done in the field of textual description of video using text mined knowledge [4]. All of the work done has one common thing is that they all rely on computer vision detections to produce natural language descriptions. In image description state of the art object detection methods are used to detect objects(NP) and in video description state of the art activity detection methods with object detection methods are used to detect activity and objects in the video frame(VP and NP). Words are clustered on the basis of their semantics and heirarchy in Wordnet and these are used to generate descriptions. The difference in most of these literartures are the way the match the detected objects to words hierarchy. Also due to various recent advancements in the field of deep learning, multimodal deep learning methods have been evolved to generate caption for the image [5].

Methodology

Recent work by Mitchell et al [1] propose a new method for generating image description and implemented this method in his system "Midge" which is more robust than other state of the art method. The generated description by this system is in present tense. Both, the image and image description is described as triple. The problem is to map one triple to the another. Likelihood estimates of the syntactic structures and different object co-ocurring together gives a bit robustness to the system and deals with little computer vision detection errors. There still lack of evaluation methods for such type of problems.

Datasets

Flicker has enormous amount of captioned image data.
1. SBU Captioned Photo Dataset
2. Pascal datset(For Evaluation)
3. Youtube dataset
Example captions of images available in the dataset:
"A wooden chair in the living room
white horse near avebury
Yellow flower surrounded by scorched black stalks - Moore Nature Reserve
King Arthur's beheading rock - right on the sidewalk in the middle of town
This is a shot of the Brittanic flag flying atop a farmhouse beside a field of megaliths
It was taken when the season was running out and only this lonely flower was left in the field.
Coma sleeping on her bed in the front bedroom (2008)
Photos from a trip to a castle tower above Ljubljana.
Frost in my bathroom window as seen on a cold winter day.Cabri, Saskatchewan in February 2011
this was the sun and some tree branches in front of my bedroom window.. i inverted and edited."

Future Work

After generating the image description we are planning to extend this work to get video description.

References

[1]Mitchell, Margaret, et al. "Midge: Generating image descriptions from computer vision detections." Proceedings of the 13th Conference of the European Chapter of the Association for Computational Linguistics. Association for Computational Linguistics, 2012.

[2] Ali Farhadi, Mohsen Hejrati, Mohammad Amin Sadeghi, Peter Young, Cyrus Rashtchian, Julia Hockenmaier, and David Forsyth. 2010. Every picture tells a story: generating sentences for images. Proceedings of ECCV 2010

[3]Siming Li, Girish Kulkarni, Tamara L. Berg, Alexander C. Berg, and Yejin Choi. 2011. Composing simple image descriptions using web-scale n-grams. Proceedings of CoNLL 2011.

[4] Krishnamoorthy, Niveda, et al. "Generating Natural-Language Video Descriptions Using Text-Mined Knowledge." NAACL HLT 2013 (2013): 10.

[5]Srivastava, Nitish, and Ruslan Salakhutdinov. "Multimodal learning with deep Boltzmann machines." Advances in Neural Information Processing Systems 25. 2012.

[6] Vicente Ordonez, Girish Kulkarni, and Tamara L Berg. 2011. Im2text: Describing images using 1 million captioned photographs. Proceedings of NIPS 2011.