Generating Text Descriptions From Images / Video

Anant Raj and Rohun Tripati

{anantraj, rohunt}@iitk.ac.in

Abstract

In recent years a lot of work has been done in understanding the connection betweent the two modalities computer vision and natural language processing. Due to exponentially increasing web scale images and their available captions new opportunities have been evolved for integrative models bridging natural language processing and computer vision. Generating the text description of images or video has been an important part of this whole ongoing research. This project is motivated from the same research problem. The goal of this project is to genetate syntactically correct sentence [1] or sentences for a given image or video by generating syntactic trees which semantically describes the content of images or video.

colosseum_teaser

Proposal and Mid-Term Project Presentation

Proposal                                                        Presentation

Poster Presentation

Poster

Report

Final report

Code and Resources

Computer Vision Resources
Berkley Parser
Code

References

[1]Mitchell, Margaret, et al. "Midge: Generating image descriptions from computer vision detections." Proceedings of the 13th Conference of the European Chapter of the Association for Computational Linguistics. Association for Computational Linguistics, 2012.

[2] Ali Farhadi, Mohsen Hejrati, Mohammad Amin Sadeghi, Peter Young, Cyrus Rashtchian, Julia Hockenmaier, and David Forsyth. 2010. Every picture tells a story: generating sentences for images. Proceedings of ECCV 2010

[3]Siming Li, Girish Kulkarni, Tamara L. Berg, Alexander C. Berg, and Yejin Choi. 2011. Composing simple image descriptions using web-scale n-grams. Proceedings of CoNLL 2011.

[4] Krishnamoorthy, Niveda, et al. "Generating Natural-Language Video Descriptions Using Text-Mined Knowledge." NAACL HLT 2013 (2013): 10.

[5]Srivastava, Nitish, and Ruslan Salakhutdinov. "Multimodal learning with deep Boltzmann machines." Advances in Neural Information Processing Systems 25. 2012.

[6] Vicente Ordonez, Girish Kulkarni, and Tamara L Berg. 2011. Im2text: Describing images using 1 million captioned photographs. Proceedings of NIPS 2011.