Im2Text: Describing Images Using 1 Million Captioned Photographs

Vicente Ordonez; Girish Kulkarni; Tamara L. Berg

2011 NIPS NeurIPS 2011

Im2Text: Describing Images Using 1 Million Captioned Photographs

Abstract

We develop and demonstrate automatic image description methods using a large captioned photo collection. One contribution is our technique for the automatic collection of this new dataset -- performing a huge number of Flickr queries and then filtering the noisy results down to 1 million images with associated visually relevant captions. Such a collection allows us to approach the extremely challenging problem of description generation using relatively simple non-parametric methods and produces surprisingly effective results. We also develop methods incorporating many state of the art, but fairly noisy, estimates of image content to produce even more pleasing results. Finally we introduce a new objective performance measure for image captioning.

🌉 Interdisciplinary Bridge — Computer Vision and Machine Learning

🧭 Keyword Pioneer — image captioning

🐝 Cross-Pollinator — Artificial Intelligence, Computer Vision, Data Science & Analytics, Deep Learning, Interdisciplinary, Machine Learning, Mathematics & Optimization, Natural Language Processing, Reinforcement Learning

🌱 Topic Pioneer — Text Generation

📈 Trend Setter — Image Captioning

🐣 Hot Topic Early Bird — multimodal learning

Authors

Vicente Ordonez , Girish Kulkarni , Tamara L. Berg

Topics

Machine Learning > Application Areas > Domain Adaptation Computer Vision > Generation > Image Captioning Natural Language Processing > Generation > Text Generation Deep Learning > Learning Types > Multi-Modal Learning

Keywords

multimodal learning image captioning visual description dataset collection non-parametric method image description visual content estimation

Download PDF

Related papers

Co-Training for Domain Adaptation 2011

The Local Rademacher Complexity of Lp-Norm Multiple Kernel Learning 2011

Learning to Agglomerate Superpixel Hierarchies 2011

A Reinforcement Learning Theory for Homeostatic Regulation 2011

A Global Structural EM Algorithm for a Model of Cancer Progression 2011