← all repositories

peteanderson80/bottom-up-attention

A bottom-up attention model based on Faster R-CNN with ResNet-101 that extracts salient image region features for visual question answering and image captioning.

1.5k stars Jupyter Notebook Computer Vision
bottom-up-attention
Not currently ranked — collecting fresh signals.
star history

This repository provides code for training a bottom-up attention model using multi-GPU Faster R-CNN with ResNet-101 backbone, trained on Visual Genome object and attribute annotations. The pretrained model generates spatial features for salient image regions that can replace traditional CNN features in attention-based image captioning and VQA systems. The approach achieved state-of-the-art performance on MSCOCO captioning (CIDEr 117.9, BLEU_4 36.9) and won the 2017 VQA Challenge with 70.3% overall accuracy.

Frequently asked

What is peteanderson80/bottom-up-attention?
A bottom-up attention model based on Faster R-CNN with ResNet-101 that extracts salient image region features for visual question answering and image captioning.
Is bottom-up-attention open source?
Yes — peteanderson80/bottom-up-attention is open source, released under the MIT license.
What language is bottom-up-attention written in?
peteanderson80/bottom-up-attention is primarily written in Jupyter Notebook.
How popular is bottom-up-attention?
peteanderson80/bottom-up-attention has 1.5k stars on GitHub.
Where can I find bottom-up-attention?
peteanderson80/bottom-up-attention is on GitHub at https://github.com/peteanderson80/bottom-up-attention.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.