—We present bilinear CNNs, an architecture that efficiently represents an image as a pooled outer product of two CNN features, that is effective at fine-grained recognition tasks. These models capture localized part-feature interactions similar to those in part-based models, but can also be seen as an orderless texture representation. Based on this(More)
