Dual-Dilated Large Kernel Convolution for Visual Attention Network
- 1 School of Communication, The Hang Seng University of Hong Kong, Hong Kong, China
- 2 School of Communication, The Hang Seng University of Hong Kong, Hong Kong, China
- 3 School of Communication, The Hang Seng University of Hong Kong, Hong Kong, China
Abstract
Visual Attention Networks (VANs) leveraging Large Kernel Attention (LKA) have demonstrated remarkable performance in diverse computer vision tasks, often outperforming Vision Transformers (ViTs) in some cases. LKA strategically combines the strengths of Convolutional Neural Networks (CNNs), such as local structure information, with the long-range dependency and adaptability of self-attention mechanisms, while maintaining linear computational complexity. This paper introduces Dual-Dilated Large Kernel (D2LK), a novel attention mechanism designed to enhance LKA’s kernel decomposition. D2LK improves upon LKA by incorporating an additional depth-wise dilation convolution layer, which enables the approximation of larger kernel convolutions with further reduced computational requirements. This decomposition allows for a more efficient representation of larger effective receptive fields. Our experiments demonstrate that D2LK achieves a superior balance between efficiency and performance. For instance, a D2LK module configured with a kernel size of 29 and 32 channels reduces parameters by 11% (3,008 parameters) compared to an LKA module with the same specifications (3,392 parameters). When integrated into the VAN-B0 architecture, D2LK with a larger kernel size of 29 yields a Top-1 accuracy of 85.1% on ImageNet100 classification, a slight improvement over the LKA baseline (kernel size 21), which achieved 85.0%. Critically, this performance gain is accomplished with a marginally reduced overall parameter count (3.8649 million for D2LK vs. 3.8745 million for LKA). These results validate D2LK as an efficient and effective attention mechanism for Visual Attention Networks, enabling enhanced receptive fields at lower computational overhead.
- Guo, M., Lu, C., Liu, Z., Cheng, M. and Hu, S. (2023) Visual Attention Network. Computational Visual Media , 9, 733-752. https://doi.org/10.1007/s41095-023-0364-2
- Lau, K., Po, L. and Rehman, Y. (2023) Large Separable Kernel Attention: Rethinking the Large Kernel Attention Design in CNN. https://doi.org/10.2139/ssrn.4463661
- Liu, S., Wei, J., Liu, G. and Zhou, B. (2023) Image Classification Model Based on Large Kernel Attention Mechanism and Relative Position Self-Attention Mechanism. PeerJ Computer Science , 9, e1344. https://doi.org/10.7717/peerj-cs.1344
- Wang, J., Wang, Y., Sun, A. and Zhang, Y. (2024) A Lightweight Network FLA-Detect for Steel Surface Defect Detection. https://doi.org/10.21203/rs.3.rs-4581669/v1
- Hu, J., Shen, L., Albanie, S., Sun, G. and Wu, E. (2019) Squeeze-and-Excitation Networks. https://doi.org/10.48550/arXiv.1709.01507
- Wang, Q., Wu, B., Zhu, P., Li, P., Zuo, W. and Hu, Q. (2020) ECA-Net: Efficient Channel Attention for Deep Convolutional Neural Networks. 2020 IEEE / CVF Conference on Computer Vision and Pattern Recognition ( CVPR ), Seattle, 13-19 June 2020, 11531-11539. https://doi.org/10.1109/cvpr42600.2020.01155
- Wang, X., Girshick, R., Gupta, A. and He, K. (2018) Non-Local Neural Networks. 2018 IEEE / CVF Conference on Computer Vision and Pattern Recognition , Salt Lake City, 18-23 June 2018, 7794-7803. https://doi.org/10.1109/cvpr.2018.00813
- Woo, S., Park, J., Lee, J. and Kweon, I.S. (2018) CBAM: Convolutional Block Attention Module. In: Ferrari, V., et al ., Eds., Computer Vision — ECCV 2018, Springer International Publishing, 3-19. https://doi.org/10.1007/978-3-030-01234-2_1
- Wang, L., Xiong, Y., Wang, Z., Qiao, Y., Lin, D., Tang, X., et al . (2016) Temporal Segment Networks: Towards Good Practices for Deep Action Recognition. In: Leibe, B., et al ., Eds., Computer Vision — ECCV 2016 , Springer International Publishing, 20-36. https://doi.org/10.1007/978-3-319-46484-8_2
- Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., et al . (2021) An Image Is Worth 16x16 Words: Transformers for Image Recognition at Scale. https://doi.org/10.48550/arXiv.2010.11929
- Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., et al . (2021) Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows. 2021 IEEE / CVF International Conference on Computer Vision ( ICCV ), Montreal, 10-17 October 2021, 9992-10002. https://doi.org/10.1109/iccv48922.2021.00986