<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.3 20210610//EN" "JATS-journalpublishing1-3.dtd">
<article article-type="research-article" dtd-version="1.3" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xml:lang="ru"><front><journal-meta><journal-id journal-id-type="publisher-id">radioelectronics</journal-id><journal-title-group><journal-title xml:lang="ru">Известия высших учебных заведений России. Радиоэлектроника</journal-title><trans-title-group xml:lang="en"><trans-title>Journal of the Russian Universities. Radioelectronics</trans-title></trans-title-group></journal-title-group><issn pub-type="ppub">1993-8985</issn><issn pub-type="epub">2658-4794</issn><publisher><publisher-name>Saint Petersburg Electrotechnical University</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.32603/1993-8985-2020-23-6-6-16</article-id><article-id custom-type="elpub" pub-id-type="custom">radioelectronics-474</article-id><article-categories><subj-group subj-group-type="heading"><subject>Research Article</subject></subj-group><subj-group subj-group-type="section-heading" xml:lang="ru"><subject>ТЕЛЕВИДЕНИЕ И ОБРАБОТКА ИЗОБРАЖЕНИЙ</subject></subj-group><subj-group subj-group-type="section-heading" xml:lang="en"><subject>TELEVISION AND IMAGE PROCESSING</subject></subj-group></article-categories><title-group><article-title>Метод автоматического определения ключевых точек объекта на изображении</article-title><trans-title-group xml:lang="en"><trans-title>An Automatic Method for Interest Point Detection</trans-title></trans-title-group></title-group><contrib-group><contrib contrib-type="author" corresp="yes"><name-alternatives><name name-style="eastern" xml:lang="ru"><surname>Зубов</surname><given-names>И. Г.</given-names></name><name name-style="western" xml:lang="en"><surname>Zubov</surname><given-names>I. G.</given-names></name></name-alternatives><bio xml:lang="ru"><p>Зубов Илья Геннадьевич – магистр техники и технологий (2016), программист-алгоритмист компании ООО "НЕКСТ". Автор шести научных публикаций. Сфера научных интересов – цифровая обработка изображений; прикладные телевизионные системы.</p><p>ул. Рочдельская, д. 15, стр. 13, Москва, 123022</p></bio><bio xml:lang="en"/><email xlink:type="simple">ZubovIG@gmail.com</email><xref ref-type="aff" rid="aff-1"/></contrib></contrib-group><aff-alternatives id="aff-1"><aff xml:lang="ru"><institution>ООО "НЕКСТ"</institution><country>Россия</country></aff><aff xml:lang="en"><institution>Ltd "Next"</institution><country>Russian Federation</country></aff></aff-alternatives><pub-date pub-type="collection"><year>2020</year></pub-date><pub-date pub-type="epub"><day>28</day><month>12</month><year>2020</year></pub-date><volume>23</volume><issue>6</issue><fpage>6</fpage><lpage>16</lpage><permissions><copyright-statement>Copyright &amp;#x00A9; Зубов И.Г., 2020</copyright-statement><copyright-year>2020</copyright-year><copyright-holder xml:lang="ru">Зубов И.Г.</copyright-holder><copyright-holder xml:lang="en">Zubov I.G.</copyright-holder><license xml:lang="ru" license-type="creative-commons-attribution" xlink:href="https://creativecommons.org/licenses/by/4.0/" xlink:type="simple"><license-p>Данная работа распространяется под лицензией Creative Commons Attribution 4.0.</license-p></license><license xml:lang="en" license-type="creative-commons-attribution" xlink:href="https://creativecommons.org/licenses/by/4.0/" xlink:type="simple"><license-p>This work is licensed under a Creative Commons Attribution 4.0 License.</license-p></license></permissions><self-uri xlink:href="https://re.eltech.ru/jour/article/view/474">https://re.eltech.ru/jour/article/view/474</self-uri><abstract><sec><title>Введение</title><p>Введение. Внедрение систем технического зрения в повседневную жизнь становится все более популярным. Использование систем на основе монокулярной камеры позволяет решить большой спектр задач. Анализ монокулярных изображений является наиболее развивающимся направлением в области машинного зрения. Это обусловлено общедоступностью цифровых камер и больших наборов аннотированных данных, а также мощностью современной вычислительной техники. Для того чтобы система компьютерного зрения описывала объекты и предсказывала их действия в физическом пространстве сцены, необходимо интерпретировать анализируемое изображение с точки зрения базовой 3D-сцены. Этого можно достичь, анализируя жесткий объект как совокупность взаимно связанных частей, что представляет мощный контекст и структуру для рассуждений о физическом взаимодействии.</p></sec><sec><title>Цель работы</title><p>Цель работы. Разработка автоматического метода ключевых точек объекта интереса на изображении.</p></sec><sec><title>Методы и материалы</title><p>Методы и материалы. Предложен автоматический метод локализации ключевых точек транспортных средств на изображении, в частности номерного знака. Представленный метод позволяет зафиксировать ключевые точки объекта интереса на основе анализа сигналов внутренних слоев сверточных нейронных сетей, обученных для классификации изображений, и выделения объектов на изображении. Также метод позволяет детектировать части объекта без больших затрат на аннотацию данных и обучение.</p></sec><sec><title>Результаты</title><p>Результаты. Эксперименты подтвердили корректность выделения ключевой точки объекта интереса на основе предложенного метода. Точность выделения ключевой точки на номерном знаке составила 97 %.</p></sec><sec><title>Заключение</title><p>Заключение. Представлен новый метод выделения ключевых точек объекта интереса на основе анализа сигналов внутренних слоев сверточных нейронных сетей. Метод обладает точностью выделения ключевых точек объекта интереса на уровне современных методов, а в отдельных случаях превосходит их.</p></sec></abstract><trans-abstract xml:lang="en"><sec><title>Introduction</title><p>Introduction. Computer vision systems are finding widespread application in various life domains. Monocularcamera based systems can be used to solve a wide range of problems. The availability of digital cameras and large sets of annotated data, as well as the power of modern computing technologies, render monocular image analysis a dynamically developing direction in the field of machine vision. In order for any computer vision system to describe objects and predict their actions in the physical space of a scene, the image under analysis should be interpreted from the standpoint of the basic 3D scene. This can be achieved by analysing a rigid object as a set of mutually arranged parts, which represents a powerful framework for reasoning about physical interaction.</p></sec><sec><title>Objective</title><p>Objective. Development of an automatic method for detecting interest points of an object in an image.</p></sec><sec><title>Materials and methods</title><p>Materials and methods. An automatic method for identifying interest points of vehicles, such as license plates, in an image is proposed. This method allows localization of interest points by analysing the inner layers of convolutional neural networks trained for the classification of images and detection of objects in an image. The proposed method allows identification of interest points without incurring additional costs of data annotation and training.</p></sec><sec><title>Results</title><p>Results. The conducted experiments confirmed the correctness of the proposed method in identifying interest points. Thus, the accuracy of identifying a point on a license plate achieved 97%.</p></sec><sec><title>Conclusion</title><p>Conclusion. A new method for detecting interest points of an object by analysing the inner layers of convolutional neural networks is proposed. This method provides an accuracy similar to or exceeding that of other modern methods.</p></sec></trans-abstract><kwd-group xml:lang="ru"><kwd>сверточные нейронные сети</kwd><kwd>анализ карт активаций</kwd><kwd>выделение ключевых точек</kwd></kwd-group><kwd-group xml:lang="en"><kwd>convolutional neural networks</kwd><kwd>activation map analysis</kwd><kwd>interest point detection</kwd></kwd-group></article-meta></front><back><ref-list><title>References</title><ref id="cit1"><label>1</label><citation-alternatives><mixed-citation xml:lang="ru">Zeiler M. D., Fergus R. Visualizing and understanding convolutional networks // Proc. of the 13th Europ. Conf. on Computer Vision, Zurich, Switzerland, 6–12 Sept. 2014. Berlin: Springer, 2014. P. 818–833. doi: 10.1007/978-3-319-10590-1_53</mixed-citation><mixed-citation xml:lang="en">Zeiler M. D., Fergus R. Visualizing and understanding convolutional networks. Proc. of the 13th Europ. Conf. on Computer Vision, Zurich, Switzerland, 6–12 Sept. 2014. Berlin: Springer, 2014, pp. 818–833. doi: 10.1007/978-3-319-10590-1_53</mixed-citation></citation-alternatives></ref><ref id="cit2"><label>2</label><citation-alternatives><mixed-citation xml:lang="ru">Simonyan K., Vedaldi A., Zisserman A. Deep inside convolutional networks: Visualising image classification models and saliency maps // Proc. of the ICLR Intern. Conf. on Learning Representations, Banff, Canada, Apr. 2014. URL: https://arxiv.org/abs/1312.6034 (дата обращения 15.11.2020)</mixed-citation><mixed-citation xml:lang="en">Simonyan K., Vedaldi A., Zisserman A. Deep inside convolutional networks: Visualising image classification models and saliency maps. Proc. of the ICLR Intern. Conf. on Learning Representations, Banff, Canada, apr. 2014. Available at: https://arxiv.org/abs/1312.6034 (accessed 15.11.2020)</mixed-citation></citation-alternatives></ref><ref id="cit3"><label>3</label><citation-alternatives><mixed-citation xml:lang="ru">Simon M., Rodner E., Denzler J. Part Detector Discovery in Deep Convolutional Neural Networks // Proc. of the ACCV Asian Conf. on Computer Vision, Singapore, 1–5 Nov. 2014. Berlin: Springer, 2014. Pt. 2. P. 162–177. doi: 10.1007/978-3-319-16808-1_12</mixed-citation><mixed-citation xml:lang="en">Simon M., Rodner E., Denzler J. Part Detector Discovery in Deep Convolutional Neural Networks. Proc. of the ACCV Asian Conf. on Computer Vision, Singapore, 1–5 Nov. 2014. Berlin: Springer, 2014. Pt. 2, pp. 162–177. doi: 10.1007/978-3-319-16808-1_12</mixed-citation></citation-alternatives></ref><ref id="cit4"><label>4</label><citation-alternatives><mixed-citation xml:lang="ru">Object Detectors Emerge in Deep Scene CNNs / B. Zhou, A. Khosla, A. Lapedriza, A. Oliva, A. Torralba // Intern. Conf. on Learning Representations, San Diego, USA, May 2015. URL: http://hdl.handle.net/1721.1/96942 (дата обращения 15.11.2020)</mixed-citation><mixed-citation xml:lang="en">Zhou B., Khosla A., Lapedriza A., Oliva A., Torralba A. Object Detectors Emerge in Deep Scene CNNs. Intern. Conf. on Learning Representations, San Diego, USA, May 2015. Available at: http://hdl.handle.net/1721.1/96942 (accessed 15.11.2020)</mixed-citation></citation-alternatives></ref><ref id="cit5"><label>5</label><citation-alternatives><mixed-citation xml:lang="ru">MMDetection: open MMLab Detection Toolbox and Benchmark / K. Chen, J. Wang, J. Pang, Yu. Cao, Yu Xiong, X. Li, Sh. Sun, W. Feng, Z. Liu, J. Xu, Zh. Zhang, D. Cheng, Ch. Zhu, T. Cheng, Q. Zhao, B. Li, X. Lu, R. Zhu, Y. Wu, J. Dai, J. Wang, J. Shi, W. Ouyang, Ch. Change Loy, D. Lin. 13 p. URL: https://arxiv.org/pdf/1906.07155.pdf (дата обращения 02.06.2020)</mixed-citation><mixed-citation xml:lang="en">Chen K., Wang J., Pang J., Cao Yu., Xiong Yu, Li X., Sun Sh., Feng W., Liu Z., Xu J., Zhang Zh., Cheng D., Zhu Ch., Cheng T., Zhao Q., Li B., Lu X., Zhu R., Wu Y., Dai J., Wang J., Shi J., Ouyang W., Change Loy Ch., Lin D. MMDetection: open MMLab Detection Toolbox and Benchmark, 13 p. Available at: https://arxiv.org/pdf/1906.07155.pdf (accessed 02.06.2020)</mixed-citation></citation-alternatives></ref><ref id="cit6"><label>6</label><citation-alternatives><mixed-citation xml:lang="ru">Deep Residual Learning for Image Recognition / K. He, X. Zhang, S. Ren, J. Sun // Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition, Las Vegas, USA, 27–30 June 2016. Piscataway: IEEE, 2016. Art. 16541111. doi: 10.1109/CVPR.2016.90</mixed-citation><mixed-citation xml:lang="en">He K., Zhang X., Ren S., Sun J. Deep Residual Learning for Image Recognition. Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition, Las Vegas, USA, 27–30 June 2016. Piscataway: IEEE, 2016, art. 16541111. doi: 10.1109/CVPR.2016.90</mixed-citation></citation-alternatives></ref><ref id="cit7"><label>7</label><citation-alternatives><mixed-citation xml:lang="ru">Simonyan K., Zisserman A. Very Deep Convolutional Networks for large-Scale Image Recognition. Apr 2015. 14 p. URL: https://arxiv.org/abs/1409.1556 (дата обращения 02.06.2020)</mixed-citation><mixed-citation xml:lang="en">Simonyan K., Zisserman A. Very Deep Convolutional Networks for large-Scale Image Recognition. Apr. 2015, 14 p. Available at: https://arxiv.org/abs/1409.1556 (accessed 02.06.2020)</mixed-citation></citation-alternatives></ref><ref id="cit8"><label>8</label><citation-alternatives><mixed-citation xml:lang="ru">MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications / A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, H. Adam. URL: https://arxiv.org/abs/1704.04861 (дата обращения 02.06.2020)</mixed-citation><mixed-citation xml:lang="en">Howard A. G., Zhu M., Chen B., Kalenichenko D., Wang W., Weyand T., Andreetto M., Adam H. MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications. Available at: https://arxiv.org/abs/1704.04861 (accessed 02.06.2020)</mixed-citation></citation-alternatives></ref><ref id="cit9"><label>9</label><citation-alternatives><mixed-citation xml:lang="ru">SqueezeNet: AlexNet-Level Accuracy with 50x Fewer Parameters and &lt;0.5MB Model Size / F. N. Iandola, S. Han, M. W. Moskewicz, K. Ashraf, W. J. Dally, K. Keutzer. URL: https://arxiv.org /abs/1602.07360 (дата обращения 02.06.2020)</mixed-citation><mixed-citation xml:lang="en">Iandola F. N., Han S., Moskewicz M. W., Ashraf K., Dally W. J., Keutzer K. SqueezeNet: AlexNet-Level Accuracy with 50x Few-er Parameters and &lt;0.5MB Model Size. Available at: https://arxiv.org/abs/1602.07360 (accessed 02.06.2020)</mixed-citation></citation-alternatives></ref><ref id="cit10"><label>10</label><citation-alternatives><mixed-citation xml:lang="ru">Redmon J., Farhadi A. YOLOv3: An Incremental Improvement. URL: https://arxiv.org/abs/1804.02767 (дата обращения 02.06.2020)</mixed-citation><mixed-citation xml:lang="en">Redmon J., Farhadi A. YOLOv3: An Incremental Improvement. Available at: https://arxiv.org/abs/1804.02767 (accessed 02.06.2020)</mixed-citation></citation-alternatives></ref><ref id="cit11"><label>11</label><citation-alternatives><mixed-citation xml:lang="ru">Carvana Image Masking Challenge. URL: https://www.kaggle.com/c/carvana-image-masking-challenge (дата обращения 02.06.2020)</mixed-citation><mixed-citation xml:lang="en">Carvana Image Masking Challenge. Available at: https://www.kaggle.com/c/carvana-image-masking-challenge (accessed 02.06.2020)</mixed-citation></citation-alternatives></ref><ref id="cit12"><label>12</label><citation-alternatives><mixed-citation xml:lang="ru">Viola P., Jones M. Rapid object detection using a boosted cascade of simple features / Proc. of the IEEE Computer Society Conf. on Computer Vision and Pattern Recognition, Kauai, USA, 8–14 Dec. 2001. Piscataway: IEEE. Art. 7176899. doi: 10.1109/CVPR.2001.990517</mixed-citation><mixed-citation xml:lang="en">Viola P., Jones M. Rapid object detection using a boosted cascade of simple features. Proc. of the IEEE Computer Society Conf. on Computer Vision and Pattern Recognition, Kauai, USA, 8–14 Dec. 2001. Piscataway: IEEE, 2001, art. 7176899. doi: 10.1109/CVPR.2001.990517</mixed-citation></citation-alternatives></ref><ref id="cit13"><label>13</label><citation-alternatives><mixed-citation xml:lang="ru">Silva S. M., Jung C. R. License Plate Detection and Recognition in unconstrained Scenarios // Computer Vision – ECCV 2018. 15th Europ. Conf., Munich, Germany. 8–14 Sept. 2018. Berlin: Springer, 2018. P. 593–609. doi: 10.1007/978-3-030-01258-8_36</mixed-citation><mixed-citation xml:lang="en">Silva S. M., Jung C. R. License Plate Detection and Recognition in unconstrained Scenarios // Computer Vision, ECCV 2018, 15th Europ. Conf. 8–14 Sept. 2018, Munich, Germany. Berlin: Springer, 2018, pp. 593–609. doi: 10.1007/978-3-030-01258-8_36</mixed-citation></citation-alternatives></ref><ref id="cit14"><label>14</label><citation-alternatives><mixed-citation xml:lang="ru">Mask R-Cnn / K. He, G. Gkioxari, P. Dollar, R. Girshick // IEEE Intern. Conf. on Computer Vision (ICCV), Venice, Italy, 22–29 Oct. 2017. Piscataway: IEEE, 2017. Art. 17467816. doi: 10.1109 /ICCV.2017.322</mixed-citation><mixed-citation xml:lang="en">He K., Gkioxari G., Dollar P., Girshick R. Mask R-Cnn.IEEE Int. Conf. on Computer Vision (ICCV). Venice, Italy, 22–29 Oct. 2017. Piscataway: IEEE, 2017, art. 17467816. doi: 10.1109 /ICCV.2017.322</mixed-citation></citation-alternatives></ref><ref id="cit15"><label>15</label><citation-alternatives><mixed-citation xml:lang="ru">Haar Cascad license plate detection. Weights for the model. URL: https://github.com/opencv/opencv/blob/master/data/haarcascades/haarcascade_licence_plate_rus_16stages.xml (дата обращения 15.11.2020)</mixed-citation><mixed-citation xml:lang="en">Haar Cascad license plate detection. Weights for the model. Available at: https://github.com/opencv/opencv/blob/master/data/haarcascades/haarcascade_licence_plate_rus_16stages.xml (accessed 15.11.2020)</mixed-citation></citation-alternatives></ref><ref id="cit16"><label>16</label><citation-alternatives><mixed-citation xml:lang="ru">Silva S. M., Jung C. R. License Plate Detection and Recognition in Unconstrained Scenarios. URL: http://sergiomsilva.com/pubs/alpr-unconstrained/ (дата обращения 31.05.2020)</mixed-citation><mixed-citation xml:lang="en">Silva S. M., Jung C. R. License Plate Detection and Recognition in Unconstrained Scenarios. Available at: http://sergiomsilva.com/pubs/alpr-unconstrained/ (accessed 31.05.2020)</mixed-citation></citation-alternatives></ref><ref id="cit17"><label>17</label><citation-alternatives><mixed-citation xml:lang="ru">Nomeroff Net. A Open Source Python License Plate Recognition Framework. URL: https://nomeroff.net.ua/ (дата обращения 31.05.2020)</mixed-citation><mixed-citation xml:lang="en">Nomeroff Net. A Open Source Python License Plate Recognition Framework. Available at: https://nomeroff.net.ua/ (accessed 31.05.2020)</mixed-citation></citation-alternatives></ref><ref id="cit18"><label>18</label><citation-alternatives><mixed-citation xml:lang="ru">PyTorch implementation of YOLOv3. URL: http://docs.openvinotoolkit.org/2019_R2/_intel_models_person_vehicle_bike_detection_crossroad_1016_description_person_vehicle_bike_detection_crossroad_1016.html (дата обращения 31.05.2020).</mixed-citation><mixed-citation xml:lang="en">PyTorch implementation of YOLOv3. Available at: https://docs.openvinotoolkit.org/latestomz_models_intel_person_vehicle_bike_detection_crossroad_1016_description_person_vehicle_bike_detection_crossroad_1016.html (accessed 31.05.2020)</mixed-citation></citation-alternatives></ref><ref id="cit19"><label>19</label><citation-alternatives><mixed-citation xml:lang="ru">Towards End-to-End License Plate Detection and Recognition: A Large Dataset and Baseline / Z. Xu, W. Yang, A. Meng, N. Lu, H. Huang, C. Ying, L. Huang // Computer Vision. ECCV 2018. 15th Europ. Conf., Munich, Germany, 8–14 Sep. 2018. Berlin: Springer, 2018. P. 261–277. doi: 10.1007/978-3-030-01261-8_16</mixed-citation><mixed-citation xml:lang="en">Xu Z., Yang W., Meng A., Lu N., Huang H., Ying C., Huang L. Towards End-to-End License Plate Detection and Recognition: A Large Dataset and Baseline. Com-puter Vision, ECCV 2018, 15th Europ. Conf. Munich, Germany, 8–14 Sep. 2018. Berlin: Springer, pp. 261–277. doi: 10.1007/978-3-030-01261-8_16</mixed-citation></citation-alternatives></ref></ref-list><fn-group><fn fn-type="conflict"><p>The authors declare that there are no conflicts of interest present.</p></fn></fn-group></back></article>
