Cities are in constant change and city managers aim to keep an updated digital model of the city for city governance. There are a lot of images uploaded daily on image sharing platforms (as “Flickr”, “Twitter”, etc.). These images feature a rough localization and no orientation information. Nevertheless, they can help to populate an active collaborative database of street images usable to maintain a city 3D model, but their localization and orientation need to be known. Based on these images, we propose the Data Gathering system for image Pose Estimation (DGPE) that helps to find the pose (position and orientation) of the camera used to shoot them with better accuracy than the sole GPS localization that may be embedded in the image header. DGPE uses both visual and semantic information, existing in a single image processed by a fully automatic chain composed of three main layers: Data retrieval and preprocessing layer, Features extraction layer, Decision Making layer. In this article, we present the whole system details and compare its detection results with a state of the art method. Finally, we show the obtained localization, and often orientation results, combining both semantic and visual information processing on 47 images. Our multilayer system succeeds in 26% of our test cases in finding a better localization and orientation of the original photo. This is achieved by using only the image content and associated metadata. The use of semantic information found on social media such as comments, hash tags, etc. has doubled the success rate to 59%. It has reduced the search area and thus made the visual search more accurate.
KeywordsPose RecognitionBuilding DetectionSingle Image2D MapCollaborative CartographySocial Media
Grindgis (2015) 67 Important GIS Applications and Uses. http://grindgis.com/blog/gis-applications-uses
The City of New York. NYC Open Data, 2017. https://data.cityofnewyork.us/City-Government/gis/x8zf-jmep
INSPIRE. INSPIRE European Directive, 2017. http://inspire.ec.europa.eu
Lyon City. Lyon Open Data Portal, 2018. https://data.grandlyon.com
Auffray, C. (2015) Infographie—Portrait de l’utilisateurde smartphone francais. http://www.zdnet.fr/actualites/infographie-portrait-de-l-utilisateur-de-smartphone-francais-39796286.htm
Pew Research Center (2017) Mobile Fact Sheet. https://www.pewresearch.org/internet/fact-sheet/mobile/
NainaKhedekar. We Now Upload and Share over 1.8 Billion Photos Each Day: Meeker Internet Report, 2014. https://www.firstpost.com/tech/news-analysis/now-upload-share-1-8-billion-photos-everyday-meeker-report-3652169.html
Goodchild, M.F. (2007) Citizens as Sensors: The World of Volunteered Geography. GeoJournal, 69, 211-221. https://doi.org/10.1007/s10708-007-9111-y
Google Inc. (2017) Local Guides. https://maps.google.com/localguides/home
Wing, M.G., Eklund, A. and Kellogg, L.D. (2005) Consumer-Grade Global Positioning System (GPS) Accuracy and Reliability. Journal of Forestry, 103, 169-173. https://doi.org/10.1093/jof/103.4.169
Chen, X.L., Shrivastava, A. and Gupta, A. (2013) NEIL: Extracting Visual Knowledge from Web Data. 2013 IEEE International Conference on Computer Vision (ICCV), Sydney, 1-8 December 2013, 1409-1416. https://doi.org/10.1109/ICCV.2013.178
Weyand, T., Kostrikov, I. and Philbin, J. (2016) Planet-Photo Geolocation with Convolutional Neural Networks. In: European Conference on Computer Vision, Springer, Berlin, 37-55. https://doi.org/10.1007/978-3-319-46484-8_3
Suleiman, W., Favier, E. and Joliveau, T. (2011) Buildings Recognition and Camera Localization Using Image Texture Description. International Journal of Computer Vision, 61, 159-184.
Lowe, D.G. (2004) Distinctive Image Features from Scale-Invariant Keypoints. International Journal of Computer Vision, 60, 91-110. https://doi.org/10.1023/B:VISI.0000029664.99615.94
Walch, F. (2016) Deep Learning for Image-Based Localization.
Bioret, N., Servières, M. and Moreau, G. (2008) Outdoor Localization Based on Image/GIS Correspondence Using a Simple 2d Building Layer. 2nd International Workshop on Mobile Geospatial Augmented Reality, LNGC, Québec, 28-29 August 2008.
Chu, H., Gallagher, A. and Chen, T. (2014) GPS Refinement and Camera Orientation Estimation from a Single Image and a 2D Map. 2014 IEEE Conference on Computer Vision and Pattern Recognition Workshops, Columbus, 23-28 Jun 2014, 171-178. https://doi.org/10.1109/CVPRW.2014.31
Hochreiter, S. and Schmidhuber, J. (1997) Long Short-Term Memory. Neural Computation, 9, 1735-1780. https://doi.org/10.1162/neco.1997.9.8.1735
Taketomi, T., Sato, T. and Yokoya, N. (2011) Real-Time and Accurate Extrinsic Camera Parameter Estimation Using Feature Landmark Database for Augmented Reality. Computers: Graphics, 35, 768-777. https://doi.org/10.1016/j.cag.2011.04.007
Ventura, J., Arth, C., Reitmayr, G. and Schmalstieg, D. (2014) Global Localization from Monocular SLAM on a Mobile Phone. Visualization and Computer Graphics, IEEE Transactions, 20, 531-539. https://doi.org/10.1109/TVCG.2014.27
Zamir, A.R., Darino, A. and Shah, M. (2011) Street View Challenge: Identification of Commercial Entities in Street View Imagery. 10th International Conference on Machine Learning and Applications and Workshops (ICMLA), Volume 2, 380-383. https://doi.org/10.1109/ICMLA.2011.181
Gallup, D., Frahm, J.-M. and Pollefeys, M. (2010) Piecewise Planar and Non-Planar Stereo for Urban Scene Reconstruction. 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, San Francisco, 13-18 June 2010, 1418-1425. https://doi.org/10.1109/CVPR.2010.5539804
Lafarge, F., Keriven, R., Bredif, M. and Hiep Vu, H. (2013) A Hybrid Multiview Stereo Algorithm for Modeling Urban Scenes. IEEE Transactions on Pattern Analysis and Machine Intelligence, 35, 5-17. https://doi.org/10.1109/TPAMI.2012.84
Hane, C., Zach, C., Cohen, A., Angst, R. and Pollefeys, M. (2013) Joint 3D Scene Reconstruction and Class Segmentation. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, Portland, 23-28 June 2013, 97-104.
Hoiem, D., Efros, A.A. and Hebert, M. (2005) Geometric Context from a Single Image. Tenth IEEE International Conference on Computer Vision, Vol. 1, 654-661. https://doi.org/10.1109/ICCV.2005.107
Hoiem, D., Efros, A.A. and Hebert, M. (2005) Automatic Photo Pop-Up. ACM Transactions on Graphics, 24, 577. https://doi.org/10.1145/1073204.1073232
Achanta, R., Shaji, A., Smith, K., Lucchi, A., Fua, P. and Susstrunk, S. (2012) SLIC Superpixels Compared to State-of-the-Art Superpixel Methods. IEEE Transactions on Pattern Analysis and Machine Intelligence, 34, 2274-2281. https://doi.org/10.1109/TPAMI.2012.120
Liu, B.Y., Gould, S. and Koller, D. (2010) Single Image Depth Estimation from Predicted Semantic Labels. 2010 IEEE Conference Computer Vision and Pattern Recognition (CVPR), San Francisco, 13-18 June 2010, 1253-1260.
Wang, G.H., Chen, X.J. and Chen, S. (2014) Cut-and-Fold: Automatic 3D Modeling from a Single Image. 2014 IEEE International Conference Multimedia and Expo Workshops (ICMEW), Chengdu, 14-18 July 2014, 1-6. https://doi.org/10.1109/ICMEW.2014.6890555
U.S. Government (2017) GPS Accuracy. http://www.gps.gov/systems/gps/performance/accuracy
Chen, H.Z., Tsai, S.S., Schroth, G., Chen, D.M., Grzeszczuk, R. and Girod, B. (2011) Robust Text Detection in Natural Images with Edge-Enhanced Maximally Stable Extremal Regions. Proceedings International Conference on Image Processing, ICIP, Brussels, 11-14 September 2011, 2609-2612. https://doi.org/10.1109/ICIP.2011.6116200
Neumann, L. and Matas, J. (2012) Real-Time Scene Text Localization and Recognition. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, Providence, RI, 16-21 June 2012, 3538-3545.
Gómez, L. and Karatzas, D. (2014) MSER-Based Real-Time Text Detection and Tracking. 22nd International Conference on Pattern Recognition, Stockholm, 24-28 August 2014, 3110-3115. https://doi.org/10.1109/ICPR.2014.536
Von Gioi, R.G., Jakubowicz, J., Morel, J.-M. and Randall, G. (2012) LSD: A Line Segment Detector. Image Processing on Line, 2, 35-55. https://doi.org/10.5201/ipol.2012.gjmr-lsd
Rother, C. (2002) A New Approach to Vanishing Point Detection in Architectural Environments. Image and Vision Computing, 20, 647-655. https://doi.org/10.1016/S0262-8856(02)00054-9
Zamir, A.R. and Shah, M. (2010) Accurate Image Localization Based on Google Maps Street View. In: European Conference on Computer Vision, Springer, Berlin, 255-268. https://doi.org/10.1007/978-3-642-15561-1_19
Kirillov, A., Girshick, R., He, K.M. and Dollár, P. (2019) Panoptic Feature Pyramid Networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, 15-20 June 2019, 6399-6408. https://doi.org/10.1109/CVPR.2019.00656
Huang, Z.J., Huang, L.C., Gong, Y.C., Huang, C. and Wang, X.G. (2019) Mask Scoring R-CNN. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, 15-20 June 2019, 6409-6418. https://doi.org/10.1109/CVPR.2019.00657