In this podcast episode, Jennings Anderson, a research scientist at Meta, discusses the Overture Maps Foundation, a downstream product of OpenStreetMap.
He explains his background in open map data and his interest in studying collaboration within the OpenStreetMap community.
Jennings then dives into the Daylight Distribution, an open data product produced by Meta, and how it combines building data sets from various sources into one unified theme.
Jennings emphasizes the importance of a stable ID system within the Overture Maps Foundation and the potential for easy conflation and integration of third-party data.
Jennings also explains the relationship between OpenStreetMap and Overture Maps, highlighting how they complement each other.
Relevant podcast episodes
OpenStreetMap Is A Community Of Communities
Cloud Native Geospatial
Cloud Optimized Point Clouds
The Rapid Editor
With regards to accessing Overture Map data, you might find this YouTube video helpful
https://youtu.be/fZj6kTwXN1U?feature=shared
Just in case you are interested in the Google building footprints here is a link to that :)
https://sites.research.google/open-buildings/
100 billion Points Every Day
100 billion is a very large number, let's say that I gave you a spreadsheet with 100 billion rows in it, each row consisted of five columns Latitude, Longitude, Device ID, A Timestamp, and a column telling the name of the data provider
What would you do with that?
How would you clean it? Make sense of it? Extract value from it? What would people use it for? And how would you do this in a way that could be systematized?
FourSquare does this every day with the help of something they call a movement engine.
To help understand more about how they do this I have invited Gabriel Durkin the director of data science on the podcast. This is the last in a series of episodes I have worked on together with FourSquare and I have to say it's been really enjoyable working with them.
If you are interested in hearing some of the previous episodes just check out the links below!
From Pixels To Patterns AI In Spatial Analysis
https://mapscaping.com/podcast/from-pixels-to-patterns-ai-in-spatial-analysis/
Big Data In The Browser
https://mapscaping.com/podcast/big-data-in-the-browser/
Spatial Knowledge Graphs
https://mapscaping.com/podcast/spatial-knowledge-graphs/
Designing For Location Privacy
https://mapscaping.com/podcast/designing-for-location-privacy/
All Of The Places In The World
https://mapscaping.com/podcast/all-of-the-places-in-the-world/
Geospatial Jobs
There are a few new jobs on our Job Board! The most interesting one is the role of Social Media Manager at Felt - United States (Remote)
( If you want to apply for this one, it might be a good idea to listen to this episode first ;)
https://mapscaping.com/podcast/felt-upload-anything/ )
See more at https://mapscaping.com/jobs/
As a bonus for reading all the way to the end :)
If you are looking for free terrain data for anywhere in the world you might find this useful
https://github.com/openterrain/openterrain/wiki/Terrain-Data
Computer vision is everywhere! But teaching an algorithm to identify objects requires a lot of data and this is definitely the case when we think about GeoAI
But it is not enough to have a lot of data we also need data that is labeled
If we are looking for cars in images we need a lot of images of cars and we need to know which pixels are the car!
Of course, I am oversimplifying but I hope you get the idea,
Now imagine that you can automatically generate a large labeled data set of realistic images of cars based on the specifications of a specific sensor.
These data sets are often referred to as synthetic data or fake data and to help us understand more about this I have invited Chris Andrews from Rendered AI on the podcast.
Here are a few previous episodes you might find interesting
Computer Vision And GeoAI
https://mapscaping.com/podcast/computer-vision-and-geoai/
In this episode, the discussion is aimed at an increased understanding of the differences between computer vision and the AI that is used in the Earth Observation world.
Labels Matter
https://mapscaping.com/podcast/labels-matter/
What it takes to create labeled training data manually. If you are new to the idea of labeled data sets this is a good place to start.
Fake Satellite Imagery
https://mapscaping.com/podcast/fake-satellite-imagery/
This is a good episode if you want to know more about Generative AI and Generative Adversarial Networks.
Also, check out this website https://thisxdoesnotexist.com/ to get an idea of where and how these Generative Adversarial Networks can be used. Look for a website called This City Does Not Exist http://thiscitydoesnotexist.com/
On a silently similar note try uploading an image to https://bard.google.com/ … it's pretty interesting!
In this episode of the Mapscaping Podcast, we had the pleasure of speaking with Sam Hashemi, the CEO and co-founder of Felt, a browser-based mapping tool that is making waves in the geospatial community. Felt is not just another mapping tool; it's a platform that's revolutionizing the way we interact with geospatial data, thanks to its unique "upload anything" feature.
The "Upload Anything" Feature: Felt's "upload anything" button is a game-changer. It allows users to upload any data without specifying the type, and the system figures out the data on the backend. This feature sets Felt apart from traditional mapping tools, which often require users to specify the type of data they are uploading. The "upload anything" feature is a testament to Felt's mission to make mapping more accessible and user-friendly, reducing the technical barriers often associated with geospatial data handling.
Felt's Contribution to Open Source Projects: Felt is not just about creating a user-friendly mapping tool; it's also about giving back to the community. Felt is the first and only flagship sustaining member of the QGIS project. They are building an open-source tiling engine called Tippecanoe and also support Proto maps and the development of PM tiles. They contribute code to MapLibre and GDAL as well.
The Genesis of Felt: Sam Hashemi, a product designer by profession, started Felt after noticing the difficulty in creating a basic map, let alone a complex one. While there were powerful desktop software like QGIS and Esri, there was no modern, fun, and playful tool on the internet. This led to the creation of Felt, aiming to be the Google Sheets for maps.
Felt's Vision: Felt aims to make map creation fun, playful, and delightful. They believe in increasing the "GDP of maps," i.e., the number of maps created every year, by making the process more accessible and enjoyable. They envision a world where many more people are creating many more maps because it's fun, creative, and useful in their day-to-day lives.
Felt's Business Model: Starting in 2024, Felt will charge a monthly fee for using the software. However, there will always be a free tier for consumers. They take privacy seriously and do not sell or look at user data.
Conclusion: Felt is still in its early stages, having launched the company two years ago and the software one year ago. They are starting to see exciting ways people are using the software and believe they are at the beginning of something big. As Sam Hashemi puts it, "If I was to impress one thing upon your viewers, it's Felt. It's free, go to felt.com and you can sign up and try it. No need to listen to my words or think about it, go feel it in your hands, go get that Felt experience." So, why not give Felt a try and experience the future of geospatial mapping?
Rapid is a free open-source web-based editor for an OpenStreetMap. In the past the focus was on conflating AI-generated datasets with OpenStreetMap data but the future for this editor is conflating authoritative datasets with OpenStreetMap.
Humans are in the loop, people reviewing data authoritative datasets and adding them to OpenStreetMap with a few clicks!
So you might be wondering, what is Authoritative data? And perhaps it doesn’t even matter what authoritative means maybe the most important thing is it correct.
If you are interested in OpenStreetMap you might enjoy this episode
https://mapscaping.com/podcast/openstreetmap-is-a-community-of-communities/ which is a great introduction to OpenStreetMap as a project but also explains some of the commercial interest in updating the map which adds a lot of context to Rapid and its development and future.
Humanitarian OpenStreetMap Team
If you have not heard of Humanitarian OpenStreetMap Team this is well worth checking out!
https://www.hotosm.org/
Segment Anything
(SAM) can segment objects by simply clicking or interactively selecting points to include or exclude from the object. This makes it a user-friendly tool for image segmentation.
https://segment-anything.com/
Mapillary
Mapillary is a platform that provides street-level imagery and map data from all over the world. The platform is powered by collaboration and computer vision, which helps in generating and maintaining up-to-date, detailed maps.
https://mapscaping.com/podcast/scaling-map-data-generation-using-computer-vision/
GeoSpatial Jobs!
Drone Deploy’s Earthworks team is looking for an experienced Back End Engineer Full time / Remote
NV5 is looking for a Senior GIS Analyst
https://mapscaping.com/jobs/
Want to help? I could really use some support!
The promise of digital mapping is to provide a shared and real-time view of the state of the underlying system.
pg_eventserv is a free and open-source component that helps fulfill the promise of real-time event modeling and shared views in PostgreSQL.
By connecting to PostgreSQL and listening on specified channels, pg_eventserv captures database notifications and forwards them to web clients, enabling real-time updates and synchronization of data displayed on maps or other web interfaces.
pg_eventserv does one thing and one thing only: take events generated by the PostgreSQL NOTIFY command and passes the payload along to waiting WebSockets clients.
pg_eventserv is free and easy to install and you can find it here: https://github.com/CrunchyData/pg_eventserv
What this means is that any client can watch for notifications and update as changes in the database happen.
Real-time data!
Here is a link to a Youtube video demonstration of pg_eventserv in action!
https://youtu.be/UakRtYmoWow
I will let Paul Ramsey the creator of pg_eventserv explain all this in more detail in this episode.
If you want to reach out to Paul the best place to do that is http://blog.cleverelephant.ca/
Or if you want to listen to previous episodes with Paul you might find these interesting
Raster in the database?
https://mapscaping.com/podcast/rasters-in-a-database/
Dynamic Vector Tiles Straight From The Database
https://mapscaping.com/podcast/dynamic-vector-tiles-straight-from-the-database/
Spatial SQL- GIS Without The GIS
https://mapscaping.com/podcast/spatial-sql-gis-without-the-gis/
also ... If you are interested in spatial databases at scale ... you might find this episode interesting
https://mapscaping.com/podcast/distributing-geospatial-data/
How do we get data from a satellite down to Earth? How do we task a satellite?
Today the answer is likely to be via radios and a system of downlink sites or ground stations. As the satellites pass overhead or within “line of sight” data can be sent via radio from the satellite to the receiver on the ground.
If you don’t want to wait until the satellite can see the ground station, you can send your data to a geostationary satellite that can always see a ground station and let it send the data back to Earth.
Radios are tried and tested, they have been used for this purpose since the inception of satellite communication and radio waves can pass through Earth's atmosphere without significant loss!
But … the frequency spectrum for radio waves is strictly regulated, which can limit available channels for communication, and the bandwidth of radio frequencies is limited, which can reduce the volume of data transmission.
What about lasers?
You can send more data faster with a laser, you don’t need to worry about interfering with someone else part of the radio spectrum, and ground stations can be much smaller even human-portable!
But … lasers struggle with clouds and the technology is still relatively new
So what is the best way to communicate with satellites? Radio or Laser? The answer is … it depends ;)
Jordan Wachs, Director of Business Development for SpaceRake.net does a great job adding context to this discussion but perhaps the bigger question here is what will we do when satellites become internet devices, part of the Internet of Things?
What if they were always on always connected in the same way your phone is always on, always connected? What will this enable?
This episode was sponsored by Sponsored by Sinergise, as part of Copernicus Data Space Ecosystem knowledge sharing
People who liked this episode also liked …
How to keep your satellite pointing at earth
https://mapscaping.com/podcast/how-to-keep-your-satellite-pointing-at-earth/
Hyperspectral v’s Multispectral
https://mapscaping.com/podcast/hyperspectral-vs-multispectral/
Sentinel Hub
https://mapscaping.com/podcast/sentinel-hub/
Swing by our website sometime https://mapscaping.com/
pygeoapi is a Python server implementation of the OGC API suite of standards ... which might be really useful if you are thinking about upgrading from the first-generation OGC standards to the second-generation OGC standards
... or if need to implement a custom data source or custom functionality to your web services.
If you are using MapServer, GeoServer, MapProxy, QGIS server, or Deegree you might find this episode interesting!
Relevant previous episodes
Cloud-native Geospatial
Geoserver
Geonode
So why would anyone want to put alot of data into a browser? Well, for a lot of the same reasons that edge computing and distributed computing have become so popular.
You get the data a lot closer to the user and you don’t have to pay for the compute ;)
… this sounds great but as I found out during this conversation it's not as easy as it might seem!
There are a lot of trade-offs that need to be evaluated when moving data and analytics to the client.
Nick Rabinowitz Senior Staff Software Engineer at Foursquare has a ton of experience with this so he volunteered his time to help us understand more about it.
https://location.foursquare.com/
https://studio.foursquare.com/home
If you are not familiar with the Arrow data format it might be worth checking out
Apache Arrow defines a language-independent columnar memory format for flat and hierarchical data, organized for efficient analytic operations on modern hardware like CPUs and GPUs. The Arrow memory format also supports zero-copy reads for lightning-fast data access without serialization overhead
Related podcast episodes that you might find interesting include
H3 grid system
https://mapscaping.com/podcast/h3-geospatial-indexing-system/
The H3 geospatial indexing system is a discrete global grid system consisting of a multi-precision hexagonal tiling of the sphere with hierarchical indexes. H3 is a really interesting approach to tiling data that was developed by UBER and has been open-sourced.
Hex Tiles
https://mapscaping.com/podcast/hex-tiles/
If you have not heard of the H3 grid system before listen to that episode first before listening to this one it will add a lot of useful context!
Spatial Knowledge Graphs
https://mapscaping.com/podcast/spatial-knowledge-graphs/
Foursquare is moving away from spatial joins and focusing on building a knowledge graph. If you are not familiar with graphs this might be a good place to start, also its interesting to hear the reasons for the move from spatial joins to another data structure.
Distribution Geospatial Data
https://mapscaping.com/podcast/distributing-geospatial-data/
This is interesting if you want to understand more about distributed databases and some of the strategies for doing this. It sounds complicated but this episode is a really good introduction!
Cloud Native Geospatial
https://mapscaping.com/podcast/cloud-native-geospatial/
This episode give a solid overview of what cloud-native means and some of the current geospatial cloud native formats out there today
I am constantly thinking about how I can make this podcast better for you so if you have any ideas or suggestions please let me know!
Also, I am thinking of recording a behind-the-scenes episode, is that something you might be interested in? if so what questions do you have?
Sounds like a great idea right?
In this episode, Paul Ramsey explains why you shouldn't ... unless you want to ... and how you can ... if you have to.
You can find Paul's blog here: http://blog.cleverelephant.ca/about
Previous episodes with Paul
Spatial SQL
https://mapscaping.com/podcast/spatial-sql-gis-without-the-gis/
GDAL
https://mapscaping.com/podcast/gdal-geospatial-data-abstraction-library/
Dynamic Vector Tiles
https://mapscaping.com/podcast/dynamic-vector-tiles-straight-from-the-database/
Blog posts by Paul about Rasters in the Database
https://www.crunchydata.com/blog/postgres-raster-query-basics
https://www.crunchydata.com/blog/waiting-for-postgis-3.2-secure-cloud-raster-access
Check Out Our Geospatial Job Board!
https://mapscaping.com/jobs/
I am sure you have heard of ChatGPT by now so the hope of this episode is to give you some more context about what is it built on and how it works.
To do that I invited Daniel Whitneck back on the podcast
You can connect with Daniel here
https://datadan.io/
and listen to his previous episode here:
https://mapscaping.com/podcast/an-introduction-to-artificial-intelligence/
This is perhaps the quote for the episode that I have spent the most time thinking about
"We always thought AI would be logical and lack creativity - but it is almost the exact opposite"
This reframes the idea of being wrong to being creative which I think you could argue really depends on the context!
If you have not already played around with ChatGPT it's well worth spending the time to experiment with it ... while its still free ;)
https://chat.openai.com/auth/login
Further listening
If you have not already listened to this episode about computer vision and GeoAI you might find it interesting. Listen out for the discussion around plausible / realistic data and real measurements - I think this gives more context to the use cases for generative AI
https://mapscaping.com/podcast/computer-vision-and-geoai/
You might also enjoy this episode about fake satellite imagery
https://mapscaping.com/podcast/fake-satellite-imagery/
BTW I have started a job board for geospatial people
feel free to check it out!
Computer vision is a field of artificial intelligence (AI) that enables computers and systems to derive meaningful information from digital images.
You might think that this is exactly what we are doing in earth observation but there are a few important differences between computer vision and what some people refer to as GeoAI.
This week Jordi inglada is going to help you understand what those differences are and why it's not always possible to use Computer vision techniques in the field of Remote Sensing.
Listen out for these key points during the conversation!
Sponsored by Sinergise, as part of Copernicus Data Space Ecosystem knowledge sharing. dataspace.copernicus.eu/ http://dataspace.copernicus.eu/
Related Podcast Episodes
Super Resolution
https://mapscaping.com/podcast/super-resolution-smarter-upsampling/
Fake Satellite Imagery
https://mapscaping.com/podcast/fake-satellite-imagery/
Sentinal Hub
https://mapscaping.com/podcast/sentinel-hub/
Google Earth Engine
https://mapscaping.com/podcast/introducing-google-earth-engine/
Microsofts Planetary Computer
https://mapscaping.com/podcast/the-planetary-computer/
BTW MapScaping has started a Job Board!
it's in the early stages but it's live
Jobs - Mapscaping.com
When comparing multispectral and hyperspectral data it is not simply a case of “more data more better”!
With hyperspectral you have “The curse of Dimensionality” but you also get more flexibility to pick exactly what bands you want to use!
With multispectral you have less noise but you also have less data!
This episode is designed to be a beginner's guide to the differences between hyperspectral and multispectral satellite data. Sponsored by Sinergise, as part of Copernicus Data Space Ecosystem knowledge sharing. dataspace.copernicus.eu/ http://dataspace.copernicus.eu/ You can reach out to Gordon Logie here: https://sparkgeo.com/blog/team/gordon/
Here are some courses that focused on hyperspectral and offer further training
https://eo-college.org/courses/beyond-the-visible/ https://eo-college.org/courses/beyond-the-visible-imaging-spectroscopy-for-agricultural-applications/ https://www.enmap.org/events_education/hyperedu/
Foursquare's ongoing work to map every place in the world
About The Guest Kyle Fowler is the senior director of engineering at Foursquare. Before joining Foursquare in 2011, he was a third party user of Foursquare developer tools. Ever since, he has worked on different projects, in varying roles to solve a variety of location problems.
What Is Foursquare? Foursquare is a cloud-based location technology platform that supports building solutions based on a deep understanding of location. For the general public, Foursquare is a location-based social network that facilitates meeting up with friends and discovering new places. Users can check-in to places and show their network where they are, and what they are doing. On the other hand, commercial companies and organizations use Foursquare APIs and advertising enablement products to gain insights from the physical world.
Foursquare Swarm Swarm is a life-logging app that keeps a user’s history of interaction with the world. It is a perfect diary for keeping a personal log of the places you have visited, alongside photos, and information about the friends you have been with, and the experiences you had. Logging places into Swarm can happen passively in the background or through active interaction with the app. In active logging, you have to interact with the app in every place you visit. But passive loggers can also actively build their history by periodically reviewing their history feed to confirm the places they have visited.
How to Download Your Swarm Data One way to download your Swarm data is by requesting it through your profile. The data will include all your check-ins and any other information you logged. The other way is through the Foursquare API after signing up to be a Foursquare developer (anyone can sign up). Using an API is the best way to pull your data if you want to use it for running analysis or building your own features.
How Is Foursquare Mapping All Of The Places In The World? Crowdsourcing is a huge contributor to Foursquare location data. But Foursquare may not have users in every corner of the world. To get a good amount of coverage the data is obtained from several other sources that include machine-generated approaches and from other companies. For completeness of coverage, Foursquare purchases licenses for regional data generated by other trusted companies that have a particular interest in a specific area.
How Does Foursquare Control the Quality of Location Data? All the places submitted at Foursquare go through a quality audit process to ensure that the location is real. The initial screening removes low accuracy points that may introduce skewed locations.
Machine learning models are also used to review the input sources and assign a confidence score of whether a place exists. Further controls also involve sending samples of places to human moderators for review.
Foursquare understands that the negative outcomes of being wrong about a place are worse than the place not being listed at all. If people using that location information actually go there and find that it does not exist, it results to a bad user experience and a negative perception on the app.
A quality auditing of the data helps to take out all the bad sources and ensures the recorded locations are an accurate reflection of the real world.
Maximizing Accuracy with the Geosummarizer Foursquare’s Geosummarizer is a model that analyzes the inputs for a certain POI (Point of Interest) and selects the right coordinates.
The Geosummarizer helps to solve different challenges such as when people record slightly different coordinates for the same place they check into. This is more likely to happen in dense urban environments especially when the user is inside a building.
To try and find the right coordinate, the Geosummarizer compares the input coordinate to the geocodes of that place’s address and infer whether the coordinate should exist within that place. The process then picks the best coordinate that has the most corroborating features with that address.
If there are multiple clusters of coordinates for a given location, the Geo-summarizer tries to assign them appropriately to different geocodes. A place can have multiple geocodes; and the users will use the one appropriate for them based on the use case.
For instance, an app may want a pickup and drop-off location of a POI (Point of Interest). The Geosummarizer also outlines various levels of venue hierarchy for describing relationships of places within places such as a gate inside an airport terminal.
Locations for mobile things like food trucks are treated separately from other categories. The summarization process for generating a new coordinate for that venue takes the latest data into account much more heavily than it would for a brick-and-mortar store that is not expected to move.
Mobile POIs can get updated in real-time – on every single check-in – as long as the coordinates are believed to be legitimate.
Foursquare’s Taxonomy of Places Foursquare has over 1200 categories for classifying places across the world. In an effort to make the places locally familiar, some of those categories are only visible in certain countries. For instance, one would not expect a Chinese Food category in China, or Italian restaurants to show up in Italy.
Foursquare tries to adapt the categories system to reflect the depth of reflectiveness for a particular country. Categories are re-evaluated on a monthly basis to ensure they are showing up appropriately, and whether new categories should be added.
Mapping all the places in the world and making them locally familiar is a challenging task. However, Foursquare’s knowledge of the world is continually increasing. As more locations are added and their model is continuously improving, Foursquare is steadily moving towards achieving this goal.
es I am going to publish in partnership with Foursquare and the idea is to use it as a reference for later episodes about Privacy and location data, Knowledge Graphs, AI, Location Based Marketing and Big geospatial Data in the Browser.
Protomaps is a serverless system for planet-scale maps, it's an umbrella project consisting of a few different components one of which is PMtiles.
PMtiles is “Cloud Optimise Geotiff” for web mapping, what this means is that you can build a base map and host it without the need for a server!
PMtiles is a single file that you can access via HTTP range requests in the same way that you can access data within a Cloud Optimised Geotiff with the important difference that PMtiles can also contain vector data!
What this means is that you can create your own base map, and host it on something like Amazon S3 object storage at a fraction of the cost of other base map solutions!
During this episode, you will hear Brandon, the founder, and creator of Protomaps, talk about scarcity, and well I have never really thought about base maps as being a scarce resource I can definitely see how a product like PMtiles could remove some of the barriers to entry for a lot of creativity in terms of base maps.
More information on Protomaps is here: https://protomaps.com/
Tippecanoe
https://github.com/felt/tippecanoe.git
https://bertt.wordpress.com/2023/01/06/creating-vector-pmtiles-with-tippecanoe/
Relevant podcast episodes
Cloud Optimized Point Clouds
https://mapscaping.com/podcast/cloud-optimized-point-clouds/
Cloud Native Geospatial
https://mapscaping.com/podcast/cloud-native-geospatial/
Microsoft’s Planetary computer
https://mapscaping.com/podcast/the-planetary-computer/
Stamen Design - Full Stack Cartography
https://mapscaping.com/podcast/full-stack-cartography/
If you have any questions or comments, let me know, I would love to hear from you!
Modeling The Vision Of A Better City With LIDAR
LIDAR is a fast and efficient method to capture high-resolution data with accurate measurements. But in this episode, the focus is not on using LIDAR to make better measurements. Rather the talk is about how point cloud data can be used to tell a story. The guest tells how creating a film with 3D LIDAR data of a city provided a medium that could be consumed by the public to understand how a city could be transformed.
About The Guest Benjamin Muller is a team leader at Leica Geosystems in the R&D section. Doing mechatronics, they combine disciplines such as software development, mechanics, and electronics to develop devices such as drive chains for laser scanners and total stations.
Benjamin was part of a project that used LIDAR data to create a film that communicated the vision of a better city. He emphasizes the importance of using easily understandable films to communicate complex details. In his words:
“Facts do not always speak for themselves. If a visualization can be wrapped around the truth, and packaged in a story, it can reach a lot more people. And this could be a driver for change.”
How Is LIDAR Useful In 3D City Modelling? Getting insights into a whole city is an important element in the process of transforming a city into a future-proof city. It is not possible to get this kind of insight when walking through a city because at that moment you are only viewing the city from a very small perspective.
LIDAR makes it possible to make 3D models of the whole city, with accurate measurements that preserve the integrity of the city features. Visualizing a city in this way provides a wider perspective on how various changes would impact a city. Viewing a city and its surroundings in a new perspective unlocks new ideas and better strategies of transforming a city. Additionally, LIDAR data collection can be done whenever required to capture any new changes that have happened in the city.
Drawbacks in Working with LIDAR During LIDAR data collection, massive amounts of point clouds are captured. Even for a small city, the size of data captured can easily grow into several terabytes. With data of this size, many LIDAR experts have reported that a great part of their work, over 80%, is filtering the data. It can take months to process LIDAR data for a city.
Processing LIDAR data for a city is costly both in terms of time and money – which can run into several hundreds of thousands of dollars. Using AI in processing can considerably reduce processing times, but even then working with the data requires expensive and powerful equipment. For most people, who may just need a one-time visualization of a part of a city, it is uneconomical to buy such expensive infrastructure just to visualize a small part of the city.
How LIDAR Was Used To Visualize a Better and More Sustainable City A project undertaken in the city of St Gallen in Switzerland used LIDAR data to create 3D models of the city. It added contributions from different experts to create subsequent versions of how the city can be transformed. These versions were pieced together to create a 30-minute long film that helped to communicate to people a vision of a greener, more sustainable city.
The advantage of using a film as a communication tool is that it is easier to understand. It effectively communicates the message to people who may not deliberately read detailed books on how to transform a city. Additionally, a film can efficiently reach a wider audience in society and inspire people to become a part of the change for a better future city.
From A Vision to A Strategy The film of the city was created to communicate a vision of a better future city. But it also transformed to a strategizing tool. With this visualization, decision-makers could formulate a step-by-step strategy that can be applied to make the city more sustainable and healthy for both the people and the whole environment. This led to the creation of 16 books alongside the film, which contained detailed information on how to do the desired transformations in the city.
Other Uses of 3D City Models Various sectors including tourism can benefit from 3D models of a city. Using augmented reality the 3D models can be used to support digital tourism. While visiting a site, tourists can also view the site from different perspectives, and experience things that would otherwise not be obviously visible.
For cities where you need to report the changes you are doing to structures on your property to city authorities, visualizing the changes in a 3D model provides an opportunity for understanding the impact of the changes to the city and the environment.
How Can We Democratize 3D City Data? With AI it is possible to build services that process LIDAR data on the cloud. These services can then be accessed through a browser by the public. In this way, more people can visualize the data without having to buy very expensive infrastructure or become an expert in visualization. Providing easier and cheaper ways for the public to consume LIDAR data is an important step in its democratization. But if only a few specialists are able to use the data, it will still remain a very expensive luxury to have it.
HxDR (Hexagon Digital Reality) HxDR is one example of a publicly available platform for viewing 3D models of cities. You have to register first before you can view detailed 3D city models. At the moment, there are models of several cities in Europe and the US on the platform. Some of the cities have high-precision geospatial data and you can also purchase parts of the city that you are interested in.
A platform like this is only the beginning of a transformation to have access to 3D data, but we expect there is more coming in the future.
If you are interested in more technical episodes about point clouds you might enjoy these! The Point Data Abstraction Library Cloud Optimized Point Clouds Bathymetric Lidar Lidar from Drones Lidar from Space
Personally, I don't feel like aerial imagery gets the attention it deserves! So I invited Michael Bewley - Senior Director of AI Systems at Nearmap back on the podcast to help bring us up to speed on the state of the art of capturing, processing, and building a business around aerial imagery.
If you don’t care about aerial imagery, think of this as a story about turning unstructured data into structured data into insights and building a business around that.
You can connect with Micheal on Twitter and LinkedIn
Listen out for the following highlights
Previous Interview with Michael Bewley
Stratospheric Balloons As Remote Sensing Platforms
How Does Your Phone Know Its Location?Today’s show explains how your mobile device determines your location, commonly displayed on a map using the popular ‘blue dot’. Our guest is Ed Parsons, Google's Geospatial Technologist. He has been at Google for over 15 years; but before that he came from academia, and even helped to set up one of the first GIS courses taught at Kingston University. Prior to that, he worked for the Ordnance Survey and the National Mapping Agency in the UK.
The Blue Dot and GPSThe blue dot on a map shows your location as determined by your mobile device. With the blue dot, you no longer have to manually figure out where you are, as you would have to if you were using a paper map.
The very first mobile devices with the capability to determine location were using GPS (Global Positioning System). To this day, many people still think that GPS is the only satellite constellation that all mobile devices use to determine location. GPS is actually only one of the several GNSS (Global Navigation Satellite System) constellations used to determine location. GPS is an American system; other GNSS include GLONASS of Russia, Galileo of the European Union, and Baidu of China, among others.
Why Does Your Phone Need To Know Its Location?Despite helping you in navigation, your phone also needs to know its location for its basic functions. Knowing a phone’s location, service providers are able to route its calls through the closest cell towers. Messaging and calling is not possible if a phone’s location is not known. Therefore, all mobile networks have to maintain a rough location of where a device is in relation to their cell towers. This is a key way that a device is able to know its location.
Setbacks in GNSSWhen a device is using GNSS to determine its location, it needs a line of sight to the satellite sending the timing signals so as to calculate an accurate location. If there's an obstruction between the device and the satellite, (such as a tree, a building, or a mountain) the device may not get a signal at all, or may suffer multipath problems, which is when the signal bounces off the objects during its journey, resulting in inaccurate timing.
For a device to use GNSS, its radio receivers have to be turned on so as to receive the satellite signals. The power needed to keep these radios on for prolonged periods, as well as the power needed to do the actual computation of location from the satellite signals, can drain the phone’s battery rather quickly. Power management is a major concern in using GNSS, and several alternatives such as using Wi-Fi hotspots to determine location have been developed over the years to try and minimize power consumption.
Location Determination with Wi-Fi HotspotsDetermining location based on Wi-Fi hotspot proximity is widely used today due to the popularity and presence of Wi-Fi. From a home broadband to public Wi-Fi in railway stations, airports, or coffee shops – there are Wi-Fi hotspots all over. Each Wi Fi hotspot has a unique identifier – an address that is part of the internet’s network infrastructure. The Wi-Fi hotspots and their locations can be built into a database, and cached into a mobile device for use in location determination. There are a number of providers of the databases that match Wi -Fi hotspots to locations. A device just needs to keep its database updated to capture the changes when Wi-Fi hotspots are moved around.
Using Wi-Fi hotspot to determine location is more efficient in terms of power consumption than using GNSS technology. In most cases, people already have their Wi-Fi on as that is how most people access the internet today. This means the device does not use a large deal of extra battery power to find its location. Since the phone switches between the location determination technologies in the background, where you may be thinking that your device is giving you a location coming from GNSS technology, it is most likely the Wi-Fi infrastructure and technology identifying where you are.
Error CorrectionSometimes a device displays an incorrect location before correcting itself in the background and updating to the correct location. A common cause for this sudden shift is when your device initially picks up a stale Wi-Fi location, especially on startup. By checking against the nearest cell tower, it disregards the location from the stale Wi-Fi and updates to your correct location.
For multipath problems when using GNSS, using a model of a city’s morphology can help in error correction. The model gives insight on how the satellite signals bounce off objects in the area. The multipath errors caused by these bounces can then be canceled out and corrected for, increasing the location accuracy.
Fused Location ServiceCell towers, Wi-Fi, and GNSS constellations are the three main technologies used to determine a device’s location. A fused location service means that the choice of the technology being used at any particular point in time is largely invisible to the user, and in many cases from the application developer as well.
The device’s choice of which technology to use is a capability at the operating system level. All that is needed at the application level is making a single call to the operating system with a request for the location level that is required. Depending on whether the application needs a precise location or a relative location, the operating system will decide on the best technology for providing the location at the required accuracy. Hence, it is very likely that a device will only use GNSS to get location when precise location is needed, for instance, if an application is giving turn by turn directions. If proximal location is needed, Wi-Fi will most likely be used due to its power efficiency.
Visual PositioningVisual positioning uses the phone’s camera to sense the environment around it. It makes a comparison between what the camera sees against a database containing images of that street. The device can then orient itself to the correct position and orientation during navigation. Google in particular has a rich resource of street view imagery of many cities and locations around the world. Google Maps users can use visual positioning to orient themselves correctly and know which side of the street they are on, and which direction they are facing. The visual positioning system supplements other location services to not only identify your location, but also the orientation (i.e. the compass direction to which you are facing or heading).
Do You Always Need A Precise Location All the Time?Precise location is not always what we need for most of our day to day location needs. Now, if you are doing turn by turn navigation as a pedestrian, then you definitely need precise location. If your goal is to get a local weather forecast, or find your local Starbucks, then a proximity location that puts you on the right street would suffice.
COPC – A Cloud Efficient Data FormatIn this episode, the discussion revolves around cloud-optimized point clouds. Our guest is Martin Dobias, the CTO at Lutra Consulting. Coming from a background in computer science, and a passion for geospatial, Martin has been part of a team that has done a ton of interesting work in the open source geospatial world. Today, he shares about their latest, state of the art developments in working with point clouds on the web; Cloud Optimized Point Clouds (COPC).
What Does Cloud Optimized Point Clouds Even Mean?Point clouds are sets of individual points plotted in 3D space. They are typically very large datasets, as they must capture a real space in great detail. For instance, point clouds for a whole country may easily be several trillions of points, and many terabytes of data. Handling these large datasets on the web requires a lot of bandwidth to download and process.
The files for Cloud Optimized Point Clouds are structured with indexes for each part of the dataset. The index structure makes it possible to stream only the parts of the data that are required, without having to download the entire dataset.
What is Point Cloud Indexing?Point cloud indexing structures a point cloud file, making it possible to find any particular point of interest in the file without having to scan through the entire dataset. COPC files are internally indexed using a 3D structure of cubes called Oak trees. One part of the file is the data itself, and the other part contains the hierarchical information of where to find each cube.
At the root level of the hierarchy is a single cube, which at the next level is split into eight smaller cubes. The splitting continues subsequently up to the highest hierarchical level. As the cubes get smaller, they contain a smaller amount of data, which saves bandwidth if a user is only interested in a small part of the data. It is similar to traditional tiling, but with the added 3D context.
How COPC Files Are Accessed Using HTTP Range RequestsRange requesting is a feature of the HTTP protocol used to access information more efficiently from web servers. Instead of a server sending an entire file to a browser, the range request feature allows the browser (client) to define a specific part of the data that the user is interested in. Subsequently, only the requested part is sent by the server.
For cloud-optimized point clouds, the server will simply go through the hierarchy and find the cube or multiple cubes that satisfy the request, and return these. This process is much faster than having to search through the entire dataset. Moreover, only the point cloud cubes that satisfy the browser request are sent by the server, which reduces the bandwidth used.
Converting LAS Files to COPCLAS is a standard open source format for point cloud data interchange. However, it is less efficient to work with on the cloud since the data in it is not indexed. This means the whole LAS file must be loaded before it can be queried, sidestepping the efficiency we see with COPC indexing.
Conversion of LAS to COPC can be done in QGIS using Entwine. Practically, when a point cloud file is loaded to QGIS, it is automatically converted to COPC. QGIS structures the data automatically in order to make operations more efficient as opposed to working with unorganized datasets.
When point cloud files are converted to COPC in QGIS, the new COPC file contains all the information in the original dataset. Unlike other software that may discard some information when processing a file, for COPC nothing is discarded. This makes the COPC format great not only for visualization purposes, but for analysis as well.
What Infrastructure is Required to Serve COPC Data?Serving cloud-optimized point clouds does not require any special infrastructure between the server and the client. It is easy to host and get the data to the client without complex infrastructure, i.e. there is no need for something like GeoServer, MapServer, or QGIS. Just having the COPC data in blob storage somewhere is all the infrastructure that may be needed.
Compatibility of COPC FormatA compressed LAS file is called a LAZ file. COPC is much like a LAZ file. This means that applications that accept these formats will also be able to work with COPC files without having to implement special support. The only difference is that they will not be able to use the extra features of internal indexing in the COPC file.
Where Can You View COPC data?
QGIS offers support for viewing COPC – both stored locally on your device, or remotely in the cloud. Using a link that points to a server containing COPC data, QGIS will load the data on demand according to the queried range. The data is further cached in QGIS, which makes subsequent data loads and views much faster.
There are also a couple of projects coming to life that explicitly support cloud-optimized point clouds. An example is the web viewer built by Hobu. With the link to a COPC data server, the web viewer will fetch the relevant COPC files and render them in the browser.
The PDAL LibraryPDAL (Point Data Abstraction Library) is a library that contains a set of tools for working with point cloud data. In the QGIS environment, many users are familiar with PDAL’s feature for the data access of point clouds. Many of the library’s other functionalities are unknown to a lot of users. It contains a dozen features to classify, filter, export, convert point cloud to raster or meshes, amongst others.
The main reason why many functionalities in PDAL are not popular among ordinary users is due to the complexity in using it. The PDAL library uses pipelines that need to be crafted manually when working with point cloud data. While this may work well for advanced users, ordinary users find it a bit too complicated.
After a successful crowdfunding campaign, Lutra Consulting and several other partners are working to reduce this complexity, and make the functionalities in PDAL more user friendly. The project seeks to build a simple integrated toolbox within QGIS for point cloud data processing. The same way there are integrated toolboxes in QGIS for working with vector or raster data, there will also be one for working with point cloud data. These developments may be expected across the next two QGIS releases; in February and June of 2023.
What is the STAC Protocol?STAC (SpatioTemporal Asset Catalog) is a protocol for easy access to spatial and temporal data. It makes it easier to index, discover, and work with geospatial information. STAC is commonly used with satellite imagery but recently it is increasingly being used for distribution of point clouds as well.
How Were Point Clouds Streamed Before COPC?No doubt, before COPC there were some existing formats for streaming point cloud data. In the open source world, one of them is the EPT format, built by Hobu Inc.. The EPT format closely compares to raster tiles; but for point clouds. It is structured in a big directory with individual files (tiles). COPC files have an advantage over EPT format, as opposed to having thousands or even millions of files in a folder structure, COPC is just a single file – which is much easier to work with. In the proprietary world there are a couple of formats as well, one being the I3S format from ESRI, which supports point cloud data and other 3D data. There is no doubt we will continue to see explosive growth supporting point cloud data management. Stay tuned with the MapScaping Podcast to make sure you stay current on the latest and greatest developments!
PDAL - Point Data Abstraction Library
Full Stack Cartography – Think Like a Designer for Better MapsThis episode features Alan McConchie, a leading cartographer at Stamen. The discussion centers on mapping and visualization; and explores the usefulness of adopting a design-first thinking approach to the map making process. Alan has over 8 years experience making maps that are as beautiful as they are useful. Together, we will gain a deeper insight into full stack cartography, and how we can leverage it in our own work.
Spatial Data Visualization vs. Non-Spatial Data VisualizationAlthough they are both ultimately design tasks, there are some distinct differences between spatial and non-spatial data visualization. For one, there is lesser liberty in spatial visualizations than there is in non-spatial data visualization. Unlike in the latter where the designer can take liberty in choosing where to place objects and what shapes to use, spatial visualizations usually have to maintain location integrity. If the locations in a map are distorted, the map will be misleading and fail to serve its purpose. The flexibility in maps is mainly found in the choice of the symbology, color ramps, and fonts used. Even these areas are often guided by standard cartography rules, so it's important to make the most of the creative space provided if you want your maps to stand out from the rest.
What is Full Stack Cartography?Full stack cartography is seeing the world as a challenge of design, and how it can be effectively mapped for a particular use case, or to communicate your topic of interest. To back it up a bit, full stack is the idea of taking your entire workflow, and all of the areas of your system that it touches into consideration together. Full stack cartography incorporates design thinking into the map making process right from the beginning, throughout the process, and until the end. Writing the code and scripts used to process the map data is heavily guided by the use case of the end product.
Ideally, every decision, at each stage of the process should contribute towards making the end product the best it can be for its use case. All the decisions made when making a map are important; from how the data is stored in a database, to how it is processed, to how the end product is displayed. The ripple effect of all these decisions throughout the process can make achieving the end goal easier (if done right), where ignoring them can make the process more difficult, inefficient, and altogether frustrating.
Why Are Base Maps Pre-Generated? A base map is a pre-generated background map that provides context to other map layers that are overlaid on top of it. They generally contain generalized land uses, streets, water features, and building footprints, amongst other landmarks. Pre-generating base maps makes loading web maps faster, and more efficient, by leveraging imagery tiles. Since creating base maps is a complicated process that usually requires a huge amount of data, generating them on the go every time is computationally intensive and slow for devices with low to standard computing power – which is the case for the majority of users.
In order to improve the user experience even further, base maps are divided into many small pieces, called map tiles, so that a browser does not have to load the entire base map, but rather only loads the pieces (tiles) of the area that a viewer is interested in. This is faster and more efficient than loading the full dataset.
How Do Map Tiles Work?Using map tiles is a popular and common technique for displaying maps on the web. After generating a map, it is divided into many small pieces (usually square shaped, be sure to look into Hex Tiles); which, when joined, form a seamless single map.
Tiling maps improves the user experience by reducing the size of the dataset that a browser has to load when displaying a map. Less data to load means displaying a map is faster, and uses less bandwidth. This is significant as web maps are already data and display intensive due to the many additional features and layers that need to be loaded alongside the base map, especially if those layers contain a lot of attribute data. To view an area of a map, the browser will only load the tiles that cover that area. When a user pans the map, more tiles that cover that area are returned. Previously loaded tiles may be cached so returning to your previous extent will be faster than if the map was loaded fresh.
How Are Map Tiles Created?After creating a map, tiles can be generated using map tiling tools or services such as Mapbox Tiling Service. Tiles can be in a vector or raster format. The tiles are generated for different zoom levels to allow users to zoom in or out of a map as they may need. At higher zoom levels, the map tiles are further split into smaller equal tiles that display more details. Each zoom level has its pre-generated tiles that are easily loaded to a browser upon request. Most programs, like ArcGIS, have a default tiling scheme and thus create tiles for specific zoom levels that are most optimized for their viewers. If, however, you want your map to be drawn at different intervals, you can create and publish a custom tiling scheme.
Geometric SimplificationCreating map tiles includes deciding which details will be included in the tiles, and which will be filtered out. In order to improve efficiency even further, cartographers try to reduce the size of data stored in each tile as much as possible. Some simplification is applied in the lower zoom levels to exclude details that may not be visible or useful in that zoom level. For instance, an application to show bike trails in a park may only display the trail lines at lower zoom levels, but then include the attributes like the trail names, or topographic lines at higher zoom levels. A cartographer decides which details to include at each zoom level and cuts out what is not required to avoid storing unnecessary data in the tiles. Having light weight tiles makes them even more efficient, fast, and useful for many mobile applications.
Styling Map TilesAfter deciding what map details appear at each zoom level, it’s time to decide how they should look. The tiles only contain the shapes and attributes of the map features. In order to apply styles such as colors and textures, a stylesheet is used. The stylesheet defines how the data stored in the tiles will be displayed. Design tools such as Mapbox Studio, or the open source option Maputnik, can change the color, line thickness, and other styling characteristics of a feature and display them instantaneously for review.
What is a “Good Map” in Cartography?A good map may mean different things in different use cases, and to different people. Generally, a good map is one that finds the right balance between beauty and usefulness. No doubt we all love a beautiful map, but if it is too beautiful then users spend more time admiring it than actually using it. This is not a good map as it does not successfully and efficiently convey its message. Keeping the distractions to a minimum is one important trait of a good map.
Design-first thinking helps to create good maps that are useful, as well as visually appealing. Everything from the choice of colors, to fonts, to avoiding emphasis on less important features will contribute to whether a map is successful in its application. The ultimate measure of a good map is how well the map serves its intent. This of course includes making it visually appealing so that people are willing and able to use it in the first place.
Is Cartography Changing?In the world of cartography, a lot has changed already. The tools used to make maps decades ago are not the ones being used today. The way maps are displayed today is different from how they were displayed in the past as well. The way maps are used have evolved too, and this is why the approach to cartography has also evolved over the years.
Today, maps are not only created for humans but for machines as well. For instance, self-driving cars need maps for navigation, and have to process a lot of simultaneous inputs.
In the future, more changes can be expected as further advancements are realized. Technology is getting better by day and soon we may see even more new products in cartography. For one, since personalization is a big thing these days, we may start seeing more personalized maps. This will come hand in hand with the loss of privacy that seems commonplace, and even expected nowadays. The developments seen in indoor mapping suggest we will soon see seamless outdoor-to-indoor navigation.
There is much we can expect from the future. In order to avoid being surprised, get involved now.
Recommended Podcast Episodes about Cartography
Artificial Intelligence In Simpler TermsThe guest on this episode is Daniel Whitenack. He is a data scientist at SIL International, a teacher of AI, and a co-host of the Practical AI podcast. With his background in computational physics and his day to day role as a data scientist, Daniel has developed expert-level modelling and math skills. Currently, he is working on AI technology that benefits local language communities (i.e. speech recognition for local languages, machine translation, and different natural language processing techniques). On today’s show, he helps to demystify questions around Artificial Intelligence (AI), machine learning, and deep learning.
What Is AI?Put simply, Artificial Intelligence can be thought of as a function of a software. In programming, a function is the logic or algorithm that performs a certain transformation on an input, and gives a certain output. Usually, a function is created by a developer who curates the logic associated with that function. The developer specifies exactly what happens to the input, how it is transformed, and how the output is provided. In the case of AI, instead of a human specifying all the logic of a function, they create a function whose internal parameters are not set. These parameters are left to be set by the computer itself later on. This is the main distinction of AI functions from other conventional functions in software engineering, the computer is trained to select its own parameters to achieve a goal.
What Is Algorithm Training?When an AI model is created, its functions are un-parameterized. This means that although the structure for handling the desired data transformation is defined, the parameters that should be used are not specified. The computer has to learn and figure out the right parameters that will transform the input into the desired output so that it can fill them out itself for future runs.
In algorithm training, the computer learns to set the optimal parameters for an AI function through a trial and error cycle. The training data provided to the AI algorithms includes samples of the expected input to the model, as well as samples of the outcome that the model is expected to produce. Using the sample data, the computer sets the parameters for the algorithm. Tests are conducted to see which parameters produce the best results in comparison to the sample output, and the model’s settings are refined through an iterative process until this is achieved.
What Are Transferable Models?In the world of AI models, transferability means that an existing model that solves a certain problem, and can be used to solve another different, but very similar problem. Tweaking the parameters a little bit can make the model useful to another use case. For instance, by just tweaking the parameters of a model trained to recognize dogs in images, it could be used to recognize cats in images. Transferable models are more computationally favorable since training does not start from scratch, but instead makes use of the knowledge and efforts already developed in the parameter and training sample set.
What Are Model Architectures?An architecture of models is a configuration of a neural network that is composed of different layers bolted together. Each layer in the neural network is better than its counterparts in processing particular types of data. Model architectures are useful in applications like Natural Language Processing (NLP) which often uses recurrent layers to process the sequences of text. In NLP, text is treated as a sequence of words and characters. The neural networks process the order in the sequence and the relationship between parts of the sequence in order to pull meaning from the individual words to create a sentence, and impose context.
When Should AI Be Used?There are typically two instances that would tell whether AI is beneficial for a certain application. The first is if the data transformation that needs to happen is so complicated that a human cannot seem to be able to do it reliably every time. An example of this is detecting mental health issues from a person’s voice. While it is nearly impossible for a human to do this due to the very complicated data transformation that needs to happen, an AI model would do it very efficiently.
The second is the scale of the task. When the scale of the task is too large and repetitive, a human may not be the most efficient option for doing it. For instance, it would be tedious for a human to have to classify two million aerial images, breaking them down into different land use types. At such large scales, it would be best to automate with a model.
What is the Difference Between Machine Learning and Deep Learning?The main distinction between machine learning and deep learning is the scope at which the models are operating. Machine learning models operate within a smaller scope, and do not usually utilize neural network architectures (e.g. decision trees, random forests, naive Bayes, among many more). On the other hand, deep learning occurs at a larger scope, and requires less human intervention than machine learning.
Deep learning usually involves neural network architectures that can take in millions or even billions of parameters. Significantly more data is required to properly fit the parameters of functions in deep learning than what would be required in machine learning models. Similarly, more computing power is needed in deep learning than in machine learning.
Making the Choice Between Machine Learning and Deep LearningThere are some situations where deep learning, though applicable, may not be the best option for your use case. One is when we require clear interpretability of the decisions we make. Certain industries, i.e. in finance or health care, there could be a high burden or government regulation that requires you to be able to explain and audit the decisions that you are taking. This means clear documentation on every step taken in a process. In cases such as this, a simpler model may be more appropriate since it is possible to actually tell how a decision was made, you aren’t plugging everything into a black box.
The other element is the cost involved. In certain cases, a lot of specialized, expensive hardware may be required in a deep learning AI project which you may not already have access to. If a simpler machine learning model can be trained on a laptop to solve the same problem, then there shouldn’t be a need to spend thousands of dollars on specialized AI hardware clusters for deep learning.
Do You Need Specialized Hardware to Use AI?The phases of working with AI models can be broadly put into two categories – training and inferencing. The training side is often the more time and resource intensive side that would require specialized hardware, especially if it is deep learning. A lot of computations are required to set the parameters of a model and extra Graphical Processing Units (GPUs) may be necessary depending on the scale of the data.
Inferencing is when the model is making predictions after it has already been trained. Specialized hardware is not usually required, save only for certain use cases. In real time processing and data feeds, where instant results are required, specialized inference hardware might be required. Otherwise, many AI models can run on CPUs since inferencing is not as computationally intensive as training.
Demystifying AIFor quite some time, the AI industry has faced two extremes: a heightened hype of unwarranted expectations, and on the other end, a total shutdown on the applicability of AI in certain cases. This is well captured in the words of Daniel Whitenack;
“If you are on the side of things where you are thinking AI is the solution to every problem, then you are overestimating its utility and where it is at the moment. On the other hand, if you are running a business of any size, operating at least some type of technology, and you think AI cannot solve any of your problems, then you might be mistaken as well.”
This is a story about a peer-to-peer mapping technology that is enabling people to "fight maps with maps"
https://www.digital-democracy.org/
You can find Mapeo here
https://www.digital-democracy.org/mapeo/
Promoting OSM projects
OpenCage and Mapscaping are working together to help projects based on OpenStreetMap reach a wider audience.
Projects will be selected to be featured on upcoming Mapscaping podcast episodes, with all costs covered by OpenCage!
Apply here!
https://forms.reform.app/NAvQ41/opencage-mapscaping/af9MFV
Recommended Podcast episodes about Off-line mapping
MerginMaps - QGIS offline and in the feild
https://mapscaping.com/podcast/qgis-offline-and-in-the-field/
https://mapscaping.com/podcast/two-mobile-data-collection-apps-you-need-to-know-about/
More podcast episodes on GIS and GIS careers can be found on our website https://mapscaping.com/podcasts/
Consider supporting this podcast on Patreon
https://www.patreon.com/MapScaping?
Or go to MapScaping.com to find out about sponsoring our website
reach out on Twitter https://twitter.com/MapScaping
or LinkedIn https://www.linkedin.com/in/danielodonohue/
Applications of thermal imaging from space include monitoring wildfires, urban heat islands, economic activity, and the built environment.
But it's not easy ;)
Connect with Robin Cole at
https://robmarkcole.com/
Check out the Earth Observation Hub!
https://geoawesomeness.com/eo-hub/
More Geospatial Podcasts episodes
Recommended Podcast Earth Observation Podcasts
Fake Satellite Imagery
https://mapscaping.com/podcast/fake-satellite-imagery/
The LandSat Program
https://mapscaping.com/podcast/the-landsat-program/
How To Keep Your Satellite Pointing At Earth
https://mapscaping.com/podcast/the-landsat-program/
It turns out that monitoring atmospheric pollution from space is really hard! But if you can do it will help you understand air quality, solar energy, ozone Layer and UV radiation, emissions and surfaces fluxes, and climate forcing
Contact Mark on Twitter
https://twitter.com/m_parrington
Copernicus Atmospheric Monitoring Service
https://atmosphere.copernicus.eu/
https://twitter.com/CopernicusECMWF
Check out the Earth Observation Hub!
https://geoawesomeness.com/eo-hub/
More Geospatial Podcasts episodes
Recommended Podcast Earth Observation Podcasts
Fake Satellite Imagery
https://mapscaping.com/podcast/fake-satellite-imagery/
The LandSat Program
https://mapscaping.com/podcast/the-landsat-program/
How To Keep Your Satellite Pointing At Earth
https://mapscaping.com/podcast/the-landsat-program/
The problem of unification: Spatial data comes in many different sizes, shapes, and formats making it a difficult and time-consuming process to join data for visualization, exploration, and analysis.
Enter the Hex Tile system!
Contact Foursquare: connect.foursquare.com/mapscaping
Check out Unfolded https://foursquare.com/products/unfolded/ https://www.unfolded.ai/
Introducing Hex Tiles: https://foursquare.com/article/introducing-hex-tiles-our-next-gen-tiling-system/
Example Maps
https://studio.unfolded.ai/public/e86787a1-4871-4fbb-8eb8-33445e81ca73
https://studio.unfolded.ai/public/e3a6d2df-1784-4686-972b-9d3793c2515d
https://studio.unfolded.ai/public/beb3cf0d-211d-4e5c-b08d-663da9a0631d
Recommended Podcast Episodes
Dynamic Vector Tiles Straight From The Database
This is a story about a hobby project called Scribble Maps that grew into business… but it's also a story about the opportunities and challenges of creating geospatial tools for non-geospatial professionals.
Jonathan Wagner
“Its easy to get people to make a map today, but how do you get them to come back and make a map tomorrow?”
Special Offer for Listeners for the MapScaping Audience
ScibbleMaps.com/mapscaping
More podcast episodes on GIS and GIS careers can be found on our website https://mapscaping.com/podcasts/ Consider supporting this podcast on Patreon https://www.patreon.com/MapScaping?
Or go to MapScaping.com to find out about sponsoring our website
reach out on Twitter https://twitter.com/MapScaping
or LinkedIn https://www.linkedin.com/in/danielodonohue/
SAR Technology: Looking Beneath the Earth from Space
The guest on today’s show is Lauren Guy, the CTO and founder of Asterra. He has a background in Geophysics and discovered the technology on which Asterra is built while during his masters. Lauren was involved in a project that used radar sensors orbiting Mars to search for water on the planet. The Synthetic Aperture Radar (SAR) signals used penetrate the ground to give a clearer view of what lies beneath and whether there is any indication of the presence of water.
How Bad Are Water Leakages?
Most of our drinking water is transported from far-off locations like reservoirs or even desalination plants near the oceans.
About 30-40% of the water that is moved around the world is lost through leaky pipes.
It is not only water that is lost in leakages – a huge amount of energy (used to pump the water) is lost as well. Considering the growing impacts of climate change, and looming future water scarcity issues, this is incredibly significant.
How Does Synthetic Aperture Radar (SAR ) Find Water Leakages on Earth?
Different materials have varying dielectric constants and therefore, electrical conductivity. These properties cause materials to reflect SAR signals differently. Treating drinking water gives it a distinct salinity level from other makeups of water.
Since salinity affects conductivity, it reflects a distinct SAR signal that differentiates drinking water from other kinds of water sources in the ground.
The assumption made is that there is no other source of drinking water in the ground – it has to come from pipes. This makes it possible to use SAR to create an underground map illustrating leakages.
Even with the capability to accurately isolate drinking water from other kinds of water, there are still possibilities for false positives as the same water is used for a lot of different uses like watering lawns, gardens, or filling swimming pools. The isolation has to go a notch higher to be able to distinguish the drinking water that is coming from pipes.
One way to do this is by calibrating the algorithm to only show moisture in the ground that has been accumulating for more than 48 hours. The thinking behind this is that most people do not water their lawn for more than 48 hours at once. This helps to avoid wasting time on false triggers.
How Far Can SAR Penetrate into the Ground?
SAR can only penetrate a few meters into the ground, depending on the soil type, and the top covering (i.e. asphalt, pavement, etc.). Generally, the depth of SAR penetration is about 2m in cities, 5m in more rural locations, and up to 10m in very sandy soils. The penetration depth of SAR is suitable for this application since water pipes are usually laid at a depth of 1m.
Georeferencing SAR Images
Georeferencing is a critical part of working with SAR images. The images need to be georeferenced in order to figure out where they are on the surface. Finding the exact location where there is a leak as shown in a SAR image is very important to avoid sending a crew to the wrong place.
Since SAR sensors are usually pointed at Earth at an oblique angle, georeferencing SAR images can be difficult. Georeferencing algorithms have to undergo a robust training phase in order to achieve the level of accuracy required. As we continue to see the popularity of SAR technology grow, we may see this get easier.
Tackling Signal Noise in Urban Environments
Telecommunications in urban environments creates a lot of noise for SAR sensors as they use the same frequencies. Reflective surfaces also create a lot of noise. This problem can be overcome by using different polarizations. A SAR signal can be sent in three ways: Vertical, horizontal, or as an alternating combination of the two in some cases.
When the signal is bounces off different materials or noises, the polarised signal goes from vertical to horizontal and vice versa. Detecting these polarization changes, and measuring their magnitude makes it possible to identify the source of noise that caused the change of polarizations and correct for it.
Building a Business in Around SAR Tech
The problems that Asterra faced while building a business around SAR technology are quite the same for many companies that are developing new solutions in tech. Despite putting in a lot of effort to convince clients that the technology actually works, it is even more difficult to convince a client to create a new budget to buy that solution.
For Asterra, these were passive utility companies that were not actively looking for leakages in their infrastructure, but primarily relied on citizens to report suspected leakages. One of the more common ways a customer might notice a leak is a sudden spike or even gradual increase in their water bill, without having changed their habits. As utility companies did not have an existing budget for a field team that looked for leakages, it was difficult for them to justify a new expense.
Why You Should “Speak the Same Language” As Your Clients
The best kind of clients are the ones that are already solving a problem. It is easier to convince a client to buy your solution if you can acknowledge their existing ideas and solutions, and then offer your ‘new’ solution as a complement to theirs. It will not sound realistic to them to ask them to throw everything out – what is often very expensive equipment – and replace it with your solution. A general good tip in convincing anyone of something, is to make them feel like they came up with it themselves.
For Asterra, it was easier to convince utilities that already had a field team that were actively looking for leakages in their infrastructure, and already had a budget for it. The companies could just reappropriate the budget towards buying the new solution, and even save costs as Asterra’s solution is often cheaper than the original way.
There were also advanced utilities that not only have a field team, but also very expensive IoT equipment for finding leakages. Due to the high costs involved, it is not economical to install the equipment throughout an entire city. For these cases, Asterra’s solution could be adopted as a complement to these devices in order to identify the most problematic areas in the city, and shift the more costly equipment to those locations where it is needed most.
Embracing Competition in Tech Businesses
For most businesses, the thought of competitors may be dreadful, but in tech businesses, competition might just be the force that drives the business forward. Being the only player in the game may be viewed by clients and investors to mean that the market is not viable. It leaves them with many questions of why there are no other solutions in that area already if there is so much money to be made. Competition fuels more discussion around a technology, which is an opportunity for the newest solutions to gain exposure.
Other Applications of SAR Technology
SAR technology is also useful in monitoring other important infrastructure like highways and railways for potential issues. Since SAR is looking underground, the issues it identifies may not be apparent yet from ground level. Accumulated water causes most of the issues in infrastructure. Pinpointing locations with very high soil moisture can help railway companies, or a country’s Department of Transportation to identify infrastructure that may fail soon. SAR’s ability to penetrate the ground at night and in any kind of weather can be harnessed in many other applications, such as mineral explorations, defence, and identifying contaminated soils. Who will be the ones to make it happen?
Recommended Podcast episodes
How to Keep Your Satellite Pointing at Earth
The guest on this show is Jack Reed, a PhD student at the MIT Media Lab. He started out in mechanical engineering and then later on moved into aerospace engineering. At the MIT Media Lab, he is part of an interdisciplinary research group alongside lab mates with backgrounds in data science, ethics, and art. Together they work on making space sustainable, and using space-based assets and imagery to help promote sustainability on Earth.
Why do We Need to Orient Satellites?
We tend to assume that since space is a vacuum, then there is nothing that would make the satellite drift from its original course. If this is the case, then there should really be no need to worry about orientation since the satellite should keep pointing to the same direction.
In reality, satellites are affected by forces like atmospheric drag, especially in low Earth orbits. Atmospheric drag slows down satellites and pulls them out of orbit.
Another reason to orient satellites is that as they revolve around the Earth, at 180 degrees they would be pointing directly away from the Earth. Unless they are rotated it will have to go another 180 degrees before they point directly at earth again.
In order to keep the satellite pointing at Earth at all times, they need to be rotated constantly, otherwise, they lose half of their utility.
How Satellites Orientate Themselves in Space?
The ability to control a satellite’s position in three dimensions, and where it is pointing is critical to increasing its lifespan. Controlling it in three-dimensional space helps to keep the satellite in the correct orbit, while controlling where it is pointing to ensure that the satellite is capturing the right data, and sends communications back in the right direction.
If it goes off its orbit, or points in the wrong direction and is unable to get back, its lifespan would be cut short, and a huge amount of resources would be wasted.
Propulsion is used to move a satellite through three dimensions. It enables satellites to get back to orbit if they get off. Propulsion can be achieved by having rockets or thrusters at different corners of a spacecraft for turning. Large spacecraft also use propulsion to control where they are pointing. Ideally, engineers want to achieve propulsion with the least possible amount of fuel, as the weight of the fuel at launch can make the endeavour far more expensive and complicated.
When talking about flight, attitude is the information about an object’s orientation and position on its axis in relation to the plane below it. These properties are pitch, roll, and yaw.For attitude control, most satellites use reaction wheels. Speeding up or slowing down a reaction wheel on a spacecraft, will cause the spacecraft to rotate in the opposite direction. By having at least three wheels, a satellite can precisely orient itself towards wherever it needs to be. You could hypothetically call this an attitude adjustment.
How Do Satellites Determine Their Attitude?
In order to keep pointing at Earth, satellites need to know how they are oriented relative to the Earth’s location at any particular moment. A spacecraft subsystem, the ADCS (Attitude Determination and Control System) serves the purpose of maintaining a satellite's optimal orientation. It consists of IMUs (Inertial Measurement Units) and a variety of other sensors working together to tell where a satellite is pointing in relation to the Earth’s location.
Some of these sensors include:
Horizon Trackers
Since the Earth is warmer than space, infrared cameras can see the cut off of the earth’s horizon in space. In this way, a satellite is able to figure out where the Earth is. The downside though is that due to the horizon’s huge size, it does not give a precise point as to where all satellites will be pointing, producing a more generalized target direction.
Sun Trackers
The sun can be tracked through sensors that look for the hottest thing in the sky. Unlike the Earth’s horizon, which may appear as a huge circle of horizon, the sun tends to be in a very particular spot in the sky that the satellite can point at. By tracking the sun, a satellite can determine how they are oriented in space. During the times in its orbit when the Earth is between the sun and the satellite, sun tracking cannot be used.
Star Trackers
Star tracking is the most accurate method that satellites use to determine their attitude and figure out the location of the Earth. In fact, even the early astronauts using sextants, an instrument used for measuring the positions of stars since the 18th century. Star trackers compare the position of very particular stars against a star catalogue in order to tell the direction in which a satellite is pointing.
Magnetometers
Satellites on Earth orbits can use the Earth’s magnetic field to orient themselves. A magnetometer which detects the Earth’s magnetic field can tell a satellite which direction is north, and with this, a satellite can determine its orientation.
Earth based Ground Stations
Broadcasted signals from ground stations on Earth can be used by satellites to figure out where the Earth is relative to their current position. A major limitation to this method is that the signals’ strength fades out the further the satellite travels, and many ground stations are needed to keep continuous coverage and not lose data.
Can GNSS Satellites Be Used to Orientate Low Orbit Satellites?
GNSS constellations live about 20,000 km above Earth, and hypothetically they can potentially help orient low orbit satellites, which are usually below 2000 Km. While it is possible, there are certain caveats that limit using GNSS satellites from orienting lower satellites. Firstly, GNSS satellites are broadcasting specifically to Earth, so it is possible to receive a signal from at least four of them at any point on Earth, and triangulate an accurate position. This is how we use GNSS for earth navigation systems, and GPS. As you move higher, the field of view gets narrower and satellites which are very high up may not receive signals from enough GNSS satellites to triangulate an accurate position. Additionally, GNSS predominantly gets a satellite’s position, but not its attitude or orientation.
Want to learn more about using GNSS on earth? Listen here to learn about Where Does the Blue Dot Come From?
Edge Computing
Tones of high-resolution imagery is being captured by earth observation satellites every day, and edge computing has an important role to play here. It is not necessarily feasible to transmit everything that a satellite captures, since some of the data may not be usable. For instance, running algorithms on captured images out in space will help to sort out the ones that have too much cloud in them, and only downlink the images that would be usable. Algorithms can also complete some light processing and correction of the images, so the data is already in a more usable state before being sent down to a ground station on earth.
Space and Geospatial
Compared to previous generations, we are certainly living in a golden age of both geospatial data, and the space industry. There are a lot of commercial players that are designing, building, and launching new Earth observation satellites; and many others who are figuring out new and innovative ways of processing that data and turning it into useful products for various applications. Earth observation satellites are surely becoming an important pillar for many geospatial applications.
Want to learn more about the growth of earth observation, remote sensing, and imagery? Listen to our podcast with Dr. Aliastair Graham
Recommended Podcast episodes
This is a story about bathymetric Lidar... and how geo-tagged sharks led to the discovery of a huge nature-based carbon sink in the Bahamas.
More on this case study here: https://r-evolution.com/r-initiatives/oceans
The Ocean of Things
https://www.darpa.mil/program/ocean-of-things
Recommended Podcast episodes
Mapping The Ocean Floor
Mapping Oceans With Sound And Mapping The Sound In The Oceans
PDAL -Point Data Abstraction Library
Whitebox Tools Is The Backend To Many Frontends
consider supporting this podcast on Patreon
https://www.patreon.com/MapScaping?
Or go to MapScaping.com to find out about sponsoring our website
reach out on Twitter https://twitter.com/MapScaping
or LinkedIn https://www.linkedin.com/in/danielodonohue/
The Open Geospatial Consortium (OGC) is an international consortium of more than 500 businesses, government agencies, research organizations, and universities ... but why should you care?
Maybe because this community is defining the future of geospatial infrastructure across all verticals!
Maybe because they are determining a set of "the best practices" when it comes to solving big problems like climate change.
Or ... if you have a business ... maybe because joining the OGC membership might just be the greatest marketing hack in the geospatial industry ;)
Here is a list of current members https://www.ogc.org/ogc/members
Recommended Episodes
Open Geospatial Standards – shared standards to solve shared problems
https://mapscaping.com/podcast/open-geospatial-standards-shared-standards-to-solve-shared-problems/
Cloud Native Geospatial
https://mapscaping.com/podcast/cloud-native-geospatial/
Building a Business from an Open Source Project: A Case of White Box Tools The guest on the show is John Lindsay, the creator of Whitebox Tools and a professor of Geomatics at the University of Guelph. John introduced us to Whitebox Tools in a previous episode and described it as a collection of open source geospatial analysis tools that can be embedded in other applications. In this episode, John is sharing with us how they are building ways to generate revenue from Whitebox Tools, while maintaining its status as an open source project - yes, this is possible! When is a Business the Right Solution for an Open Source Project? Being a developer for an open source project could be an experience of two extremes. On one hand, while it may be immensely rewarding to encounter day to day users of the software you have developed, but the experience could be quite stressful as well, especially as the user community expands.
Usually, open source developers have their day job, and only spare their personal time to work on the open source project. When the user community grows, the demands of this larger community may take a toll on the time of the developer. It may be a daunting and potentially draining task to have to respond to too many requests from users. This could lead to developer burnout, which could lead to a project dying out.
A growing user community is a good thing, but for the developer this growing demand can be taxing. At some point, it may be difficult for the developer to provide a satisfactory level of support to the user community as it scales. This means that the project would need more hands to support its userbase.
Generating revenue for an open source project builds more stability for the project, and means that it might be there in the future, with the right support and resources.
If you are having doubts about the idea of building a business from an open source project, then think of the 1000s of users of open source software who not only find them useful for education and research, but also in generating revenue for companies. If there are companies that can use open source tools to generate revenue, then it makes sense that the developer of the project should be able to leverage it in generating revenue as well. The Process of Building a Business from an Open Source Project Building a business from an open source project is not very different from how enterprises build their businesses. The only may be the challenge of preserving the integrity of the project as strictly open source. Outside of this, the business world presents a lot of the same challenges. Finding a Co-Founder Doing business alone could be really lonely, especially for a developer who is already overwhelmed by the demands of the project and community. It may be useful to find a co-founder who can help with the business side of things. At this point, it is important to know what your limitations are, and use them as a guide for the things you should be looking for in a co-founder. The right partnership is one in which the strengths of the partners fill in on the weaknesses of the other and contributes to achieving the goals of the company. Laying Out the Business Foundation In exploring the ways to generate revenue for the project, there are several things that should are imperative for the project’s success: maintaining its integrity and defining its success Keeping the Integrity of Your Goals A growing user community may be a good indication that the original goals of the project are serving the users well. When developing a business model for the project, it is critical to prioritize maintaining the integrity of those goals as much as possible, especially in the regards of it being an open source project. If the user community feels short changed, the project users may shrink and the possibility of the business being a success will diminish. It is very important to uphold the goals that keep the project intact, even when setting new goals. Defining Your Own Success For businesses built from open source projects, success may not mean the same thing as those built in commercial enterprise. Founders can have their own definition of what success is for the business. Sometimes this could mean the ability to sustain a workforce that supports the user community, and keeps the project alive.
Whitebox Tools are exploring ways to generate revenue that will be useful in keeping the project alive in the long term with adequate user support, learning materials, and improving and developing new tools. To them, success is being able to maintain and develop the Whitebox toolsets, and support the community that has gathered around it. This has helped Whitebox keep its focus on the user community, and maintain much of its original framework of the project. Ways to Earn Revenue from an Open Source Project One way to generate revenue from an open source project is by creating new paid tools. These paid tools could be those that extend the functionality of the free tools as in the case of Whitebox Tools Extension. In this way, the status of the free tools would still be maintained as open source as revenue is generated from the paid extension. In order for this to be successful, the paid tools should be compelling. Creating Compelling Products A compelling product is one that users are interested in. It may be a challenging task to know what tools the people want, or whether they will be willing to pay for them. There are some things that can help guide the creation of paid tools. One is looking at various thematic areas that your tools are geared towards and figuring out some tools that would have a great interest in those application areas. This will give an idea of where to focus development efforts. If you can create a toolset that fixes all the problems a certain type of user has, you become essential to that user.
Engaging with existing users of your product can also be a great resource for knowing where to focus development efforts. There is much that can be extracted from the conversations where users describe how they use a product, and how they would like to use it.
Gaps in a particular field can also provide an opportunity for developing a paid tool. A good place to look is in areas of recent advancements. Spotting a gap will enable you to create a solution that will be used by people working in a particular application area, and its more likely you’ll beat someone else to it. The “Pay as Much as You Can” Pricing Model “Pay as much as you can” is a pricing model that lets people donate to free tools. The Whitebox Tools website has this pricing model in the form of buttons displaying amounts between $0 and $100 on their tools download page. This presents an opportunity for users to support the project, even if they may not be able to afford the paid tools. It may also help to put in a note explaining how the donations will be used in support of the project. Steering the Business towards Success Thinking of the right tools to create, and having the right pricing model is not all that there is to succeed in this business. Increasing Product Visibility In marketing, the ideas that spread are the ones that win. It’s hard to achieve success if the project you are working on is invisible. It is unexpected for people to care about things they do not know exist or that they cannot find. This is an especially big challenge for products that are predominantly backend.
Usually, it is the frontends that get more attention than the backends that power them. Opting to create your own frontend may be a way of making your project more visible, but if your solution is designed to power other applications, this could be the wrong use of time. Alternatively, entering strategic partnerships with essential popular frontends may be useful in growing your appeal.
Cultivating various types of non-financial support may also help increase user awareness. For instance, reaching out to long term users and asking them if they could support you by mentioning your business on their website, blog, or social medias. You could also simply ask for permission to showcased them as users of your tools on your business website. Effective Communication Communicating effectively is a key part of the success of any business. Although you may create a really wonderful tool that provides an innovative solution to a problem you know people have, if it is not communicated properly, it still may not be successful. Effective communication means using the right words and images, or packaging them in a other way that helps people to understand what the tools can do for them. Try different communication styles gracefully, while acknowledging that many businesses find it more difficult to communicate a tool to users than creating a tool itself. Commitment and Confidence Commitment is super important. If you put something out in the world and it fails, have the confidence to try again. Repackage it, find other creative ways to communicate it and try once more. If you are using avenues like Twitter, LinkedIn, or even a blog or newsletters, it will help to think that you're speaking to different, if not overlapping groups of people in each of these avenues. It is important to reach out as inclusively as possible. The Website A website can be quite central in telling the story about a business. It is the first place where people from other avenues and social media will learn about you. A lot of effort should be put into the website to make sure it effectively communicates what the business offers. Imagine that people know absolutely nothing about your business and be as specific as possible about all the capabilities offered. How a User Community Kill, or Keep an Open Source Project Alive? Without doubt, a developer has a certain responsibility that comes with creating a software and making it available for people to use. After all, a project lives and dies by its user community. Its users, whether in an educational, government, or commercial setting, have a certain responsibility to ensure that the project thrives. If the developer is responsible for everything, it can result in burnout and risk the continuity of the project in the long term. Thriving open source projects are the ones which gain the support user community. Building a Business Is an Iterative Process One thing to always have in mind when building a business is that success does not usually happen the first time. If it does not turn out well the first time, try it again with a different approach and an open mind. Iterate and improve upon your process and eventually it gets better. There may be many struggles along the way but part of the journey is learning from mistakes.
WhiteboxTools – A Toolset for Geospatial Analysis
Our guest on the show today is John Lindsay, the creator of Whitebox Tools, a toolset which contains over 500 geospatial analysis tools. John is a professor of Geomatics who got into geospatial in the early 2000s. Many of the tools in Whitebox have novel functionality that are not found in any other software and are free to use.
An Overview of WhiteboxTools
Whitebox Tools is a package for geospatial analysis that can be embedded into other software in order to facilitate building other types of applications. The journey of creating WhiteboxTools began with the development of a Terrain Analysis System (TAS). This was a full-blown desktop GIS that was fairly widely used at the time of its creation. Following further developments and tools additions, Whitebox GAT (Geospatial Analysis Tools) was created, from which WhiteboxTools emerged.
WhiteboxTools evolved from Whitebox GAT after a shift in focus from the frontend to the backend, with the aim of creating a portable toolset that could easily be embedded in other software applications.
This includes desktop GIS like QGIS and ArcGIS; as well as scripting languages like Python, R and NEM, which enable users to integrate any tool in Whitebox into their larger workflows involving other geospatial packages. Whitebox is also embedded in other geospatial applications, such as the open LIDAR toolbox, geemap, and Leafmap, and many other applications that harness the analytical power of Whitebox Tools in their analyses.
WhiteboxTools Is Open
White box is an open source software with open access. Here, the notion of open access means to remove barriers for the users of the software so that they may examine, inspect, and understand the source code itself. Users can view the source code of any of the open core tools in Whitebox, using any frontend that has the functionality to view code.
Users can also port the toolset and translate the tools into another programming language that they are more comfortable with.
Whitebox Tools Is a Standalone Software
No prerequisites are required to be able to use WhiteboxTools. Whitebox is designed to have zero dependencies. It is a standalone executable that users can download and use in their platforms.
How Many Tools Are in Whitebox Tools?
WhiteboxTools is comprised of two categories of tools: The WhiteboxTools Open Core, and the Whitebox Extensions that extend the functionality of the open core tools.
There are about 465 tools in the WhiteboxTools Open Core, and approximately 63 tools in the Whitebox Extensions. All these tools have been developed in the RUST language. RUST is a low-level programming language that was picked by Whitebox with the aim of making the tools as fast as possible, while using the least amount of memory possible.
What Can You Use Whitebox For?
Whitebox contains a broad set of tools that are useful across a variety of applications. Generally, these tools are for use in solving spatial problems. Some of the tools are intended specifically for certain applications, for example, LiDAR Toolbox contains a collection of tools which can take a typical LIDAR dataset from raw point clouds all the way to the end product needed, such as an interpolated digital elevation model (DEM).
Some geospatial applications that make use of WhiteboxTools include: Wetland mapping projects, Geological resource inventorying, Digital archaeology, Soil mapping applications, Landslide modeling, Geomorphological modeling, Forestry, Spatial hydrology, and Modelling.
Where Do the Tools in Whitebox Come From?
Most of the tools in Whitebox come from research and experience. Whenever a cutting-edge tool for a certain type of process is developed during research, Whitebox becomes the platform to disseminate the information about that innovative tool. Some tools are also developed from user requests.
If there are multiple users requesting the same feature, then Whitebox will make it a priority. At the same time, if it's just a single request, but it's a really interesting spatial problem that had not been thought about before, a unique solution can be built then it may result in a tool as well.
A number of tools are also a result of helping people solve problems they encountered. Whenever a person has a geospatial problem and emails Whitebox to see if there could be something that can be done about it, if a tool is built that solves their problem, it is published on Whitebox in the open core where everyone else can use.
This is one of the reasons why there are a number of very niche-specific tools in Whitebox, which solve a very specific need. They are usually inspired by a very specific need, by a certain user.
Environments to Use WhiteboxTools
WhiteboxTools can be used in desktops as well as high computing environments like supercomputers, web-based environments, and even mobile environments. Although it is currently not being used in mobile environments, there is no reason it couldn’t be.
At the moment, mobile GIS is still much more focused on front end data visualisation for people in the field than it is for raw data analysis, but we could easily see this change in the future.
How Should WhiteboxTools Be Used?
Users may use WhiteboxTools in whatever capacity they can envision. If a tool is intended to perform a certain task in a particular application, but its capabilities may be applied in a creative way in very different scenario, then users are encouraged to explore that capability.
Whitebox fully supports the idea of users using their capacity to apply WhiteboxTools in different ways, as evidenced by their responsiveness to user input.
Turning Science into Practice
Whitebox is a platform that communicates usable implementations of cutting-edge geospatial solutions to geospatial practitioners. In many areas of research, the ideas and solutions developed are usually published in journals within academia.
While this is fairly accessible, many practitioners may still not have access to it, or may have access to it but not to the software artifacts that implement the solutions. WhiteboxTools bridges this gap and brings science (ideas and solutions) into practice by making usable implementations accessible to the whole community, rather than just the academic community.
Does Whitebox Accept Tool Contributions from Users?
The concept of Whitebox was developed with the goal of allowing user contributions to the toolset. One way is by suggesting ways to improve and optimise the existing tools.
This is why the source code is open access. The other way is by developing a new tool and presenting it to be published in WhiteboxTools. The new tool should be developed in RUST language, and will of course be subject to review.
Who Is WhiteboxTools Designed For?
In many fields today, people are leaning towards niche solutions. If something is for everyone, then it is also for no one since it does not differentiate itself in the market.
WhiteboxTools brings a new twist to this. It has a broad set of tools cutting across different applications, but what stands out is that these tools are very niche specific. Many of these tools are developed to perform a certain task in a particular application – novel tools that provide unique solutions that cannot be found elsewhere. Whitebox tools are not niche in terms of their applications but in terms of their solution.
Recommended Podcast episodes
Mergin Maps – Offline Data Collection for QGIS Projects
The guest for this episode is Peter Petrik, the CPO of Lutra Consulting. He previously worked in the automotive industry before transitioning to geospatial about eight years ago. Ever since joining Lutra Consulting, he has made use of his programming knowledge and skills, while collaborating with others to develop products and contributions to the open geospatial community. One of their recent developments is Mergin Maps, an open-source software that allows users to take QGIS to the field.
What Is Mergin Maps?
Mergin Maps is an integration that allows users to collect geodata using their mobile devices or tablets. Teams working together can use Mergin Maps to synchronize their data seamlessly to the cloud, or a server. The QGIS plugin enables users to easily access their data in the cloud, and utilize QGIS to post-process or analyse their project.
How Mergin Maps Works
In order to use Mergin Maps to conduct field surveys for a QGIS project, the first step is installing the Mergin Maps Plugin for QGIS. The plugin will ‘package’ your whole project into a folder, and make all of its layers available for use offline. If the project uses an online map, QGIS processing tools can be used to generate its layers for offline use. The plugin also runs several validation checks to ensure that the tool will work correctly on a mobile device. Once everything is set up, a single click on the synchronise button will take the project to the cloud. The project can then be downloaded to a mobile device, and used offline within the Mergin Maps Input app.
Every time a user clicks on the sync button in the app, any new changes are pushed to the cloud, and a new version of the project containing those changes is created in the cloud. It is easy to track the changes made to the project and identify the user that made them.
Mergin Maps also has an automatic merging feature. Users working offline on the same project, will have their data automatically merged by Mergin Maps once they sync it to the cloud. This removes the need to do it manually, and can save a lot of time if there are many users in the field. Mergin Maps does not have restrictions on the number of users that can be added to a project.
Data Storage When Using Mergin Maps
Mergin Maps is open source, so users are not restricted to using just its cloud. Users may opt to deploy their own servers to store their data.
The Mergin Maps Cloud is an option for users who do not wish to have the weight of server infrastructure and maintenance on their shoulders. There are various Mergin Maps subscriptions for users depending on their storage needs. For academia and personal use, Mergin Maps is free to use for up to 100 MB of storage.
Mergin Maps has features that help users reduce the size of their data, and cut down their storage needs. For example, there is a QGIS tool where users can set the resolution of all the photos captured in the i.e. 1 MB via the app. This helps to avoid storing high resolution photos whem they are not necessary.
Mergin Maps User Management
Mergin Maps allows a project owner to manage the users that can participate in the project. A user can be added to a Mergin Maps project as a reader, writer, or administrator.
Readers can only access the project and view the data, while writers can create new versions of the project, either in the mobile application, or from QGIS. Administrators can invite other users, as well as change user permissions.
Each user will need to create a Mergin Maps account before they can be added to a project. Once they have an account, they can use the same log in credentials to sign into the mobile application and, download the project from the cloud onto their device. They can then start working in the project (i.e. digitise, edit data, syncing etc.)
What Expertise Do You Need to Use Mergin Maps?
For the Mergin Maps QGIS plugin, some knowledge of GIS is required to be able to set up a project and push it to the cloud, and later access the data and do analyses and generate reports.
No GIS knowledge is required for users of the Mergin Maps Input app. The developers acknowledge that the users of the app are likely from different fields, with varying levels of knowledge. Therefore, they kept the design simple and easy to use. On average, a new user can learn how to use the Mergin Maps Input app in about 15 minutes!
Mergin Maps is flexible and has plenty of users from various industries and applications including agriculture, utilities and fibre optic companies, consulting firms, and many more.
Mergin Maps Integration with External GPS Tools
If the location from a mobile device is not very accurate, a user can connect to an external GPS. Mergin Maps supports most of the common external GPS tools which a user can plug in or simply connect to. They will then be able to use the external GPS’s precision for their data collection. Only Android is supported for this functionality at this time.
Controlling Data Inputs on Mergin Maps
Mergin Maps offers a number of widgets that can be used to control the values accepted by the Mergin Maps Input app.
For instance, if an input field is set to only accept integer values, an integer widget will appear on the data collector’s mobile device as a slider, or text box where they can only input a number.
Further limits can be set to define whether the number must be positive, or may not exceed a certain figure. For dates, a widget will pop up on the mobile device as a calendar. These functionalities are great for helping to manage data quality and data assurance standards.
There are also widgets for lists, checkboxes, and other common widgets and constraints. Users can also use the QGIS Expression Engine to write calculations for derived fields.
There are lots of possibilities supported on both the QGIS plugin, and Mergin Maps Input app, where users can apply some basic quality assurance methodologies, and help with capturing quality data in the desired format. After data has been pushed to the cloud, users can make use of the Python integration to write scripts in QGIS, or on their server to further validate the data from an input field.
Mergin Maps Community Edition
Mergin Maps are advocates for open-source GIS and that reflects in their products. Their users are not vendor locked within Mergin Maps infrastructure. If a user is tech-savvy enough, they can set up Mergin Maps, and use it on their own. There is a public GitHub repository with a Docker compose file that users can run for themselves for their personal projects. In case they encounter problems, they can join the Slack chat and seek help.
Mergin Maps Enterprise Edition
The Enterprise edition is best suited for companies that need to hold data in their servers due to legal or security reasons, or internal rules. Mergin Maps facilitates this by offering their knowledge from their years of experience running cloud services, to help in the deployment and maintenance of each company’s server, making sure everything works smoothly and efficiently.
Recommended Podcast Episodes
Joining us today on the podcast is Gregor Willhauck, the Geospatial Cloud Strategist at Trimble, to talk about digital twins and how they integrate with geospatial. Gregor’s background and degree is in forestry, but he has spent most of his career in product management positions, working largely with satellite imagery, GNSS systems, and now, digital twins and modeling. Focused on the big picture, he is always looking for the best ways to use the cloud to further enable digital twins and geospatial tech.
What Are Digital Twins?
With the rise of commercial and recreational drones and LiDAR scanners, digital modeling has become much more commonplace in GIS and other industries. This has unlocked all kinds of cool and useful applications of models and point clouds for both everyday, and scientific uses. Standard digital models are awesome, but digital twins have entered the space in response to a unique set of needs and spatial questions.
So, what exactly is a digital twin, how is it different from a regular model? Well, a digital twin is a virtual representation and digital counterpart of a physical asset, which is updated in real time. The key defining element to a digital twin is this real time aspect. A standard 3D model is captured once, and represents a single frame in time, where digital twins are the result of regular scans, and fuel and end to end data pipeline.
Defining what qualifies as “real time” can be a bit tricky, and largely is going to rely on what you are monitoring with the scans.
An active construction site for example may need to be modeled several times daily, even hourly, to realistically capture a crew’s progress. A completed construction project, however, may only need to be scanned daily to monitor just enough change to understand when maintenance is needed.
Both of these qualify as “real time”, but one is a bit more real than the other.
Digital twins are data dense, and can require a hefty amount of resources to create and maintain. This is why they pair so well with the cloud. Cloud computing and storage allows these models to be created at a frequency that matches the source image collection, resulting in minimal lag between when the UAV is on site, and when the finished product is on screen. This adds legitimacy to the final product, and allows decisions based on these products to be made with a higher level of confidence.
How Are Digital Twins Made?
The most straightforward answer here is that a digital twin can be created with pretty much any of the sensors and vehicles already used for remote sensing. Again, the wheel is not being reinvented here, it just has more RPMs.
Some of the most common data sources for fueling digital twins are drones, handheld scanners, and even robots.
These vehicles will generally carry cameras, or LiDAR scanners, and perform a flight plan for mapping, just as if they were creating a one time model.
The resulting images may be directly streamed to the cloud, where they will undergo processing before being uploaded to whatever end application is needed for the workflow.
Robots are exciting here because they can allow the entire digital twin workflow to be automated. After an initial human-led flight to build a map of the area, which is used to plot the robot’s course, the robot can then be scheduled to complete, and upload the collection to its overlords around the clock.
Photogrammetry is the primary tool for processing that creates these models.
Photogrammetry is when software identifies “tie points”, or common points between two or more images, used to “stitch” the images together into a continuous product.
Some clever geometry related to changes in where the sensor is in space, and the Pythagorean theorem, allow 3D reconstructions, and highly accurate remote measurements.
Digital twins can be supported by GNSS/GPS devices as well, especially to track movement within the context of the larger environment. GNSS devices are the easiest platform to monitor and mimic in “real time” because the data is produced quickly, and is easy to collect and transfer. Different data sources will have different levels of latency, and a digital twin can be composed of multiple data sources.
Uses and the Future of Digital Twins
Digital twins are a very cool, futuristic technology, but with all the required investment, what can be expected in return, and is it worth it? When it comes down to it, digital twins are a logistics tool. They allow an unprecedented level of monitoring, and enable highly valuable insights through visualization, and analysis of the vast data they are capable of producing.
A real-time geospatial data model itself is not the product, the information it contains is.
End users want to be able to extract features, like structural damage, or measure change, like on an active construction site.
Being able to do this remotely results in some financial and time savings associated with traveling, and allows stakeholders to catch issues as they occur, rather than at the next scheduled manual inspection.
Another major use of digital twins is measuring traffic or “flow” throughout an area. For example, if a digital twin is built of an office complex, then internet of things (IoT) devices like smart elevators, or even sliding doors, can paint a picture of congestion, or inactivity in certain areas at various times. Similarly, this technology can be used for fleet management to track where and when various vehicles are moving through space.
Digital twins of course serve as an excellent visualization tool. They are eye-catching and engaging, and can be accessed from nearly anywhere.
Virtual reality and mixed reality views can fully immerse a user in the space, opening up the option for interactive engagement, making markups, and viewing alternative setups or configurations in real time.
All of these activities ultimately serve to help promote collaboration and efficiency in an organization. Everyone is able to access the same information as someone on site, which we may see continue to grow in importance with the advent of work from home. As computing, storage, and UAV resources become more affordable and accessible, we will likely see new unprecedented uses for digital twins continue to emerge in both real, and digital life.
The Spatial internet of things - Sensor Up
Big Query
Google Earth Engine
The Planetary Computer
Augmented Reality
Our guest today is Jakub Dziwisz, the CEO and founder of Orbify, a web-based SaaS solution for earth observation applications. Although he does not come from a GIS background, Jakub has always enjoyed problem-solving through technology. Originally, this consisted of being a software engineer focused on expanding people’s horizons and perspectives through the travel industry. Now, Orbify helps expand the geospatial capabilities of everyday people, solving not-so everyday problems.
Why Merge GIS and SaaS? First of all, what is Software as a Service (SaaS)? Well,
SaaS products are traditionally web-based software that greatly reduce the amount of time it takes for customers to get up and running with their own products by utilizing the SaaS product’s infrastructure.
Generally, these products follow a ‘pay only for what you use’ format, very similar to how some cell phone providers offer custom mobile data plans.
SaaS products can be great for GIS because they leverage the SaaS provider’s hardware and software resources, allowing customer applications to live and run in the cloud. This lowers the barrier to entry in terms of resources for creating an application. Most SaaS products allow users to select their application’s function or design from a series of preset templates, or DIY in some other modular way. This also lowers the barrier by reducing the amount of time and expertise necessary to get started on a project.
Another huge time saver with SaaS products is that, due to being web-based, there is no installation needed. Organization administrators can configure all their user settings on the backend, and easily regulate access to their application. This makes it a very popular platform with large companies who have a huge and constantly changing employment pool.
Although the pay-as-you-go method has been popularized in cloud storage, and to a certain degree in the ArcGIS platform for some analysis using credits, SaaS providers that are creating infrastructure for GIS workflows are few and far between.
The Need for SaaS in EOS: An Analogy In the 1920s, Henry Ford introduced the world to the raw efficiency of the assembly line. Before this, the process of building a car was slow, and expensive. By introducing the assembly line, cars could now be built in a streamlined fashion which was faster and easier, ultimately making vehicles accessible to a much wider market.
In the 1990s, it was customary to go to a travel agency if you needed to go on a trip. The travel agent would list for you the often limited available flights and fares, and you could take it or leave it. With the introduction of budget travel services like Expedia and Hotwire, consumers were provided with options to modify their plans to their needs. This resulted in a more competitive airline market, and democratized the space by putting the consumer in control of what they consumed.
Finally, in 2007, Steve Jobs unveiled the iPhone. People had been familiar with the telephone, even the cell phone, for decades. Similarly, they had been familiar with MP3 players for years. The idea, however, of being able to have both of these things in one changed how people judged the capabilities of technology.
For decades, GIS users and earth scientists have been working across dozens of applications, tools, script, and data sources.
By bringing SaaS to this niche, geospatial data workflows can become faster, easier, and more accessible. As it succeeds and more players enter the space, consumers will find they have more competitive options, and that their needs and dollars are in control of what happens next, and they shouldn’t be afraid to ask for what they want.
Creating a SaaS User Experience The best SaaS products are intuitive, and feel almost effortless to navigate. This is in stark contrast to the alternatives- environment-specific scripts, custom desktop installations, and coordinating and managing server resources. Planning and designing a great SaaS product is heavily dependent on knowing what your clients need, and what their customers want.
In order to build an app, there are four main elements, data, functionality, the output, and the packaging/user interface.
Normally, it takes quite a lot of time, code, and infrastructure to build an application like this. It often even requires a full team of developers. Using SaaS software, something like this can be configured by only one developer, and relatively quickly.
Let’s take a look at an example, using Orbify as our SaaS platform.
Orbify was developed with low tech skilled end users in mind, who need to solve a spatial question on a regular basis.
The Orbify platform is a bit complicated for someone with little technology experience to configure, so the most likely situation is that a developer has been contracted to build this application. Our end user in this case is an ecology graduate student who needs an application to monitor shipping vessel activity to better understand how this overlaps with known whale migration paths.
A user-friendly experience is key to the speed here. This starts with the data aspect. An interface like Orbify comes packaged with access to a catalog of existing cloud data that the developer can connect to and use in their application.
As the source code is completely customizable, they have the flexibility to connect to their own data as well.
For our scenario, the developer is able to connect to hosted AIS ship tracking data, and import their own data for their whale migration paths.
The functionality needed for our application is to visualize the ship and whale data in the context of time. This is easily accomplished in Orbify by selecting the ship tracking application template, and enabling the time slider. The output for this project is a map, which is accomplished through the template. Finally, the user interface will largely be directed by the template, but the developer can select from a number of existing no-code buttons and tools to include in the interface, and customize the look of the application as desired.
In the end, our graduate student has a simple, custom application that was created with minimal man-hours, and will only be billed for what they use.
This scenario can be extended to all kinds of potential customers for whom the traditional custom application would be too expensive, farmers, foresters, firefighters, etc.
By reducing the barriers to entry for using custom GIS applications, more people can be exposed to the wonders of geospatial, and we can collectively work towards solving problems that weigh on everyone from the individual, to the global organization.
Our guest today is Ron Hagensieker Ph.D., the founder, and CEO of OSIR.IO, an artificial intelligence and earth observation company. Ron has had an interest in remote sensing in the beginning, starting with his BA in geography, all the way to his PhD in remote sensing. All along the way, he has found ways to incorporate machine learning. One project, in particular, was thiscitydoesnotexist.com, a site that uses AI to randomly generate a false Landsat-style image of, you guessed it, a city that does not actually exist. How is Fake Aerial Imagery Created? In order to create fake imagery, first, you must obtain a great deal of real imagery. These images will be used to train your neural network to recognize the unique landscape characteristics that make a city a city. The machine learning models learn to identify and recreate the patterns we associate with small-scale views of cities (i.e Landsat or Sentinel imagery) mimicking the transitions from city center, to suburbs, to farmlands. They can even be trained to create fake DEMs.
ThisCityDoesNotExist.com’s algorithm actually involves creating two competing machine learning algorithms, the generator, and the discriminator.
The generator algorithm generates images, attempting to fool the discriminator model which dubs each image as a ‘real’ or ‘fake’ city.
Based on the responses of the discriminator model, the generator will adapt its next image, and try again. In programming, this technique is called a Generative adversarial network (GAN). GANs are especially powerful algorithms, and can even be pit against other algorithms to expose their weaknesses.
The positive feedback loop the generation and discrimination process creates can hypothetically go on forever, but will ultimately be limited by computer resources, or be influenced by the finite number of images originally used to train the model.
These algorithms use unsupervised classification, which is when the computer groups like pixels into a prescribed number of classes.
Unsupervised classification distinctly does not require drawing bounding boxes around features, and therefore may be faster to prepare than supervised methods.
It is possible to add some influence to the process with supervised classification. This involves inputting a dataset of city center points, telling the model “here, these are cities”, potentially improving results. Why Create Fake Imagery? In a data rich world with so much access to all varieties of raster data, why would we want to add fake imagery into the mix? Well, the truth is we aren’t quite sure yet. When plenty of real data exists, there is little desire to inject energy into simulating data. While there are no especially compelling uses for large amounts of fake imagery, doctored imagery is a different story.
There may be military applications for faking parts of images, but not necessarily whole landscapes. This could be identifying and disguising sensitive military operations in satellite imagery, or maybe even simulating the result of damage to landscape elements.
Randomly generated AI landscapes don’t have much use on their own, but once you add more channels to the data (Elevation models, population density, etc), and start adding constraints, things can get really interesting.
The trick here is the channels must be coregistered to be meaningful. Coregistering means that the images have all been lined up with each other spatially.
By combining information from many channels, you can create a living modeling environment. It gets even more exciting if you can pull off adding time. This allows predictive simulation of urban planning scenarios, deforestation patterns, and even of events resulting from climate change. By simulating different population counts and densities, we can begin to identify common patterns in how people grow into their cities.
Fun Fact: If you drop oat flakes into slime mold at the positions of train stations in Tokyo, the resulting fungal network grows to look just like the actual Tokyo Rail System.
The reason landscape modeling does not qualify as a compelling application of fake imagery is that it is incredibly costly in terms of computing resources. Without adequate financial backing, it does not make sense to create a program like this. For now, coders are just having fun without funding. What Are the Risks of Fake Satellite Imagery? The most obvious risk associated with creation of imagery from machine learning models is the risk of people using this technology maliciously. Deep faking images may allow you to represent extreme scenarios on a landscape. One popular one is predicting the same landscape in all four seasons, but you could also simulate flooding, or wildfires.
Luckily there is not much genuine risk of misinformation (amongst free media at least) about events as these things are so easily verifiable.
If someone actually takes the time to simulate media coverage of a catastrophic event somewhere using deep-faked images, fact-checkers can verify from half a dozen public imagery sources if the event actually occurred (or even simply just call someone there).
So if you come across a fake image in the wild, how would you be able to tell it was fake? Well, one characteristic is that fake imagery often looks downscaled. Depending on the base logic in the algorithm used, sometimes you can pick out a fake image by the lack of truly straight road segments. This also applies to the lack of natural angles on building corners, or the presence of especially blurry building footprints.
Characteristics like haze or odd colors in areas are not reliable tells of fake imagery as these elements are normal in most real imagery as well due to atmospheric effects.
Overall, fake imagery has the potential to be very powerful, but in the current technical landscape, there is not much need for it. Due to the high skill, time, hardware, and financial investment necessary, it is unlikely we will see too much development growth in this area. The one thing we have learned from artificial intelligence, however, is to expect surprises.
Our guest on the show today is Chris Holmes, the Vice President of Product and Strategy at Planet. Chris entered the geospatial arena almost 20 years ago as an early contributor to the GeoServer project. In the beginning, he was writing code, but eventually discovered his talents were better spent helping to build a community and awareness around GeoServer and the greater Open Geospatial Consortium (OGC). Chasing innovation, he now works to promote wider spread adoption of cloud native geospatial solutions such as the SpatialTemporal Asset Catalog (STAC), and use of cloud optimized geotiffs (COGs) from his position at Planet.
The Basics of Cloud Native Geospatial As one of the newest technologies on the GIS scene, cloud native geospatial can seem a bit intimidating. Realistically, it is a lot of familiar industry staples repackaged to take advantage of huge advancements in computing technology.
At its core, cloud native geospatial (CNG) is your classic geospatial infrastructure, without all those pesky computational power and storage limitations.
By leveraging the power of AWS, BigQuery, Snowflake, Google Earth Engine, or ArcGIS Server, users can access and analyze global-scale datasets without needing to purchase and maintain the physical servers traditionally used in on-premise setups.
This shift reduces the barrier to entry for scientists, and even casual users, allowing more spatial questions to be asked and answered.
CNG systems are flexible and scalable to most needs. One can choose to host some, or all of their data in the cloud, then access it for local analysis, complete that analysis in the cloud, or take a hybrid approach.
GIS work has traditionally followed a desktop-centered workflow. Using cloud native geospatial, it does not matter if you access and analyze your cloud data through a browser, or desktop application, although each platform will come with its own natural limitations. GIS Data Formats and the Cloud Cloud technology alone has existed for a while, but took a bit of a learning curve to adopt into GIS due to the specific requirements and preferences of cloud infrastructure.
This means that some people took this rare opportunity to essentially start from scratch, and build data formats that are optimized specifically for cloud systems, but often still maintain the flexibility to be backwards compatible to desktop and enterprise workflows.
It is not possible to talk about CNG without talking about Cloud Optimized GeoTIFFs (COGs). COGs are the backbone of cloud native geospatial, and are essentially responsible for starting the GIS industry’s race to the cloud. The beauty of COGs is that they can be used as a regular GeoTIFF in a desktop setting, or leveraged in the cloud to unlock fantastic real-time data streaming and analysis.
The key element of COGs is their compatibility with range requests. The efficiency of range requests is what enables streaming in a lot of our favorite applications, like Spotify, and even Netflix or YouTube. A range request is when a client reaches out to a server for information in its HTTP header to first know if the server supports range requests.
If range requests are supported, this means that the client can query what is essentially a table of contents for the data to retrieve only the data that the client is interested in, then stream it back for use.
If you are familiar with Python, you can think of this almost like slicing a list [x:y]. By pulling only the data relevant to the query, performance is greatly increased as fewer packets can be transferred back to the client.
At this time, streaming optimized geospatial formats are mostly limited to raster and point cloud data. There are nevertheless hopes that we will see options for optimized vector formats in the not too distant future. Open vs Closed Geospatial Data Standards In the past, GIS was often more or less siloed within an organization, making it reasonable to work in closed or proprietary formats. Today, in an increasingly interconnected world, sharing data and preparing it with interoperability in mind is essential.
Open data standards have been embraced by many as they can be used as-is, or maybe modified to work with existing infrastructure to create a better fit for an organization’s needs.
The Open Geospatial Consortium in particular has been instrumental in expanding the prevalence of open formats by publishing and documenting open data standards.
Having this groundwork in place gives developers a good spot to start from when creating a custom implementation, and leads to greater potential for innovation.
Although open data standards and formats have come to dominate the industry, closed data standards still have a presence. The best example is Google Earth Engine. Within its system, Google Earth Engine ultimately handles data in a closed proprietary format.
They know, however, that consumers today are generally unwilling to accept the risks of holding all of their data in a closed format.
After all, if the provider went out of business, then the consumer would lose usability of their data. Google remedies this by allowing other data formats to essentially port into their own, allowing clients the flexibility and security they would get if they utilized open standard formats. What's Next for Cloud Native Geospatial? CNG has already brought a huge paradigm shift to the industry, removing traditional access, storage, and processing barriers for those who enter the game. As more and more datasets are uploaded into the cloud, it begs the question of what the next great leap forward will be.
One potential development we may hope to see in the future is data that is optimized for retrieval by search engines. At this point, the data exists, but it needs to be described in a way that allows the search engines to find it, and match it to user needs. This means fully populated metadata, and plain text descriptions that allow data to be matched to a query.
More accessible and queryable geospatial datasets could play a huge role in bringing geospatial to the masses, rather than following the current theme of only being used by those involved in the GIS industry. The SpatioTemporal Asset Catalog is a great example of what this may look like.
Another glimpse into the future of CNG is that eventually, pretty much all the data we could want will be available in the cloud, in open data standards. This can allow the focus to shift from data aggregation, to creative data analysis and applications. As young people come into the industry, they will not be clouded with the ideas of what cannot be done, but will rather see the wealth of options and resources available, and take it and run with it, hopefully leading to the next big thing.
Our guest today is Dr. Adam Symington, the lead data scientist at Geollect. Although now he is known for his impressive geovisualization, built using Python, his professional training is as a computational chemist. After completing his schooling with University of Bath he stuck around to complete postdoc work on the development of machine learning models used to track the movement of atoms within materials. At Geollect, he uses a similarly inspired approach to track the movement of ships using AIS data. Python plays a significant part in this work, and is the driving technology behind his spatial visualization hobby project- PythonMaps. The Path to Learning Python Anytime someone new to geospatial asks the greater GIS community what the best skill they can learn to increase their value is, the answer is the same- Python. Highly praised as flexible and easy to learn, Python is an excellent choice for those looking to bolster their resume. Of course, just because a group of computer scientists say Python is easy to learn, does not mean that the process is necessarily intuitive to those learning it as their first programming language.
Everyone has their preferred style of learning. There are a huge variety of books, tutorials, and full-on web courses designed to teach Python to beginners, many of these resources are even free.
Regardless of which medium you choose, it is important to start from the beginning, and dedicate some time to really understanding the fundamentals of Python.
Internalizing explicit explanations of what is a variable, what is a function, what is an object, etc… lays the groundwork for the more complex concepts ahead, such as defining methods, loops, and arrays.
Once you are comfortable with your knowledge of the basics, it’s time to start coding. The most important part of your Python journey is to stay excited about what's next. Find a problem that you find super interesting, and then find as many ways as possible to solve it using Python. The motivation to answer your original question will be essential when it is time to power through the inevitable knowledge gaps and challenges you will encounter in your projects, and when solving real world problems. What is PythonMaps? PythonMaps began as COVID lockdown project, where Dr. Symington was experimenting with creating Python-driven geo visualizations to further his own skills. As time went on, and he developed more expertise, and consequently higher and higher quality maps, he went public. Now you can find his work on Twitter, LinkedIn, and Reddit.
PythonMaps is the perfect example of why it is valuable to learn Python through personal projects. When there is genuine interest driving a project, it is much easier to find the motivation to learn new technologies and skills to support it.
Although Python offers thousands of libraries, only four are really used for most PythonMaps projects. These are matplotlib, geopandas, rasterio, and rioxarray.
Matplotlib and geopandas are classic data visualization libraries. Matplotlib’s focus is more on traditional datasets, but it does have a number of mapping functions. You can also check out the Matplotlib Basemap Toolkit API to add additional functionality and context to your creation. Geopandas is a geospatial expansion on the highly popular pandas data science library. It provides functionality to read and write geospatial data into dataframes, as well as a slew of great options for geospatial data analysis, summation, and visualization.
Where the above options are generally geared towards vector data, fear not, there are programmatic options for working with raster data in Python as well. Two useful libraries are rasterio, and rioxarray. Both of these libraries allow the reading and writing of a variety of popular raster formats, rioxarray has the added benefit of explicitly supporting cloud optimized geotiffs (COGs). These libraries can also provide core functionalities like reprojecting, resampling, and masking rasters.
There are many more data vizualization libraries available, but it is worth keeping in mind that for beginners, the most popular libraries will be the easiest to work with.
This is due to the larger knowledge bank of documentation, and existing troubleshooting conversations on platforms like StackExchange.
Having a mission in mind and playing around with these libraries is a great way to get started on producing your own content, but if you need some specific ideas, or a bit more help, Dr.Symington has a number of tutorials here. What Are the Benefits of Mapping with Python? In today’s information age, the power of a strong social media presence cannot be underestimated. This does not mean you need to be producing daily, or even weekly content, but it does mean that you should be investing some conscious time and effort into what does go public. Curating a professional online presence pays its own dividends, but even if you don’t get famous, your time will be far from wasted.
If your goal is fame, great. As you spend some time building up your Python and cartography skills, start thinking about how you want to present your creations, and yourself. Do you want to create a brand, like MapScaping, or would you rather take the personal route and present yourself as your brand, like Joe Morrison?
Building a brand allows you to distance yourself from criticism, and grow the project beyond what you alone can accomplish. On the other hand, making your brand personal may make it easier to attract a following.
It also has the potential to amplify the successes, and failures, that you encounter as there is no buffer between you, your work, and the criticisms of the internet.
The internet is infamous for harsh criticisms. Behind the safety of a screen, people are more willing to share their unfiltered opinions, but this is not always a bad thing. If you stay open-minded (and have a thick skin), those criticisms can be read as feedback.
Paying attention to negative comments that have some substance to them can show you where you still need to improve, in the same way that praise can tell you what you are doing right.
If your objective is not fame, then what does the return on investment look like here? Well, as many say, it is about the journey, not the destination. Learning Python and improving your cartography skills can do great things for your career and portfolio. It shows that you can succeed at self-directed learning, and can create value as an independent worker. Being able to communicate effectively is one of the most valuable skill sets in data science, and maps are a classic tool for visual communication.
Remember to focus your projects on topics you enjoy, and no matter the outcome, it will never be time wasted.
Daniel:
Hi Michael. Welcome to the podcast. Thank you so much for taking the time to join me today. I really appreciate it. So in just a second, we're going to talk about open-source geospatial journalism, and we'll dive into exactly what that means in just a second for the listeners. Before we get there, can you just take the time to introduce yourself, and perhaps explain how you got involved in geospatial and journalism?
Michael Cruickshank:
I'm Mike Cruickshank, a journalist, originally from Australia, and I actually got into doing geospatial through journalism, which I suppose is a strange entry point to geospatial. This happened because I got involved in this kind of journalism called open-source intelligence. This uses a lot of geospatial data for cross-referencing things by using techniques such as geolocation.
This first brought me in contact with satellite imagery through Google Earth and that sort of thing. I just became more and more interested in what this data is, how can I use it to do more than just have something to look at? What stories are there that kind of lie beneath the pixels, so to speak, of the data and how can I use it in more interesting ways? From there I taught myself how to program, and work more deeply with this data.
I've learned how to use all these different geospatial tools and programs. I'm also doing academic research now, I'm doing a master's thesis where I'm using this kind of data to investigate links between climate change and conflict. I’m also working professionally at a company in Berlin called LiveEO, where we use geospatial data to provide information for industries on things like environmental issues or forestry, as well as monitoring industrial assets. It's quite a broad range of things that I'm involved in these days.
The Link Between Journalism and GISDaniel:
It sounds like you're moving away from journalism and into geospatial. What kind of skills can you already see that are going to be really helpful in that move, that you'll take with you from journalism and over to the geospatial side?
Michael:
I wouldn't so much say I'm moving away from journalism as I am trying to tell a different kind of story to a different kind of person. In the past, my thinking was all about how do I inform people about what is going on? I came to the conclusion that, while I still want to inform as many people as much as I can about what is going on, I also want to be able to reach more of the kind of people who are actually making the decisions in the world, and provide them with information on how things will go into the future.
I'm particularly interested in looking at the effects of climate change and its security implications. I want to be able to provide insights on this prediction, analysis, etc., using geospatial, as well as some of my background in journalism on thinking about the real effects of these things. I want to present this kind of information in a way that's engaging to people so they actually take it seriously, and just reach in a more targeted way the kind of people who can actually make a change in this world.]
Daniel:
You had this great sentence there, you said something like, "I want to tell different stories to a different kind of people." What was it about traditional journalism that you felt like wasn't working, or could be done better?
Michael:
I felt that in many ways, people were flooded by the same story. I was working in conflict journalism so perhaps you could go to a country, you can go to the front line of their wars, you can tell a story of, oh, it's very awful. It definitely is awful what is going on in these places, but the unfortunate matter of fact, is that so many people are telling these kind of stories, people just tend to switch off.
You need to find a different way of reaching them with a different kind of story, one that engages them in a different kind of way, and perhaps tells a bigger picture or a different picture.
Daniel:
I think you're absolutely right with that. I think that's an absolutely brilliant insight. I guess that the danger is when we think about engagement, we sort of fast forward this trend towards the short, incredibly clickable link bait. I really do feel like we are a world trending shallow, as opposed to deep. We see this in all the social media trends, generally. It's quick, short snippets. Do you feel like that's the way to go? If we think about engagement, the numbers speak for themselves. That seems to be what people want, what people engage with. When I think about journalism, I kind of expect something deeper.
Michael:
You're well within your rights to. I certainly think that going after the kind of engagement that you are talking about is a bad thing for publications and for journalists to do. I suppose when I talk about engagement, I don't mean someone just reading something, clicking it, sharing it, liking it, reacting to it. What I mean is someone who's really taking stock of what is written, what is seen, etc., and thinking about it and incorporating that knowledge into their life and their decision-making.
Daniel:
So I have to ask here, we are overwhelmed with media, we are flooded with information and data. How do you think your media, your style, this idea that you are going after here is going to cut through the noise?
Michael:
I'm not going to put myself up on some pedestal and say I could solve this problem on my own. What I think is it forms part of a body of information which is more verifiable, and forms more of a ground truth, or a reality of what has happened or what is going to happen, rather than just 1,000 different conflicting reports from different directions within the political sphere or different interested groups, etc. I think that part of the way that we combat this flood of information is simply providing people with better information, rather than just increasing the amount of it or telling them to consume less. Instead, we should suggest and provide the option for them to just consume better information.
The Tools of the TradeDaniel:
Yeah. I totally agree. I feel like that volume game that you're talking about inevitably leads to a lesser quality of information, or lesser quality of work because it's a volume game, right? It's about getting the amount out there as much as possible.
I know you've come a long way in your journey with geospatial, I know that you program in a couple of different languages now. Could you explain to me what it was like right at the start? Were you using tools like Google Earth just to look at the images and see what was on them and try and geolocate stuff there, or were you using something completely different than that?
Michael:
No, I was using Google Earth or Google Maps, those sorts of platforms, and just purely analyzing things on a visual level. There was no sort of looking at the data behind it, it was just all about identifying things that appeared in pictures, using video locations. I was using platforms like Google Earth, Wikimapia, and a few others to get different kinds of images or different dates of images. Beyond that, I hadn't gone particularly deep into it.
Daniel:
Okay, so that's where you came from. You started off using these very simple visualization tools like Google Maps, I believe you said. Where are you now? What tools are you using and where do you get your data from when you think about doing open-source geospatial journalistic work?
Michael:
These days I'm using a large number of open data sources, but I think the primary and most useful one is Google Earth Engine, which of course is very different from Google Earth itself. Google Earth Engine, as some of the listeners might be aware, is a multi-petabyte catalog of huge amounts of open-source satellite imagery from sources like the European Copernicus programme, Sentinel satellites, NASA's Landsat program, and many others. The beauty of Earth Engine is that it allows people to basically program with the data. You don't have to download these very large raster images and all that to your computer, you can work with it all remotely. The processing is done off your computer. So this is quite powerful because it means that you don't have to mess around with huge amounts of files and huge amounts of processing locally.
Daniel:
So maybe I made a mistake. I've been saying open-source geospatial here, Google Earth Engine is not open-source, even though some of the data might be. Are there any other tools you've been using, or have I made a mistake on my end?
Michael:
By open-source, in this case, I was referring to the data itself. Earth Engine is certainly not open-source, and you can only use it for non-commercial applications, as far as I'm aware at least. In terms of fully open-source programs, I use QGIS as well, so when I'm doing more visual-related output, more than processing, more than programming, more than analytical stuff, when I'm doing visual, I will use QGIS. For the more programming side of things, I'll usually use Google Earth Engine.
Daniel:
So let's stay with Google Earth Engine for a second here. What are you doing? I know you can make these amazing time lapses in Google Earth Engine. What kind of analysis are you doing in there?
Michael:
The kind of analysis that I'm doing, really depends on the use case. Basically, you can do any kind of transformation of the data. If you imagine these satellite images are effectively raster arrays, or just data points, you can do any kind of mathematical operations on the data. You can compare bands or images over time. You can create custom visualizations, and programs that just access this data and then work with it in other ways. You can create front-end and backend applications of this data. Recently, for instance, I wrote a small web program that enabled users to work with Sentinel-1 imagery of some of these Russian military bases where they were building up troops before they invaded Ukraine.
This was taking data from Earth Engine, specifically from the Sentinel-1 satellites, and then creating time lapses out and creating different kinds of visualizations out of this that were presented in such a way that a user could do this without any background knowledge of how to actually interpret the data. Often, especially with synthetic aperture radar (SAR), it's very difficult to get down to what you're actually looking at unless you are experienced with working with that kind of data.
Using Geospatial Intelligence Techniques for JournalismDaniel:
When we say Sentinel-1, are we talking about the SAR band?
Michael:
So Sentinel-1 is actually two satellites. One of them is currently non-operational. I'm not sure whether it's dead or just not working currently. So it's two satellites and they're both synthetic aperture radar, SAR, satellites.
Daniel:
What were you doing with this data? What kind of analysis could you do on SAR data that would give you information about what's happening with the Russian military in this case?
Michael:
The resolution of Sentinel-1 is quite low, it's approximately 10 meters per pixel. Obviously, you can't see something like a tank with this; a tank is smaller than a 10 meter square. If you know the areas where the tanks are being stored, at these bases, you have the polygons of the base. When more vehicles move into the base, the mean reflectance in the synthetic aperture radar increases. You can plot the mean of a time to get an idea of whether the number of vehicles in the base is increasing or decreasing. Moreover, if you have high resolution optical imagery, you can get a baseline of what a certain level of mean reflectance represents in terms of actual real numbers of vehicles.
Daniel:
How do you ground truth, something like that? Or can you use any other data sources to get an idea of perhaps a more precise number of how many tanks, in this case, that there are in an area? Can you incorporate other data sources into this kind of analysis?
Michael:
Absolutely. In my case, I have been using very high-resolution optical data as a ground truth that was taken on the same day. Obviously, this is not always possible because of clouds, especially in the more recent months in winter in Europe, it's very cloudy. You have to kind of wait until you catch a break. This usually can be done with a very high-resolution optical image.
Daniel:
How do you document this, and what are the results of this kind of journalism? Has it been cited anywhere? Are people using it as evidence anywhere? I guess what I'm looking for is- are people referring back to this, almost like a peer review?
Michael:
Looking at these Russian military bases, this was getting a lot of play in the media. I wasn't the only one doing it, so obviously I'm not going to take all of the credit for this. There were a large community of people within the open-source intelligence community who were sharing this kind of analysis. Many of the initial stories about this buildup were coming from within this community, and then were taken up later by the more mainstream media who were then tasking even higher resolution satellite imagery for their stories.
This really did get picked up in the lead up to the invasion. It helped in many ways to lend credence to what countries like the United States were saying when they were saying, "There's definitely going to be an invasion. This is how it's going to happen. They have all these troops here." The advantage of having people doing this kind of open-source intelligence journalism, using things like geospatial data is that we can then have a second data point and say, "Well, they're saying this, but now we can check if there really is data to back this up." In this case, there certainly was, and we've seen what's happened since then.
Using Social Media as a Data SourceDaniel:
Twitter is my social media drug of choice. At the moment, my account is flooded with images and short video clips of tanks and military operations in the Ukraine. At the same time, I'm constantly wondering, is this real, this thing that I'm looking at? Does this make sense? Is it misinformation? Is it disinformation? Can I believe this? Maybe before we dive into this as a topic, perhaps you could explain the difference between misinformation and disinformation.
Michael:
I'll start with misinformation. Misinformation is when someone shares a piece of information that is untrue, however, they don't necessarily know that that information is untrue. More often than not, they will probably believe that it is true. Disinformation is the case when someone is maliciously and with intent sharing something that they know to be untrue. The differentiation lies in whether the person sharing knows whether what they're sharing is true or not.
Daniel:
So this has been a huge problem, obviously, not just in the current situation we are in now with Russia and Ukraine, but also in other political theaters around the world. Is there any opportunity here to sort of prove or disprove some of this kind of information that we are seeing in social media feeds, for example, by doing the kind of journalistic work that you are doing? Or can we use this data together with what you are doing?
Michael:
Absolutely. Using open-source intelligence is one of the best ways that we can actually prove or disprove whether something is indeed misinformation. The primary tool is geolocation, where you can compare elements within, say, a video or an image to satellite imagery to work out whether it really is where it is said to have happened. This often makes it very easy to filter out if something is misinformation or disinformation. You know something is misrepresented, because you can immediately see that, okay, this is indeed where it's said to have happened, which then lends credence to the idea that that is indeed true.
You can take geolocation further and move on to chronolocation, where you start using things like elements of the shadows within the videos and backing this up with geospatial data on the exact position to calculate the exact time that a video or an image was taken. This further leads you down the path of whether this is true, or something which is misrepresented.
Daniel:
I can definitely see the power of geolocation here, but how do you do that? Are you looking for a place name in the video? Are you looking at some sort of metadata in the video, perhaps the geolocation tag, that kind of thing? Or are you doing something completely different to locate that image or that video?
Michael:
If there is a name, place name, shop name, or a business name within the video, this obviously makes it very easy. More often than not, that's not present, or the video has such low quality that it's difficult to actually read what's written. The more common way this is done is by cross-referencing visual elements within the video to visual elements within satellite imagery. This would be things like the architecture of the buildings, the specific layout of streets, of plants, and that kind of thing. (Much like Geoguessr). You then think about what that same scene would look like from above, and then cross-reference that scene with satellite imagery of the location where the video or image is said to have taken place. You then see if you can work out exactly where this happened. More often than not, you can work out not just where, but what angle the camera was pointing in, what time of the day, what time of the year, etc..
Daniel:
This reminds me of trying to orientate myself to a map. Looking at a map, looking for features that I can identify on a map, and then looking up in the real world and saying, okay, where are those features relative to me? Can I use them to orientate myself if I'm out hiking, for example? So this makes sense to me, but we live in a time of deep fakes. You will have seen these videos on the internet before. Is it possible to fake a location, when we think about the kind of information we can glean from social media?
Michael:
It is possible to fake a location, or at least fake a plausible looking location. Geolocation is possible 99% of the time in any video. If you can't find where somewhere is, if you can't get a fix on it, you have a pretty strong indication that this may indeed be a fake location. I haven't specifically heard any reports of videos that have been shared from things like conflict zones where the background has been completely faked using these kind of deep fake techniques, but it's certainly possible.
There's a website called This City Does Not Exist that sort of generates these deep fake landscapes, and there's no reason someone couldn't do something like this. It would take a lot of effort, and looking at some of the misinformation or disinformation that's being spread around, most recently in the conflict between Russia and Ukraine, most of it is pretty low effort. This sort of stuff takes a lot more effort and I have yet to see it being used, especially at scale.
Doing the Analytics and ResearchDaniel:
I'm imagining that the power of this kind of journalism is when you can start joining the dots. You see something happening over here, and you can connect it to an event somewhere else, or you can see that things are moving. How do you keep track of these events? How do you know for example that, oh, okay, I'm looking at tanks here now. I've used my SAR data. I can see that the values are changing over time, so there's change over time. I've confirmed this one location using social media, and georeference some of the videos or images that I've seen there.This is sort of giving me a picture in my mind and understanding of what's happening. Tomorrow, the tanks are gone, things have moved. How do you link that with other events, or how do you find where they've gone?
Michael:
I approach it kind of like a scientific, or a proof statement almost. You're trying to prove that something happened or didn't happen, or trying to work out what the limits of what you can know are. You try to create a chain of evidence going from the open data that you have, and then see how those things link together, how you bring this all together, and how each of these things mutually reinforce, or weaken the argument of the other. In terms of your question you asked, how would you know where the tanks have gone to? Well, maybe you've seen the mean SAR values at this base decrease, then maybe a day or two later, you start seeing videos pop up on TikTok of large columns of military vehicles moving around another town, somewhere else in the same region. So you think, okay, well maybe the vehicles are moving towards that area, but you don't know where they're going.
In my case, recently, I created a map of an entire region, Oblast in Russia, just doing basically change detection in SAR and seeing which areas specifically had had large increases within the last week or two. Obviously you get a lot of false positives; you get things like snow melting, lakes melting, or different landscape changes. Some of these places, however, you might see a very large area of change; a polygon that looks perhaps somewhat artificial. Then you can sort of cross-reference this with other data and think, okay, well maybe there is something there. Then you can download very high-resolution satellite imagery, optical imagery of this area to try to cross reference and determine, is this really something which is interesting or suspicious?
Daniel:
I've got to tell you that this idea of using geospatial in this way for journalistic research is kind of new to me. It makes perfect sense though. It seems like a great tool for this kind of application. Is this different from what you were doing when you were working as a journalist? Is this completely different from the kind of work that journalists would generally be doing?
Michael:
I think in some ways, it is very different from traditional journalism. Traditional journalism is all about building human sources, and your strength of your argument, or of the story that you're writing. It is almost based on your level of access that you have to reputable or trusted or well-positioned human sources. That is the skill; it's a networking skill in many ways, it's being in the right place at the right time, etc..This kind of journalism is about pulling together a whole lot of things that everyone knows, or everyone could know, and then seeing how those things all work together to prove something and drawing connections that other people aren't.
This is unlike traditional journalism, where it's all based on trusting that this is a reputable publication that wouldn't lie to you. You've got to trust that this anonymous source is really a real person, etc.. The opposite is true with open-source intelligence. Everything is effectively a proof statement. You don't need to trust the journalist. You can read through the way they've brought together their argument and their evidence and see if you agree with it. There's nothing that's hidden. Everything is in the open.
Daniel:
For me, this begs the question, why aren't big publishing, media, journalistic companies doing this? Or are they doing this? Do they have their open-source and intelligence center within their magazine or their media company?
Michael:
Over the last few years, there have been more and more mainstream media outlets that have been starting to use these techniques. It is growing in popularity, but it's also a relatively new thing. This really only started probably in 2012, 2013. That was when the ball started rolling. It has just grown more in importance over time, and in popularity and sort of the credibility that it has.
Taking the Story to the PeopleDaniel:
I'm pleased you brought up credibility. What do people say when they push back on this, when they don't believe it's a great idea? What kinds of arguments do they come with?
Michael:
It depends. A lot of the time they simply just don't really understand what you are doing. They'll just say, "Well, this isn't forensic enough, this isn't scientific. You're letting your own biases get in the way. Your analysis is wrong." They don't really understand how these things prove each other and so they just say, "Well, you're just hand waving things around and saying that this equals this," and they don't really understand why this equals this, so they dismiss it. To be fair, most of the criticism I've come across is more or less just bad faith criticism in the sense that these are people who have motivated reasoning and start with a position of "You're definitely wrong, because you're saying something I don't like." I find that there's little good faith criticism, and it's mostly just bad faith, politically-motivated criticism.
Daniel:
You brought up a great point there, that people didn't understand the analysis you did; perhaps they didn't understand the techniques, they didn't understand the data that was being used. You're telling people, "Yeah, I looked down from space and watched the change over time." I mean, it sounds pretty fantastic. How do you explain that to people?
Michael:
This is sometimes the harder part, because you have to explain a bunch of things that aren't always intuitive. I suppose for people who work in areas like geospatial, thinking in a spatial manner is very natural to us. For other people, sometimes it's not. Imagining what a scene might look like from above, imagining how things could look over time, etc., or how things can be abstracted in a 2D or 3D format, can be difficult for some people. You need to make this easy, and so having really good visualizations is incredibly important for this. I've seen some groups that make really nice videos that are sort of linking things together with nice effects to sort of show how things transition over time. It's all in the kind of visual communication you use.
Daniel:
Oftentimes, I think that the people that work in geospatial, we get stuck in only doing analysis for big companies or municipalities, organizing and maintaining the data. When I meet people like you that are doing work like this, I think, wow, here is another opportunity, here is something that people should know about, because it makes so much sense. We can use these tools that we know about, that we're comfortable with, and we can use them to tell a story. For me, the skill sounds like investigative journalism that tells a story, and backing it up with evidence. Could you imagine a time where new journalists need to learn these skills just in the same way they need to learn interviewing skills or other media skills, making that data analysis a part of their education?
Michael:
Absolutely. It will be critical going into the future. It's the only way that journalism can push back against the kind of flood of misinformation and disinformation that we're currently facing. This is a perfect way to combat it. Rather than just retreating back to ideas of, well, I'm a reputable outlet, or, you have to trust me at my word, instead, you're saying, no, don't trust me. You're right to not trust me, but here, I can prove it. I think this will hold much more weight than simply asking people to take you at your word. If more and more stories are presented in this way, then it won't be so much that people trust the media more, but the correct information will get out more. In the end, that’s what's important.
Daniel:
If you do a great job of documenting this kind of work, do you think this will help people think critically about how people draw conclusions when they see other types of journalistic work?
Michael:
Absolutely. I think this idea of thinking very forensically, in terms of how you prove something, and what proof really means, and where the edges of what can be proven and can't be proven lie really helps you narrow down where the gaps in our knowledge are, and where the area for debate lies. That way we don't spend so much time talking about things that are kind of extraneous, or are red herrings within the conversation.
Daniel:
Michael, I think we could probably round things off here. I'm curious if you have any recommendations or references. If I want to learn more about this open-source intelligence, if I want to be a part of this, is there a community I can join? Is there a newsletter I can follow? Is there somewhere I can go?
Michael:
In terms of communities, there's a great Discord community called Project Owl, which I think currently has about 25,000 users. It's growing very rapidly with hundreds of people working on all sorts of different projects all around the world, all brought together by this kind of usage of collecting information from open-sources, especially about conflict zones, but also about many other different areas- bringing this together and synthesizing this information and seeing what can we prove? What can't we prove? etc..
In terms of other sources online, I think the best group in the world who's doing this is Bellingcat. They're a UK-based publication started by Eliot Higgins, and they've done some really great investigative work. I believe they've also won a Pulitzer Prize, if my memory serves me correctly. They were looking at MH-17, some of the Novichok poisoning attacks in the UK, they've done all sorts of interesting looks at different conflict zones around the world; in Syria, in Yemen. I've also written a few articles for them myself, looking specifically at Yemen and Ukraine. There are lots of great people associated with this and they're doing really good work all the time.
Daniel:
So I really hope you enjoyed that episode with Michael Cruickshank. I found a few other articles that I want to link to and share with you. And one of them is called The Growing Problem with Deep Fake Geography: How AI Falsifies Satellite Pictures, or Satellite Images. There's another one here along the same lines, When is Satellite Imagery Fake? And the third one, Why Newsroom People Need Expertise in Remote Sensing. I'm also going to include a link to a newsletter on Substack that I found that I think you might find interesting. It's called actualcontrol.substack.com. If I read the headline of this blog newsletter, it says A Blog About Satellite Imagery, Social Media, and Other open-source Information from All Corners of the Internet. There will also be a link to the publishing house that Michael mentioned, called Bellingcat.
Before I let you go, I just want to highlight one of Michael's insights, and that was at some stage during the start of the conversation he said something like, "I guess I wanted to tell a different story to a different kind of person." He used this idea of telling better stories. I think this is really important: telling better stories. We talk about data stories and we talk about customer journey stories, but we never talk about telling better stories. Michael wanted to tell better stories. He wanted to do that because he discovered that the old stories weren't working anymore, they were being drowned out. They were too similar. Everyone was telling the same story so it wasn't sinking in anymore. It was blending in, it was becoming part of the background noise. That sounds really simple, right? Tell a better story. What does a better story look like? I'm not completely convinced you need to know what a better story looks like, or what the right story looks like. I'm more convinced that you need to try a new story. If you try enough new stories, you'll find the right story.
Okay. That's it for me. That's it for another episode of the Mapscaping podcast. I'll be back again next week with a new story. I hope that you'll join me then.
Daniel:
Hi, Jeff. Welcome to the podcast. So you are the director of user experience at a company called Element 84, and you've got this amazingly rich background in design, and now you're working in geospatial. For the sake of context, could you just take a couple of minutes to introduce yourself to the audience and perhaps explain to us how you got involved with design, your background in it, and how that led you to work with a geospatial company?
Jeff Siarto:
Absolutely. I'm Jeff Siarto. I'm the director of user experience (UX) at Element 84. E84 is a small software engineering company that specializes in geospatial and science systems. I got here through a sort of roundabout way. My educational background is not in geospatial, or science. I studied general design and user experience in college. After that, I did freelance web design for a while. During that time, I wrote a couple of books for O'Reilly and their Head First series. In that time, I met Dan and Tracey Pilone who run Element 84. They were co-authors in that same series so I was able to get to know them early on in my career.
Before I came to E84, I started a small social media analytics company, which I ran for about five or six years. We did early social media analytics for consumer product brands. I worked with companies like Wrigley, and New Balance, and Radio Flyer, and we essentially monitored their social media traffic, and then helped their marketing teams kind of figure out messaging and how to engage on early social media channels. That was sort of my initial background. I worked on the software there, and I worked on the design of those projects. After I was done with that, Element 84 reached out to me and asked if I'd be interested in working on this NASA project. I said, "I don't actually have anything going on right now. That sounds really interesting." That's sort of how I got here. I came onto that project knowing very little about remote sensing and earth observation, but bringing my design background to bear on their problems. They had never actually worked with any designers before, so I was kind of new to them as well. I sort of just learned as I went through some of their early projects.
Design and Earth ObservationD:
So we're going to talk a lot more about design and earth observation in just a second here. I'm curious, when you think back to when you started at Element 84, you said that they'd never worked with a designer before. What was it like for you coming in? Were they doing the right things? I guess a lot of the semantics around earth observation was new to you. When you looked at the products that they were making, the things that they were doing, was it all brand new, or could you see a lot of the design elements that you were used to working with, or was it something totally different?
J:
They were getting there. I was actually the first designer at Element 84, and then I was also the first designer to come onto this NASA program, so they had some foundational pieces there. They had great backend systems. They were working in the right direction, and I was able to come in and sort of help them improve the interface, or get their interfaces to the point where they were in the same league as some of the backend systems that they had. They were kind of getting to that point where they had solved a lot of the early problems, but now more of those user experience problems were starting to surface. That has ended up being a trend throughout my career here is once those early problems get solved, those user experience problems sort of creep up and become the major players.
D:
Now that you've been working in the industry for a while, when you think about earth observation and design, do you think a lot of the companies that are doing similar work to what you are doing are facing the same problems? I mean, does it surprise you in any way that we weren’t thinking about design earlier?
J:
I do think that more companies are starting to look at that. As we start to solve those kind of middleware early problems, the piping, and the user interface problems start to creep up. I don't necessarily think that it's surprising. If you step back and look at the timeline of the early internet and where all of that went, design became more important as the web matured. A lot of people sort of moved from that graphic design discipline onto the web. They brought those skill sets to bear in an area that was highly technical before it was a focused design medium.
I think we're seeing that in geospatial, particularly as the tooling gets more web-based and cloud-based, or you're interacting with these applications and APIs through the internet, you're going to encounter more user interfaces. I think the importance of those user interfaces is in making the job of whoever that end-user is easier.
Reducing the Time to ScienceD:
In one of your tweets, I saw this great catchphrase, and it was reducing the time to science. My first question is, how is design going to reduce the time to science? I think design could mean a lot of different things for a lot of different people, depending on what it is that we are working on. If we think about the geospatial stack, we could divide it up broadly into the backend and the frontend. I'd like to start with the backend. You've been talking about it as pipes and infrastructure up until this point. How do we bring design to bear on that? What bits can you come in and design to make it better?
J:
Time to science is super important. To define it a little bit, it is the amount of time that a particular scientist or researcher will spend getting to the answer of the question they initially posed. In the past, scientists have spent 50, 60% of their time doing data wrangling, or data processing- getting all of the bits together so that they can answer the question.
The first part of there in reducing time to science, is sort of smoothing out that backend process. Maybe that's going from having to deal with level zero data, or level one data, up to a more refined data product that is properly subsetted or properly re-projected, and taking care of that for the scientists. Now they don't have to do all of that specialty work on the data, and that gets them one step closer. If you sort of continue on with that, particularly as we move into a cloud data paradigm, maybe we go even further. Maybe we take it and we put these data products into a format that can be easily visualized, or easily manipulated in a way that does require programming. So we've sort of eliminated some of the lower level tasks from the workflow and we sort of shorten that time to science. As a result, now maybe they only have to spend 30% of their time getting the data in the right place.
If you go back even further, some 60% of data wrangling time was just literally waiting for data to download. I would hear stories when I first started where we would go out and talk to scientists, and we'd watch them work on their projects. We'd look at their workflow and see how they interact with the data. Oftentimes, they'd just have to set up a download, and just let it go, and come back a couple days later when all their files were downloaded. Some of that time to science was literally just kind of sitting around and waiting.
Even just improving how you interact with the data, and just generally improving the speed of the internet that people have access to is user experience as well. The speed of the internet access, the way in which they have to get the data onto their machine, and what formats that data is in. They all are related to the overall user experience.
Data Science Infrastructure as a CommodityD:
When I think about the backend, apart from the data manipulation to get it to that analysis ready-state, I think about infrastructure like blob storage and the cloud. I think about infrastructure like cloud optimized GeoTIFFs and other cloud optimized data formats so that we can stream data directly to a client without having to run a database somewhere in the background. I think about things like the stack interface. It seems to me that these pieces are sort of slowly coming together, and that at some stage, they're going to be a commodity. It's going to be “of course you're running with these types of standard components”. I'm wondering how you see that and think about that in terms of being a commodity. Do you think it is at the moment? Are we there yet? Of course you're using these things. What else would you be using?
J:
Yeah, I think we're moving in that direction. I would not say we're quite at a commodity yet, but I think we're getting there. I think you can tell that we're getting there in a couple of different ways. You're starting to see the major cloud players building these systems up. I'm speaking specifically of the AWS public data sets and the Microsoft Planetary Computer. You have these big cloud players that are kind of investing in these open geospatial systems. They are building this backend, this foundational layer of data products and APIs that then can be built on top of in a standardized way.
I think we are in the building stage for that layer. I would say I don't really know where we are in that building stage, maybe halfway through. I think in the decade-ish that I've been working in this, over the last three to four years, you've really seen an uptick in cloud data adoption, and seeing investment in those platforms that are going to maybe become a commodity in the short term. I don't think we're quite there yet, we're still in the building phase.
D:
It sounds like you are thinking in perhaps the same sort of lines as I am- that this will be a commodity at some stage. When I think about earth observation, and I think about software companies that are looking to build a business around it, I think it's going to be pretty hard for them to create the data themselves. There will be satellite platforms which are going to create the data. It's going to come down to earth. It's going to go through a commoditized system, and it's going to end up in the frontend. I think the place where a lot of value would be created, would be in the frontend. Do you think that I'm on the right track there?
J:
Yeah, absolutely. Part of the issue that we're facing right now is that Earth observation and geospatial data is fairly complex to use. We're taking a step in the right direction by moving that to the cloud and creating standardized systems on top of that, and now we've sort of eliminated that initial barrier. Once we have that, now the companies, and the startups, and the government agencies that want to start spending more time answering specific questions can do that. If we've got a commercial interest or issue that we want to answer a question about, now, you have the foundational pieces in line to be able to answer those questions, and focus on taking the data that is in the standardized format and manipulating it in a way to answer a specific question for a specific customer.
All you have to do is sort of focus on what that question is and how to build or manipulate the system to answer it. But you no longer have to worry about “Okay, where am I going to get this data? What data set do I need to use?”. You won't even really need to worry about what the sensor is or what the platform is. You'll have the normalized data ready to go. So I feel like that allows you to really focus on the customer problem.
I think a good example of this is a company like Twilio, which is sort of like a communication and SMS/voice API over a bunch of lower level technologies. If I want to then build a product that utilizes some sort of SMS messaging, or some sort of voice system, I don't have to then go interact with the SMS layer, or figure out how I'm going to do voice over the internet. I can just build on top of Twilio's system. They've sort of standardized out that lower level tech for me. I can focus on my AI chat bot, or my SMS messaging tool or something like that, without having to worry about that slightly lower level. It frees you to think about the higher level problem, as opposed to where you're going to acquire the data, and how you're going to get it into the right format.
The Future of Frontend Design in Geospatial Data ScienceD:
When I think about the frontend opportunity, we have these great platforms we can build on. I think about the brand-ability of a frontend. The way you can dress something up, you can change the look and feel of it. You can add an identity to it, and that can become a selling point in itself. It doesn't feel like those in the geospatial industry are that interested in changing the way we do frontends, and designs. It seems to me that we are pretty committed to using the tools we've always used. Firstly, I'd like to know what you think about that statement. And then secondly, I'd like to say do you see any opportunities to change things here? Any things that could be done differently?
J:
Yeah, I do agree with that. There's a pretty consistent set of tooling that's been used forever. I think there's a lot of standardized patterns that we've seen used quite frequently. I'm kind of talking about a map interface where you can drag around a map, you can create a bounding box, and you can sort of interact with a Google Maps style interface, which has become fairly ubiquitous. Ubiquitous both on the sort of geospatial earth observation side, but then also the consumer side as well. From a geospatial perspective, when you're looking at a map with a bounding box and a sidebar of layers, you're still thinking- “I need to pull data from around this area. Give me all the stuff inside of my bounding box.” I think there are parts of that that are going to carry over. I do think that we're going to need to move into a direction where we are building interfaces that are answering specific questions as opposed to just building interfaces that are allowing me to grab a bunch of stuff inside of a bounding box.
There's always going to be that. In geospatial and earth observation, you're always going to have to tell the system where on earth you want to get data from. I'm not necessarily convinced that it always has to be on a map. I'm also not convinced, after spending many, many years watching folks interact with map products, particularly in the earth observation and earth science world, that that's really the best thing. I think people still struggle with it. I think they're still a little inconsistent, and I think we could do better.
Yes, these tools are ubiquitous. But we still have not ironed out all the UX issues of the map-based interface, let alone moving into a new paradigm. I'm not convinced that it's the best interface once we get into that higher level “what question do you want answered?” phase.
D:
Every time I see a new data portal, it feels like it's a total one-size-fits-all without too much thought that's gone into it. It's a map. It's the panel on the left hand side of the screen. It's all the usual things. Part of me thinks seeing standard products, like standard design, things that I know and have used before makes it easy for me. I don't have to learn something new, so I get that side of it. It just feels like this one-size-fits-all product where no one has really come along and thought- Can we do it differently? Is there something else? We even see Google Maps doing the same thing. It feels like maybe it's not a technical problem. Maybe it's more of a cultural problem.
J:
Yeah, absolutely. I don't think it's a technical problem. There have been great strides on the technical side of internet map making. There are all sorts of great tools to make creating layers and getting the map in your browser easier. I certainly would rather be building map interfaces now than map interfaces a decade ago. We've come very, very far in the developer experience with how easy it is for a software engineer to put a map interface together, but I think that's where we've been iterating. We've been iterating on- How can we make this map interface a little bit easier to use? or a little bit easier to develop? or easier to serve information on top of? But I don't think we've ever really stopped and thought, is a map really the right interface for this particular application?
I think that's sort of where I see us going- Maybe a map's not the best interface for this particular piece of software. And if not a map, then what?
D:
I've often heard people get really excited about having an output they can put directly into a spreadsheet, because that's what they are used to working with. Something that’s going to go directly into a database, or talk nicely with other spreadsheets, and they didn't actually want a visual output at all.
J:
I think that's true. You're kind of always thinking look, I can give you this picture of the land area. We always want to put things into a visualization. Doing visualizations is an important part of this., but not all science users want their output in that format. Sometimes they want it output in code. Sometimes they want it output in a spreadsheet. The flexibility in that output is important. Not assuming that your user wants it in any particular format, but making things sort of extensible in a way that they can pull it into whatever format they're most comfortable with is the way to go. Whether that's a statistical programming language or, like you said, just a standard spreadsheet.
Going back to what we were talking about earlier, with having that middle piece built out and matured, that lets us think about those interface problems. We can spend more mental energy and man hours on those types of problems, because we're not having to sort through the lower level building blocks.
UI/UX Standardization in Geospatial D:
MapScaping used to be an eCommerce website. We ran on a platform called Shopify. Are you familiar with Shopify?
J:
Yes, absolutely.
D:
So Shopify showed up and made it incredibly easy for people like me to start an eCommerce website. There is an app environment, so you can just add these plugins, these apps to your Shopify store from a standardized box. You can quickly make a custom thing, and it's absolutely amazing. It's interesting with eCommerce. Their whole thing is to get someone into the shop and out of the shop as quickly as possible. If the checkout takes more than five seconds, it's too long. There's this real focus on “get the customer what they want as quickly as they can”. I don't feel like we've got there yet with geospatial. I'm wondering if there's any lessons that you can see in eCommerce, perhaps from Shopify that we could take over to the products that we build in the geospatial world?
J:
Shopify is a great example, because they've done a handful of things really well. One of the things that I love about Shopify is that they've sort of standardized the user experience patterns around the checkout experience. When you're checking out on a Shopify site, you know that it's a Shopify site. A lot of people probably don't know it's a Shopify site, but they know that they've seen this pattern before. They know they've seen this particular way the credit card has to be entered into the form. They expect when they start to type their address, that a dropdown is going to appear, and they're going to be able to select their address, and it's going to autofill. They don't have to do all that. That experience is seamless and it's sort of expected. The end user knows what they're going to get. I think one of the problems I see with geospatial interfaces now is that you're not always sure what you're going to get.
My one example of this is the way that an area of interest selector, or a bounding box selector on a map interface behaves. Sometimes, you have to click and drag it. Sometimes, you have to click and add points like a polygon. Once you have the area selected, the way in which you perform actions on that selection are also different. I think there are ways that we can standardize that. I think the group, or company, or people that figure out the best user experience will have standardized it in a way that makes it kind of crazy to not use that. That's one thing that Shopify's done really well is you really have to convince someone to not use the Shopify platform for their store. The people that are using Shopify, they're using it because they don't want to deal with the eCommerce aspect. They want to sell their T-shirts. They want to sell their eBooks. They want to sell their information products. They don't want to have to learn a new system to do that, or put anything out there that is going to make it difficult for their customers to buy their product.
I think with geospatial interfaces we need to move to that point. We need to make it so that there's an obvious choice for some subset of these user actions, but not everything. There are certainly people who choose to roll their own eCommerce site for various reasons, and that's always going to exist. I would love to see some of these patterns mature in geospatial user interfaces so that our science users and our municipal users who are trying to get questions answered for their county or state don't have to wonder what they're going to get when they try to pull data to answer a question. They know there's going to be some standard level of interaction there.
There are two sides here. It's great for the developers because they don't have to reinvent that wheel. They don't have to come up with a new way to do a bounding box, or a new way to do a map interface. On the customer side, they don't have to wonder, "Oh my gosh, this is a new geospatial system. What am I going to get here? What is this going to be like?". I think right now, there's still some of that.
A lot of the science users that I talk to, they're certainly apprehensive about new systems, because oftentimes it takes them a long time to become proficient in the one that they're already using. They do not like to have a new system designed by these new designers that just came in. I'm totally putting myself into those shoes. At one point, I was the new designer that came in and I was like, "You guys are doing all this wrong. You need to do this, and this, and this." That kind of freaks people out, especially people whose entire career is based around doing this type of science work. That specific data is very important to them, both to their career and whatever agency they're working for. There's a lot of uncomfortableness with that change. I think some of that is growing pains, and we've got to get through that.
D:
So we talked about Shopify, and one of the beautiful things about it is that establishment of a standard set of patterns. We know what a checkout looks like. We understand this. Three steps, there's autofill forms, there's this, there's that. Then lastly, we push our credit card numbers in. We get that. Do you think the lack of standard patterns in geospatial speaks to how niche it still is?
J:
There's not as many designers working on specific geospatial problems. I think that's certainly changing. I think a lot of the new interfaces that we see coming out from a lot of the new commercial companies that are appearing, they're taking design a lot more seriously. There just hasn't been a lot of people working on these problems.
There obviously have been great designers working on mapping problems for quite a while. Look at Google Maps, or Apple Maps, or Mapbox. I think on the earth science side specifically it gets pretty niche, especially because of the complex science that's going on there. I think the consumer side is a little more mature, but I'm not entirely convinced that the patterns developed on the consumer side necessarily fit the scientific model or even the municipal GIS worker model. Those are different people that need to navigate in a car or look at an aerial imagery of their house versus a user who's trying to work out land use patterns, or monitor sea surface temperature, or things like that. There are different requirements.
Breaking the Status Quo with New UI/UX Functionality in Geospatial D:
I think this gets back to the idea that one size doesn't fit everyone. We need specialized maps or specialized interfaces, depending on what problems we're trying to solve. I want to stay with the Shopify example just for a second here. In our store, we had this option to install a plugin. That plugin was going to provide a little mini recommendation engine- People that like this also like that. I think in the past, a big argument for having the map was, "We can see the data to decide if that's the data we want." It's an incredibly quick way of filtering a massive amount of data visually. I would love to see some sort of recommendation engines come into these portals to help us filter through it. I would like to see some voice interfaces built into these things. From a design perspective, can you see anything that's stopping us from integrating some of these ideas?
J:
I don't think there's anything specifically stopping us from integrating these ideas. I think we are moving towards that in some areas. We have this notion on the earth science side of “usage based discovery”. This is finding data sets based on what you're trying to do. What is the end goal of the data set, and categorizing things in that way.
I think as the metadata systems, that middle piece that we're in the building phase of now, I think as that gets improved we can expand the metadata available for these data products. Then those alternate discovery mechanisms become easier to build.
Part of the problem is that, in order to build those higher level discovery systems, you still have to have the information about the data. You still have to have the metadata about the products. As we're building this middle piece out, those metadata are becoming more refined, more detailed. Again, because we're not dealing with the lower level problems, people have more time to refine and incorporate more specific metadata into these programs. We can start to build smarter systems that sit on top of that data and allow us to have some sort of AI, or machine learning algorithm make recommendations for data that we may not have thought about. Maybe we have an interface that isn't a map, and you're just sort of asking the system a question based on a region of the globe. Then it can suggest other questions related to that. Or it can say, "Hey, these other researchers have posed a similar question. Here are their results." I think there's a ton of space for those types of interfaces once we have the sort of data infrastructure there, but we aren’t quite there yet.
D:
Is there anybody out there at the moment that's doing inspirational work in this space that you can point to and say, "Well, if you want some examples of great design, great interfaces, look over there"?
J:
Yeah. So a couple examples off the top of my head. I think the folks at UP42 are doing some really great design and user experience work, from a product perspective as well. They're sort of an example of the higher level system that's sort of interfacing with a bunch of these other data providers and then providing an interface or a system on top of that to streamline a workflow to get a faster answer to a question. I really like what they're doing.
I also love this little project called Placemark by Tom MacWright. He's doing some really interesting user interface work in a niche space. Unfolded.ai is another really great one. Their mapping and general design interfaces are really, really good. They're built up on top of Uber's H3 geospatial library, which is a hexagon-based geospatial library. This is being pulled out of the consumer, commercial side, and adopted now into more of the data analytics and geospatial side of the house.
Obviously Mapbox, I'm sure everybody knows about Mapbox. They're doing really, really wonderful stuff. They have fantastic designers. Their 3D rendering and map rendering is really amazing. They are importing consumer and commercial tech into the data analytics and geospatial side. Those are all great examples of fantastic design and user experience.
I'm excited because all of these new companies and projects that are cropping up look great. They're all seeing design as a first class citizen. I think we can just continue to build on that momentum and raise that bar a little bit. Then those expectations for where design and user experience need to be for these systems are just going to keep getting better. We're going to see the same trajectory that we've seen on the consumer and commercial internet side. Pick your industry. All of those systems have been improved with designers having a seat at the table.
Design and/or User Experience?D:
You mentioned design and user experience a few different times there. Do you see them as being two different things?
J:
Yeah. This is like an endless debate. I feel like I was just having a conversation on Twitter with somebody about the difference between design, user experience, and user interfaces. If you're a user experience designer, does that mean you also design user interfaces? And if you're only a user interface designer, are you also a user experience designer? I'm not sure if that's exactly the right argument.
I think I see design as the whole of everything. I consider myself a designer. There are a multitude of practices under ‘designer’. I think a designer could be a pure art illustrator. I would call architects designers. I think industrial designers sort of run the gamut.
User experience for me is the entire end-to-end workflow of engaging with a product. Oftentimes that is not necessarily the way the buttons look, or the way the map looks, or the way the interface looks. It's the process by which a user gets through the system. What are the steps that they have to take to get there? What is the language that you're using in the system? What words did you pick to describe different things? What icons did you choose to represent particular things within the interface? It even goes even further than that. What is it like to engage with your company? So if you're selling geospatial data and you have to send an email to a salesman and then jump through a bunch of hoops to get any information on a data product, well that's part of the user experience too. That's the early, early interactions.
Oftentimes at Element 84, when I'm looking at the user experience, I'm not only focused on the UX of the products that we're developing, or the tools that we're building for customers. I'm looking at the user experience of the company itself. What is it like to interact with Element 84 on social media? What is it like to interact with Element 84 when you're engaging as a new customer? Is the contracting goofy, or is it really hard to get a response from someone? Does the person you had to deal with from a project management standpoint, are they treating you well?
Specifically for geospatial and earth observation stuff, the user experience sort of starts at the engagement with the organization that is brokering the data. What's the UX of going to NASA for earth observation data? It starts pretty early on in the cycle, and our user interfaces are only one tiny part of that overall experience.
D:
So it sounds like what you're talking about is the promise. What is the promise here? When I interact with a company, when I interact with a piece of software, what's the promise that is made? And is the promise kept throughout the interaction?
J:
Part of that promise is that you're setting some level of expectation. You need to be careful. It's easy to talk about a particular feature, process, whatever, and set that high expectation. Then you've got to deliver on that throughout the entire user experience of that individual interacting with your system. That system could just be a little piece of software, or a website, or an entire company. I think keeping that promise is important. When you get into complex systems, particularly in earth science, it is very hard to keep that promise all the way through. There are so many places in that process where you can sort of lose your promise. It's really important to be vigilant about that. It’s the hardest part of the job. Making the interface look good and polished honestly is the easy part. The really hard problem that we have to tackle, is really how do we keep that promise from the initial interaction all the way through to I've got the answer to the question I had.
The Future of UX/UI Design in GISD:
I feel like we've covered a lot of this already during the conversation. When we look out into the future, if you had to sum up things into some bullet points, what do you think we can expect from user interfaces when we think about geospatial?
J:
First and foremost, there's going to be kind of a layering over geospatial. I think we're going to worry less and less about the sensor technology, the resolution, all the tech specs of the geospatial data itself. I think we're going to get to a point where we're not going to have to worry about the low level earth observation data. So there's going to be a pretty good obfuscation of that. I think you're going to get to a point where you're going to get transparent sub-setting and re-projection. It's just going to be like having to understand the intricacies of SMS messaging or any of the HTTP layers on the web.
I really think we're going to be done with downloading data. I think the future data sets are going to be enormous, like terabytes, petabytes. Huge, huge data sets. It's not going to make sense to pull that down over even your best gigabit ethernet pipe. Everything is going to be an API layer. Everything is likely going to be sitting with major cloud providers. You could have a whole other discussion on how you feel about centralization of the cloud providers. That's beyond the topic of this conversation, but I think that's where we're headed, at least from what I can see.
I also think particularly on the earth science side, you're going to get into these low code or no code science tools. There are a lot of really great projects going on right now. Pangeo is one of them where they're building a lot of standardized Python tool sets and libraries to sort of deal with earth observation and geospatial data at a code level. Now we don't have to write this low level code anymore. We can write this higher level elegant Python to do some of the data processing. As we move into the future, I think we're even going to get a step above that. We're going to get to the point where I could be a data scientist or an earth scientist and not have to deal with spinning up AWS instances, or making sure I have a particular Python library. Looking 10, 20 years into the future, I think the low code, no code science tools are going to be the big step there.
D:
Do you think the community is going to push for this? Do you think the pull is going to come from outside of the geospatial community? Do you think scientists and geospatial users are wanting this today, is this going to make their lives easier? Or do you think that it's going to be people from outside the community saying, "We would love to use that, but you're going to have to do these five things first"?
J:
I don't think we're quite there yet. I think the pull right now is “Give me better code tools. Give me better Python libraries. Make this whole spinning up cloud instances go away. Get me out of that.” I think we're moving in the right direction with things like Jupyter Notebooks and JupyterHub standardizing and making it much easier to do data processing at scale in the cloud. So I think we're getting there, and I think we are answering the call to improve developer tools, but we aren’t there yet.
I think that the push is going to come from the other end though. A good example of this, the future scientists, think people that are maybe in high school or middle school now that in 15 years, they're going to be taking over desks at NASA, and NOAA, and USGS. They're going to have grown up on low code and no code tools. They're going to have expectations that the science tools are as easy to use or of a similar capacity as what they've seen in other parts of their digital life. I think the push is going to come from that direction.
In order to get the low code or no code science tools, first you've got to have really great coding tools. The developer experience has to be really, really good, then the developers can stop worrying about the lower level code. Then they can start thinking about, "How can I write this piece of software where my scientist doesn't really have to do any coding at all, and they can just be an expert at land use, or they can be an expert oceanographer without also having to be a software engineer Python expert?"
You could think of other commercial interests full of people that don't have any desire to do code. They're going to be pushing for this too. You see this in commercial web products, like website builders, and even low code or no code tools for building mobile apps and all sorts of complex web interfaces. We're getting to the point now where the infrastructure for writing code on the web has gotten so good, that we are thinking one level higher now. We're able to think about how we can build tools that'll grant that same power to an end user that the developer has without them having to be a software developer.
Daniel:
I think this is a really, really great place to round off the conversation. I want to say thank you very much for walking us through this. I've never talked to a designer before on the podcast. I've really enjoyed the discussion. It's opened my eyes to a lot of things.
Before I let you go, if people want to reach out to you, if they want to ask more questions about any of the things we've talked about today, where can they go to do that?
Jeff Siarto:
The best place to get in touch with me is probably Twitter. They can get me @jsiarto. So first initial, last name on Twitter. Or you can also find me @Element84, which is our company's handle. Both of those places are great. DMs are open as they say.