This episode is about the state of the art in how we approach aerial imagery today. From the equipment used to capture the imagery, and how it is processed, to the business model of capturing aerial imagery and what consumers are doing with the data. The guest is from Nearmap, a company that has been doing aerial imagery for over 15 years now.
About The Guest
Michael Bewley is the Senior Director of AI Systems at Nearmap. He runs a team of machine learning engineers and data scientists who use AI (Artificial Intelligence) to turn aerial imagery captures into useful insights and information. In this episode, he shares the state-of-the-art technology that is used at Nearmap to capture and process aerial imagery.
What Is The State Of The Art In Aerial Imagery Today?
With over 15 years in the aerial imagery industry, Nearmap has developed some of the best technologies for capturing and processing aerial imagery. A quick look at the technology they use gives clues to the latest developments in the world of aerial imagery.
Nearmap has adopted a highly automated and systemic approach in how they do aerial imagery, with their most recent development featuring the state-of-the-art HyperCamera3 (HC3) that was launched in 2022.
Designed in an upgrade of specs, the HC3 is efficient and captures high-quality imagery, at a larger scale, without increasing costs. From the imagery, one can clearly identify corrugated roofs and even see grid patterns in solar panels. But the imagery from the HC3 is not only limited to RGB channels, the camera also captures NIR (Near Infrared). With this addition, NIR calculations such as NDVI can be leveraged in not only recognizing the presence of vegetation in the imagery but also in knowing the vegetation’s health status.
What Is The Importance Of AI In Aerial Imagery?
With the amount of imagery captured on the rise and the need to process data fast, relying on human labor to identify features in aerial imagery is impractical. Trained to recognize features in RGB images as a human would, AI models are used to quickly recognize the features present in an aerial image. Nearmap is now at its 5th generation AI system that can recognize over 70 features in HC3 imagery. Featuring a deep learning model with about 78 layers, they have an elegant, state-of-the-art, single model – many outputs solution.
The Importance of Accurate Labelling When Training Models
Models are trained to be able to consistently recognize what a human can recognize from an RGB image. The state of the art for training models is that it is possible to reach that human performance with the right labeling and clean data. Even when using automated methods to train models, the results should still be verified by humans (human expert labeling) because if the model gets it wrong, it is not only a problem for your model training but also for your reporting of how accurate it is.
What Is The Business Model In Aerial Imagery?
Consumers of aerial imagery can be broadly split into two groups – those who consume raw imagery directly and those who are more interested in final map products. Imagery as a service customers use aerial imagery directly in a map browser tool or use APIs to pull the imagery into their own custom applications. On the other hand, there are consumers who prefer a more final product either because they lack the in-house capability to handle raw imagery, they don’t have the time to process it themselves, or it is more cost-effective for them to consume a final product. Vector maps as a service for customers who consume final product maps that have been prepared for them.
Data Fusion and Its Challenges in Machine Learning Datasets
Fusing datasets from different sources is a common practice among providers of machine learning data sets. But this is usually, a problem when trying to examine the truthfulness of the results by comparing the imagery to another source of truth i.e. a satellite image of the same location.
Since the data is a mix of different sources, there is often no way to go back in through the stack and prove where the data actually came from. With concerns about fake imagery that can easily be manufactured, this remains a big challenge for companies who rely on different sources for their machine learning datasets. Additionally, the quality of the final results may vary depending on whether there are errors from either of the sources one is relying on.
What Is The Importance Of Post Disaster AI?
Post-catastrophe imagery is captured just after a disaster such as a fire or a hurricane event. The imagery is then quickly processed by AI models to identify various types of damage that may have occurred. Post-disaster AI is powerful in supporting rapid response operations by governments or charity organizations. Insurers can also leverage the results to know where the claims are likely to be.
Post-disaster AI is not only useful in post-disaster response but can be used as a predictive model as well. A comparison between post-catastrophe imagery of an area and its imagery before the disaster reveals how different things have been affected by a disaster. This information can be used for predictive purposes of what may happen if a particular disaster hits an area.
How to Promote Transparency in Machine Learning Performance
Models have varying performances on different datasets. One of the most challenging parts for machine learning providers is the ability to choose an arbitrary location for running tests. Cherry-picking the best results of a model and showing it to a customer can easily mislead them to infer that is how the model performs everywhere – which is not the case.
To have a clearer picture of a model’s performance on a certain dataset, it is important to use samples that are meaningful for a particular use case. For instance, if it is a local government area, then use samples from within that area. Choosing random locations from within an area of interest promotes transparency in how a model can be expected to perform for that dataset.
The Benefit of Long-Term Technology Development
Building bespoke solutions for customer problems is beneficial in solving specific needs. But problems arise when you want to solve another similar but different problem. In most instances, you may have to rebuild your whole solution to suit the other problem.
However, it would be more beneficial to adopt a long-term technology development initiative, which focuses on a foundational outline of how you can efficiently provide a certain solution at scale. With that foundation in place, you can easily stack in new layers to serve all sorts of different customers without having to rebuild the whole stack.
Stratospheric Balloons As Remote Sensing Platforms
In Conversation
From Bespoke Capture to an Industrial Approach
Daniel: Michael, welcome back to the podcast. Let me set the scene with my own experience. About 10 or 12 years ago at a research centre in Christchurch, New Zealand, we wanted to make a thermal map of the city — so we took a thermal camera out of a BMW, welded it to an aluminium plate, put an IMU beside it, hooked it up to a GPS, flew it over the city and took photos. Where are we today?
Michael: That project sounds like a lot of fun, but it’s quite different to how we approach things. We’ve been capturing aerial imagery for about 15 years now, and we’ve gone from bespoke capture systems to a really industrial, highly systematic, highly automated approach. We do a repeatable coverage program with a whole set of sensors over urban areas — mostly where people live — and we do it again and again, year after year, so business customers and organisations can get eyes on the ground. A lot of those users don’t know anything about sensors or the exciting parts of putting a camera in a plane. They just want to see the pictures, and they want to see how things are changing over time.
HyperCamera 3 and the Near Infrared Channel
Daniel: What does the state-of-the-art camera look like today? Is it still a DSLR bolted to a plate?
Michael: We used to do exactly that — DSLR setups with long zoom lenses. Then HyperCamera 2 launched shortly after I joined Nearmap about five years ago, our first real custom camera system, where it doesn’t look much like a camera anymore. Just a few months ago we started running our HyperCamera 3 system in production. That thing is state-of-the-art — from lenses and sensors to stabilisation systems, all carefully designed to a tight spec to fly efficiently and give high quality, so we can capture larger scale for the same cost. Last time we spoke we were only capturing RGB. Now, in HyperCamera 3, we’ve bolted on a near infrared sensor, so we capture NIR as an additional channel on top of RGB — and of course we do structure from motion to turn the RGB imagery into 3D, so there’s an elevation channel as well. We’re building up the repertoire of different things as customers need them.
Daniel: You previously said you’d focus on better RGB rather than more channels. What’s the thinking behind adding near infrared?
Michael: Near infrared is useful in its own right — a bunch of customers are used to getting things like NDVI calculations. But it also gives us the opportunity to add more information. We get great texture from RGB, but any additional information the near infrared provides, we can start to leverage — looking at things like vegetation health, recognising not just the presence of vegetation but what state it’s in. A lot of our machine learning work has pushed the boundary of what’s possible with 5 cm aerial imagery — not just “is it a shingle roof” but “what’s the condition, is there rust, are shingles missing.” Any little extra quality improvement gets us things we simply can’t do today.
One Model, Many Outputs
Daniel: You used to identify 30 or 40 features, and now it’s 70 or 80. What changed — the new camera, or a better model?
Michael: We’re up to our fifth generation AI system now, and it has about 78 layers. The core of it is a deep learning model doing semantic segmentation, and the one model produces all 78 layers as an output — and it’s the same model that runs in Australia, New Zealand, Canada and the US. It’s a really elegant solution: one model, many outputs. Adding more outputs is really about crafting accurate, meaningful definitions, training the labeling workforce for those new things, and winding it into the same training process. The existing layers improve as well, because where there are commonalities between attributes you get improved performance and better cost efficiency. There’s a wisdom of real production machine learning models — as soon as you release a model it’s old and out of date, so it’s about continually adding and improving your training set.
The Business Model: Imagery, Vector Maps, and Answers
Daniel: What’s the business model? My guess is imagery as a service and vector maps as a service.
Michael: For 15 years we’ve had imagery as a service customers, whether viewing it in our map browser — you don’t need geospatial or software skills for that — or via API into their own applications. Two or three years ago we launched Nearmap AI: vector maps as a service, basically for every pixel of RGB imagery. We produce a zoom-21 web Mercator set of 78 pixel probabilities for each class, run it through a big post-processing engine, and make it available behind a feature API. The new thing alongside Gen 5 is going beyond that — geospatial things are pretty complicated, and while sophisticated users can integrate a payload of GeoJSON, we’re now building things on top to make that information more accessible for customers who want a more direct end solution.
Daniel: Are any of those three products cannibalising the others?
Michael: Not really. People seem to want that integrated stack. Even the customers who want the end results want to see all the way back — every result we give them comes with a little map browser link, so they can look at the picture that generated that data. It’s useful in intangible ways. Someone might have another source of truth and think the Nearmap data is wrong, then look at the picture and realise our data was captured more recently and the property got developed since — so you’re both correct, just at different points in time. It comes down to lineage, trust, and reproducible results — you can dip back into the stack as far as you need to get the best result.
Lineage, Trust, and a Post-Truth World
Daniel: How important is lineage in getting people comfortable with remote inspection, and has anyone ever challenged you, saying “prove you didn’t fake this”?
Michael: That concern is common to anyone producing machine learning datasets — often there’s no way to prove where the data came from, because it came from a variety of sources. But in critical applications, like underwriting a property for insurance, you need to be correct and to know why you made a decision. People challenge a result, go back and look at the image, and can explain it — that explainability from having full provenance back to the source of truth is critical. There’s a fascinating initiative from Adobe working with Leica and Nikon — the Content Authenticity Initiative — around guaranteeing that what was captured on the camera is what comes out the other end. That’s a hard problem when there are multiple vendors. The easiest way is: we build the cameras, run them, and do everything along the way, so we can make some pretty hard guarantees. And it’s not just malicious deepfakes — it’s errors creeping in through complexity, which is very analogous to what geospatial professionals do when data is transformed and handed off again and again.
Confidence Scores and Change Detection
Daniel: How certain do you have to be that a house is there before you present it to a customer?
Michael: We give confidence scores. For a building we give a percentage likelihood of it being a real building — you’d only get a low result for something like a garden shed half-covered by trees. We also give a fidelity score: given it’s a real building, what’s the quality of the outline, because tree overhang might throw the square footage off. We’re transparent about that, so customers can choose their own threshold — you might want to catch every possible result and accept false positives, while someone else cuts off everything below 90% confidence to knock out noise. Change detection is deceptively hard. It’s easy to do a bad job — you take our vector map on two dates and subtract them, but things shift. If meaningful changes occur 1% of the time and your model gets it wrong 1% of the time, you’ll get it wrong about half the time. It’s not hard in concept — it’s the signal-to-noise ratio that kills you, which is why you need to look at multiple dates within your models.
Post-Disaster AI and Predicting the Future
Daniel: What is post-disaster AI?
Michael: It starts with our post-catastrophe imagery capture program, called Impact Response — we put a plane up after a fire or hurricane, often within hours or a few days, and publish the imagery fast. We’ve now fully automated the AI processing, so where we used to promise about a month, most surveys now go through in less than a week. We ran a whole heap of imagery on Hurricane Ian within days of capture, and you get high-level statistics over huge areas — how many square miles of damaged trees, damaged roofs, wreckage. It becomes really powerful for rapid response. And it’s predictive too: if you capture an area before and after a disaster, get AI results from both, and have human experts validate, you build a post-catastrophe loss dataset. From that you can say, given what I know about this building beforehand, what’s the likelihood of it being damaged in a fire or wind event. You can’t get 99% certainty like you can with direct image recognition — there’s a lot more randomness — but a geospatially distributed dataset captured consistently over time is the ideal one to build a good predictive model on.
Should We Change the World for Robots?
Daniel: With self-driving cars, people talk about changing signage and road markings for machines. Do you see us changing the physical world to help AI, or AI just getting better at seeing the world the way humans do?
Michael: It has to be the latter. Look at the Tesla philosophy for autonomous cars — they take the line of using vision primarily. The hypothesis is that because roads and signage are built for humans to navigate visually, if you can’t recognise things as a human does, there’s no way to solve it. You might get 90 or 95 or 99% done, but that’s not good enough — you have to eke out all the corner cases, and you can’t expect every road to be changed. What we do with RGB imagery is recognise whatever a human can recognise consistently, because that’s necessary for human expert labeling. It’s a big ask for the world to change to support robots — I just don’t see it happening to the extent needed.
Daniel: Are we setting the bar too low when we use human ability as the baseline?
Michael: I think we’re setting it too low if we look at a single image. We’ve let people look at multiple pictures since day one, and you get a fascinating side effect with tree overhang — if a human labels where they know a building’s roof to be across multiple dates, they can draw out the area under a tree pretty accurately, and the model learns to guess how far the building extends under the tree even though you can’t see it in that one image. Worrying about accuracy against what can be seen in an image is theoretical. The real baseline that matters is accuracy against what’s physically present on the ground — your customers don’t give you prizes for saying “it was hard to spot in that image.”
Demos Versus Production Systems
Daniel: What do you wish people like me understood about AI?
Michael: The thing that would make my life easier is if people understood the difference between a demo or early-stage proof of concept and what it takes to build systems that operate consistently at scale over time. It’s very easy to pass off an early-stage demo as “we’re just one little step away from a product,” and that gap is at least an order of magnitude of work. The biggest thing to look for in a geospatial context is the ability to choose an arbitrary location. Every slide deck cherry-picks results — the beautiful house picked up perfectly, the nice bit of roof damage. You see two or three examples and infer that’s how it performs everywhere, which isn’t the case. We’re really big on giving people samples meaningful to them — if it’s a local government, give them a sample from their area; if it’s an insurer, let them choose a random subset of locations from their policy book. That’s the only fair way to do it.
Daniel: Could you have skipped the vector maps and gone straight to providing answers?
Michael: You can, and I’ve seen people do it — you start with the customer and cobble together a system that solves their problem. That works well initially, but when you want to do it again you pretty much have to rebuild your whole stack, and again the next time. If you want to solve a lot of those solutions efficiently, you can’t skip the foundation. We had the fortune of being in a company that understood long-term technology development — you don’t build a camera system overnight, and you don’t build an AI system overnight. We were given the license to build those core foundations correctly, and now we can pack in new layers for all sorts of customers. The national tree study was relatively easy to do as an add-on. Building the foundation lets you do that final step with much greater efficiency — that’s what I’m really excited about.



