Jeffrey Martin is the co-founder and CEO of Mosaic, a company building 360-degree camera systems designed specifically for mapping. He’s been obsessed with 360 imagery for over 20 years — he built one of the first websites combining panoramic images with a map back in 2005, before Google Street View existed, and he holds a Guinness World Record for a 320-gigapixel image of London stitched together from 52,000 photos.
In this episode, we get into what a modern mapping-grade 360 camera actually looks like, and why the difference between rolling shutter and global shutter sensors matters if you care about things like colorizing point clouds or photogrammetry. We also cover the surprisingly long tail of people who need up-to-date street-level imagery — everyone from departments of transport and utilities companies to playground designers and outdoor furniture salespeople.
A few things that stood out:
- Ground-level 360 capture fills gaps that drones and satellites simply can’t — occlusion, permissions, and viewing angle all favor eye-level imagery for certain infrastructure work.
- Companies are taking very different bets on data collection strategy — Mosaic focuses on high-quality, purpose-built capture, while others like Hive Mapper are betting on scale and crowdsourced coverage instead.
- The camera hardware race may be plateauing, but the real frontier now is what we do with the imagery once it’s collected — especially as large language models start being layered on top of geospatial data.
Where does this all go next? If AI can eventually parse and reason over an entire country’s worth of street-level imagery, what does that unlock — and who ends up owning that layer of the map?
In Conversation
Twenty Years of 360 Imagery
Daniel: Hi Jeff, welcome to the podcast. Thank you for showing up on such short notice. You are the founder of something called Mosaic, and we’re going to talk about that a lot more later on. I want the focus of this to be around 360-degree imagery — what we can do with it. And bear in mind, this is a geospatial podcast, so there’ll be a focus on that side of things as well. That’s a very brief introduction. Maybe you could fill in a few details for us. Who are you? What do you do?
Jeffrey: My name’s Jeffrey Martin. I’m originally from the US, living in Prague now, and I’m a co-founder and CEO of Mosaic. We make geospatial 360-degree camera systems for mapping. I’ve been involved in various aspects of 360-degree imaging for over 20 years now.
Jeffrey: My first foray into that resulted in building one of the first websites combining 360-degree images with a map. This was back in like 2005, so there was no Google Maps at all. There was no Google Street View — it was still a few years away. Then I built a company around that called 360cities.net, which is still online. It’s a publishing and licensing platform for artistic, high-resolution 360-degree images.
Jeffrey: After that I started a company called SphereCam, around 2016. If you remember the VR hype train of 2016 — Oculus getting bought by Facebook, the whole world going crazy over VR, lots of VR headsets coming out, lots of VR cameras coming out — I was in the middle of that. That didn’t end up panning out the way people thought it would, and I kind of pivoted SphereCam into Mosaic. So going from a VR 360-degree camera for consumers to a more industry-oriented, focused use case of 360-degree imaging, specifically for mapping. And that’s where I am today.
A Guinness World Record Over London
Daniel: Just before we pushed the record button, there was some mention of the Guinness Book of World Records.
Jeffrey: Oh, that’s right. I’ve been generally sort of obsessed with 360-degree imaging for a long time, just trying to push it as far as I can. Back in the day you would make these images by using a DSLR and taking a bunch of pictures in a circle and stitching them together. You still can do that — people still do that. So I kind of pushed it in two directions. One is, how can you make a video out of this? Cobbling some kind of video cameras together. The other way to push it is, how high resolution can you make this?
Jeffrey: So instead of taking just four pictures in a circle with a fisheye lens, you can put a zoom lens on your camera and take more pictures. You take 50, a hundred, a thousand, 50,000. So I started making these gigapixel images — there’s a megapixel, which is a million pixels; a gigapixel is a thousand megapixels. I started shooting from the top of various buildings in the world and got a few interesting commissions. Then BT, British Telecom, approached us and said, “Hey, we’d like to do something for the Olympics — how about we make a world record?” So we were like, okay.
Jeffrey: From the top of BT Tower in London, we were up there for about three days and got multiple data sets comprising 52,000 photos, and stitched them all together into a 320-gigapixel image of London, which is still online. You can find it if you Google “London gigapixel world record” or something like that. We also did some gigapixel images inside stadiums, like during a football match at Wembley, the FA Cup final, where you could zoom in and see every single face. Sort of like a fan camera, where someone looked at the match and then they could go back later and find themselves and share it on social media. I’ve done all kinds of different types of 360 imaging.
Daniel: 52,000 images — that’s a lot.
Jeffrey: It took quite a while. Took a few months.
What a Mapping-Grade 360 Camera Actually Looks Like
Daniel: And now here we are with Mosaic — the company you are a co-founder of, making mapping 360 cameras. What does the camera look like? We talked a little bit about resolution before. Can you give me the details of the camera?
Jeffrey: The camera is the size of your head. It’s water resistant, actively cooled. It has six cameras on it — five cameras going around in a circle and one camera pointing up. There’s an antenna for the GNSS, and cables going into your car. You can control the camera with any device that has a web browser. Generally the camera’s mounted about a meter above the roof of the car, give or take.
Jeffrey: It allows you to have a decently high resolution — not a gigapixel, but about 72 megapixels. There are six 12-megapixel camera sensors in the device. They’re all perfectly synchronized all the time, and it has been quite diligently engineered to have sharpness across the lenses and good balance of colors. We’ve tried to make it usable and easy to use for this specific type of activity.
Daniel: And when we say camera, are these RGB cameras? RGB images it’s creating, or something else?
Jeffrey: Yes, they are RGB images. We have two distinct models. One uses rolling shutter sensors — that’s the more affordable option. The other uses global shutter sensors, with bigger sensors and bigger optics. Camera sensors come in two general flavors: rolling shutter and global shutter. Global shutter means all the pixels are captured at the exact same moment. This is more difficult and more expensive as a product. It has certain trade-offs in terms of noise and other stuff, which is why it ends up being more expensive.
Jeffrey: Rolling shutter sensors read out the pixels line by line. This is what the vast majority of image sensors use — all smartphones and all mirrorless cameras. Generally you have a mechanical shutter in front of the rolling shutter sensor in the case of a professional mirrorless camera. Rolling shutter sensors generally have better dynamic range and lower noise, but as they read out the image line by line, there can be some geometric distortion. You might have seen this if you pan a camera really quickly — the vertical lines become sort of bent and no longer vertical. And for a camera that is being used eight hours a day every day, you can’t have a mechanical shutter in front of the sensor, because it will wear out very quickly.
Jeffrey: So if you want to have geometric fidelity, you need to use global shutter. Our Mosaic X, which uses global shutter sensors, is more suitable if you want to do things like colorizing point clouds or photogrammetry, doing measurements, things like that. Our more affordable camera, the Mosaic 51, uses rolling shutter sensors. If you just want Google Street View-style panoramic images, that’s perfectly sufficient. If you’re driving really fast then some vertical lines might not be vertical, but that doesn’t matter much if your goal is just to have images.
Daniel: When you talk about point clouds, are these point clouds created from the images? We’re not talking about lidar?
Jeffrey: Generally I’m talking about lidar. We build the cameras, we also do a lot of integrations, and we have one product of our own that combines the camera with lidar. We have a lot of customers who use RIEGL, Trimble and similar sensors, so we’ve put a lot of work into making these integrations flawless. All of the synchronization is very accurate and precise, the triggering is all working, so that every laser pulse emitted by the laser profilers or lidars is exactly known, and we know exactly which pixel from our camera can be used to give that laser point a correct color.
Daniel: I can’t even imagine the engineering that went into figuring all that out and making it work consistently over time.
Jeffrey: It definitely took a while, and we’re quite proud of the fact that it all works really well.
Who Actually Needs Street-Level Imagery
Daniel: So what are people doing with these cameras? We talked about it being specifically designed for mapping, but maps are a lot of different things. They’re collecting this amazing 360 imagery — are they creating structure from motion? Is it building 3D scenes out of it? Is it a Google Street View-style thing where they’re just looking at images and clicking through them? Or are they extracting objects from it?
Jeffrey: I would say the bulk of our customers are doing the somewhat underappreciated task of making sure the world doesn’t fall apart. They are documenting the built world, for or on behalf of local governments, departments of transportation, utilities companies, telecommunications companies. All of our built world — all of our roads, street signs, utility poles, electricity pylons, power lines — all of this needs to be maintained, and it would be good to fix it before it falls apart. It would be good to trim the vegetation before it hits a power line and causes a forest fire.
Jeffrey: Generally these are people doing inspections — a lot of them are surveying or engineering companies. There are still a lot of people walking around with clipboards checking the encroachment of vegetation onto power lines. There’s still a lot of people inspecting stuff by hand. This obviously takes a long time, it’s labor intensive, and requiring someone who is an expert to do this labor-intensive task is inefficient. So it would be more efficient to capture the imagery wholesale and in some automated way. Most of our customers are inspecting the roads, inspecting the telecommunications or power infrastructure, and things like that. Other customers are making entire maps.
Jeffrey: We have a customer in Europe who is capturing the entire country every two years. They do it better, more often and more thoroughly than Google, who only goes there every few years and only captures about 75% of the roads. They get about 98% of the roads and they do it every two years. Their main customer is a government agency who’s able to check that properties are what they say they are, that there haven’t been any illegal additions or construction going on. But they have a lot of other customers who can have access to this countrywide set of images that is up to date, as opposed to Google Street View, which is often not very up to date.
Jeffrey: For example, they have a customer who makes outdoor furniture, and they can check all the restaurants and see who’s got old furniture, and they can call them up and say, “Hey, you might need some new outdoor furniture.”
Daniel: You are kidding.
Jeffrey: There’s quite a long tail of different types of people who want to access street view imagery, especially if it’s up to date. There’s a very, very long tail of people who you might not expect that can get their jobs done more easily by having access to this type of data.
Daniel: What’s the strangest use case you’ve heard of?
Jeffrey: Like these outdoor furniture people, or playground designers. If you’re designing and selling playgrounds, how can you get a handle on what is out there? Well, it’s kind of hard. You don’t want to drive around forever looking for playgrounds, right? And then we’ve done some data collection for movie studios — this is more like 3D reconstruction and photogrammetry for various movies. We worked on a movie called The Gray Man, which was a Netflix big budget movie, a Bollywood movie called War 2, a sequel to War, and a short film called The Ice Cream Man.
Daniel: Some of those use cases involve an incredible volume of imagery you need to look through. Think about satellite imagery — there’s so much of it coming out of the sky that the vast majority is never, ever seen by human eyes. It’s computers looking over it, deciding that this tree is too close to the power line, let’s do something about it; here is a crack in the road. So we’re identifying features we’re interested in, extracting them and creating a line, a point, a polygon in some other database somewhere so someone can go and look at that particular thing. Is that the kind of thing you are involved with as well, or are you simply capturing the imagery?
Jeffrey: Right now we are mainly involved in the data collection, and then this detection and extraction of features is downstream from us. Currently our customers use various types of tools to extract features. They use TopoDOT, for example, or various AI automated point cloud classification software, to classify their point clouds or images.
Daniel: Mapillary?
Jeffrey: There are customers who upload their data to Mapillary. The city of Detroit has been one of the bigger power users of Mapillary. They’ve been using some kind of Trimble device and they’ve recently purchased one of our systems as well, and they’ve been mapping the whole city of Detroit, I think every year or so. They put it all on Mapillary. It’s a good proposition for certain types of users who don’t want to deal with hosting this data, which can be substantial, and where they might have a requirement or directive for their data to actually be public or accessible. Mapillary is like free hosting, and the data is available. That’s either a strength or a weakness depending on what type of user you are. But if you’re that type of user, then it’s a really great value proposition.
Daniel: And also the object detection and extraction they do is pretty fascinating. They’re improving it all the time. For governments, I think it’s probably a really good idea.
Jeffrey: It is a really good idea, and they do a very good job — you’re correct — with their object detection and extraction. It’s something they’ve been working on for many years. It’s very good quality. They have lots of categories, lots of different geographic regions, so the actual quality of the training is very good. And yes, it’s a very good deal, if you don’t mind your data being available.
Where Ground-Level Capture Beats Drones and Satellites
Daniel: What is the sweet spot for your style of imagery collection? The reason I’m asking is, we’re talking about trees hanging over power lines that could cause a fire. I’ve talked to a few companies working on doing this via satellite. I’ve seen other companies trying to solve this particular problem with drones. But I could imagine an urban setting might be a place where drones would struggle and where satellites would struggle. Could that be a sweet spot for you at Mosaic, or is there some other spot that is the category you own?
Jeffrey: You’re right about drones and satellite imagery. With satellite imagery, obviously the angle where it’s looking is fixed. It’s always getting better, but it will always be limited by the laws of physics. Drone data can often be much, much better than satellite imagery, but it still has the same issue where it is up there, and there are only certain directions where it can look. If there is some occlusion, if there are trees overhanging, it can only see what it can see. So it does have some limitations. And of course there’s the fact that getting permission to fly a drone is often difficult or impossible. Drone imagery and aerial imagery absolutely has its place — it’s a huge industry — but it has some hard limitations.
Jeffrey: So that’s where we come in. We capture our images at the human level. It’s where the people are, it’s where the infrastructure is, and it’s not looking straight down, it’s looking straight ahead and upwards, so it can see things that are just not possible ever to see using drones or aerial imagery. And you can combine these things. You can fly a drone and use our cameras at the same time, and then you can turn this data into a truly stunning, amazing quality 3D model. We’re working on improving this process all the time, and this combination is special. It’s not easy, but when you get it right, it’s really amazing.
Backpacks, Bikes, and Indoor Mapping
Daniel: Well, I look forward to seeing that. Is anyone just putting a camera on their head, carrying it around, going through pedestrian areas? We talked about it being mounted on a car, and I can see why you’d want to do that, but I’m thinking about mapping at this kind of resolution indoors, or in pedestrian areas, bike paths, that kind of thing.
Jeffrey: Our main products are the cameras that you put on a car, for large-scale mapping of large areas. This is clearly the most efficient way to do it. But we also have a backpack product called the Explorer. This also uses global shutter sensors — it’s similar to the camera you put on the car. It adds a couple of lidars, so on your back you have 72 very sharp megapixels of RGB camera data, two lidars, a GPS antenna and an IMU, and this allows you to go anywhere. You can go on foot or on bicycle or on a scooter. You can go outdoors, you can go indoors.
Jeffrey: As far as outdoor mapping goes, there are countries like India, for example, that have a lot of narrow streets where driving a car might be inconvenient or difficult or dangerous, and it’s much more sensible to mount something onto a motorcycle or scooter or bicycle and map these narrow passageways that way. Indoors is a little bit different — that’s where we get into some other industries than pure geospatial. We’ve got facility management, construction, industrial documentation of warehouses, factories, power plants, oil and gas facilities.
Jeffrey: Like the geospatial world, you need to be sure things aren’t falling apart, or that they’re as they say they are. Especially oil and gas — you’ve got oil refineries, you’ve got all kinds of facilities that are very complex and very prone to different types of failures. There’s a lot of effort put into inspecting these places very, very carefully, to be sure that there’s no corrosion and there’s not going to be any failure, because in those cases any downtime can be extraordinarily expensive.
Autonomous Vehicles and the Hive Mapper Question
Daniel: I’m wondering — if we start thinking about how we build maps for autonomous vehicles, especially that last mile problem, it seems to me the kind of imagery you are capturing, and the kind of detail you’re capturing it in, and the maps you could build off that, would be fantastic for solving that kind of problem. Giving these robots a base map to move around in.
Jeffrey: Yes, and this is a topic I’ve been diving into in my free time — the whole topic of what sort of framework a robot or embodied mind needs to use in order to function properly and safely in the world. You’re correct: the current crop of autonomous cars are very heavily dependent on having available some kind of reference by which they can judge what’s going on. You have these Waymo cars, or these Tesla cars that are crashing all the time, trying to sense where they are and what’s going on, and they all rely on a reference map that was created using cameras and lasers, for them to understand where they are and what they should do.
Jeffrey: I think this is still being solved in a fairly primitive way. It’s starting to work — you can see that Waymo is succeeding, although crucially they’re not saying how often humans are helping. And the general mechanism by which these cars are operating, I think it still needs a lot of work for them to operate safely in an unfamiliar environment. These Waymo cars are in American cities with good weather, where the drivers aren’t too crazy. Put a Waymo car in Southern Italy or in Delhi and it’s just going to stop and kind of wait for the street to clear out or something.
Daniel: Speaking of self-driving cars or autonomous vehicles — slightly different topic — have you come across a company called Hive Mapper?
Jeffrey: I have, yes.
Daniel: I’ve been talking to them on and off for years now. They originally tried to do what they’re doing with 360 cameras, and they just could not get it to work. They ended up building their own hardware, and that’s what they use now, but they could not get it to work with 360. I think they’ve got two cameras at the front now, so they get more depth. It’s an interesting approach — it’s kind of decentralized, this idea of building really detailed maps. And the reason I bring it up at this stage in the conversation is because part of what they’re doing, the big focus, is not so much documenting the world, like the infrastructure. It’s building maps for autonomous vehicles.
Jeffrey: Hive Mapper is a little bit similar to Mapillary in terms of it being this kind of large-scale type of data collection using more off-the-shelf types of cameras. So the result is that you have a lot of imagery. It might not be as comprehensive or high quality, in terms of the image quality, as a dedicated system.
Daniel: When I look at the resolutions of your cameras, and I think about the way you’ve just explained it, you’ve definitely got a higher quality product. But I think there are two different approaches as well. You’re going for the high quality product, and they’re going for “how many images can we get.”
Jeffrey: That’s right. And I guess what they’re doing — there are cameras and cars that are driving around anyway. One of the practical issues with that in terms of mapping is that you’ve got this kind of power law where most of the cars are going on only a few streets, and most of the streets you don’t ever drive on. Even in your own neighborhood, most of the streets you just never go down. You never have a need to go down. So to get a comprehensive map, it’s always kind of tempting to mount cameras on taxis and stuff like that.
Jeffrey: I guess Hive Mapper gets around that by offering some kind of cryptocurrency for people to cover areas that are needed. So yeah, it’s a very interesting take on mapping. Generally our thesis is that we want to do the best possible quality. I mean, there’s a market for both ways, and there are certainly benefits to capturing a few of the main roads very frequently.
LLMs, Foundation Models, and Life After the Megapixel Race
Daniel: How do you think your business, your technology, is going to change over the next five years? AI is changing everything. What do you think Mosaic is going to look like in five years’ time, either in terms of technology or perhaps who is using it? My guess is that more technical systems like yours are going to be more available, more accessible, to more people.
Jeffrey: I think generally the path of technology is that it becomes easier and more accessible and cheaper, and when it does, you get a bunch of interesting use cases out of it that you didn’t expect or imagine. I know that with Mapillary, one of their customers — I don’t know if it’s Washington, DC or which city — is the leaf clearing trucks, and they can keep track of what streets need their leaves cleared. This is kind of funny.
Jeffrey: In five years I think there will be a lot of interesting ways that large language models are used with the data. The whole world is still figuring out what to do with this technology. The first thing is a chatbot. The second thing is vibe coding, programming. But I think there are a lot of other interesting things that we will figure out what to do with. And I’m already seeing people who are somehow connecting a database to an LLM and then you can ask it questions.
Jeffrey: The challenge with this type of image and geographic data is, how do you parse this data so it can be ingested by a large language model in a coherent and effective way, in order for you to be able to ask questions? That is not an easy question to answer, and I think it’s going to take a while, and a lot of different people, to figure out the best way to do that for whoever the target customer is. I can imagine — imagine you map a whole city or country, and then, if you can understand what’s in the images, you can have an interface for finding things out in a way that is much more easy and human-centric.
Daniel: In the satellite world, the earth observation world, there’s a lot of focus, a lot of excitement, around these foundation models now. The whole idea is that you take all of the satellite archives and just boil them down until you get a bunch of vectors. And then you can use this to say — exactly what you’re saying — show me all the pieces of, I don’t know, beach in this country. Show me the forest. Where’s the forest? That kind of thing. Massively oversimplifying, but it seems to me that a company like yours would be in a position — that country you’re talking about that’s using your equipment to map all the streets in their country every two years, they would be in a position with that kind of data archive to create models for their country.
Jeffrey: That’s where things get interesting and exciting, and I think it will take a few years. Everybody’s still sort of figuring out and dipping their toes into this whole idea. The idea didn’t exist five years ago, and now the idea exists, but how to execute that in the best possible way is just, I think, a trial and error process.
Daniel: Here’s another future gazing question, if you don’t mind. When I think about cameras, the things I see in terms of the technology over time are that they’ve gotten a lot smaller, significantly smaller; the resolution has increased; they’ve gotten a lot cheaper; and they have magic in them now, AI magic, so everyone’s a great photographer. Where do you see this going in terms of just cameras? Is it just smaller, cheaper, faster, more of that? Do we continue along more magic, more AI in them? Or is there some sort of technology jump on the horizon that you can see that’s going to change all this?
Jeffrey: I’m very bullish on the future and I love future gazing, but I think with cameras we are approaching a plateau in terms of physics. You have smartphone sensors that are 200 megapixels, and each individual pixel is smaller than the wavelength of light. That’s a physics problem that they’re solving, but you can’t keep shrinking it down, and you can’t keep shrinking down the optics — at least with the current technologies that we have to create optics. There will be others.
Jeffrey: So in terms of pure resolution in megapixels, there’s no longer really much of a megapixel race like there was 5, 10, 15 years ago, where every year or two there’s a new camera coming out with more megapixels. It’s not really happening so much now. Smartphone photos are still about 12 megapixels. Even if you have a smartphone sensor that these days is like 48 megapixels, those are being shrunk down into 12-megapixel images with better colors and better dynamic range. The actual quantity of pixels that people need seems to be about 12 megapixels, for cameras that people have in their pocket for taking photos.
Jeffrey: I think the evolution now with cameras is not the revolution, but the quality of the image. What is the brightest and darkest stuff you can have in the same picture and still make it look good? All these kinds of HDR tricks, and yes, these computational photography tricks that are being done. So the same 12-megapixel photo that you’ve got on Google Photos from your smartphone camera of today looks dramatically better than the 12-megapixel photo that you took 12 years ago. The same number of pixels, but it does look a lot better. So the resolution has sort of plateaued. The image quality I think can continue to improve, but I think we’re kind of hitting the limit — just like in the film world before digital cameras, camera film also sort of reached a limit of what could be created. They incrementally improved it over the years, but at a certain point it was as good as it could get and you can’t really make it much better.
Daniel: One more question and then I’ll let you go. At the moment you have RGB sensors in your cameras. Is there any need — is anyone saying, “Hey, we’d really like some other kinds of sensors here, we’d like to be able to see other stuff” — or is RGB fine?
Jeffrey: I think the most frequent request we’ve gotten is for thermal cameras. If you can see where places are hot or cold — if you can see the roofs of houses, or certain parts of the power line infrastructure like transformers, if you can see which ones are hotter than other things — there have been some requests for that. So for thermal imaging, yes. Not for other areas of the spectrum, not really, like UV. For vegetation type stuff, this is better handled by top down aerial cameras or drones. But certainly adding some thermal cameras would be cool for certain folks. Not everybody.
Daniel: That would make for some really interesting data to look at.
Jeffrey: It’s always interesting to look at the whole electromagnetic spectrum and see what our eyes can see — and it’s not very much. There’s a lot more out there that we can’t see at all, and there would be something to being able to see more of that stuff.
Daniel: There’s something about thermal. I was at a sports store the other day looking at thermal binoculars. The sales rep who was showing me them said, “Oh, watch this. Look at my feet.” And he walked across the carpet, and he was wearing shoes — I could see where his foot was touching the carpet. I could also see, as his foot came down, the thermal effect intensifying until it touched the carpet, and also when it left on the other side. And it stayed there for quite some time.
Jeffrey: Wow. And he was wearing shoes, huh?
Daniel: Yeah. I don’t know how much magic is happening there in the background, but it was pretty cool. And like you were saying, there’s so much we can’t see — we’re really restricted to quite a limited band. Thermal is fascinating.
Daniel: Hey, Jeff, I want to thank you very much for your time. I appreciate you coming on and teaching us something about 360 imagery. Best of luck with the future of imaging, whatever it might look like.
Jeffrey: Thank you so much for having me.





