From Data Dump to Data Product

Roughly twice a month. Unsubscribe any time.

This conversation with Jed Sundwall, Executive Director of Radiant Earth, starts with a simple but crucial distinction: the difference between data and data products. And that distinction matters more than you might think.

We dig into why so many open data portals feel like someone just threw up a bunch of files and called it a day. Sure, the data’s technically “open,” but is it actually useful? Jed argues we need to be way more precise with our language and intentional about what we’re building.

A data product has documentation, clear licensing, consistent formatting, customer support, and most importantly — it’ll actually be there tomorrow.

From there, we explore Source Cooperative, which Jed describes as “object storage for people who should never log into a cloud console.” It’s designed to be invisible infrastructure — the kind you take for granted because it just works. We talk about cloud native concepts, why object storage matters, and what it really means to think like a product manager when publishing data.

The conversation also touches on sustainability — both the financial kind (how do you keep data products alive for 50 years?) and the cultural kind (why do we need organizations designed for the 21st century, not the 20th?). Jed introduces this idea of “gazelles” — smaller, lighter-weight institutions that can move together and actually get things done.

We wrap up talking about why shared understanding matters more than ever, and why making data easier to access and use might be one of the most important things we can do right now.


In Conversation

Radiant Earth and the Mission of Shared Understanding

Daniel: You’re the executive director of Radiant Earth. Can you introduce yourself and what Radiant Earth does?

Jed: Radiant Earth is a 501c3 nonprofit based in the US. Our stated mission is to increase shared understanding of our world by making data easier to access and use. We have two flagship initiatives: Source Cooperative, a cooperatively governed data publishing utility built on object storage, and the Cloud-Native Geospatial Forum, which brings together geospatial data practitioners — from enterprises to governments to startups — to talk about new ways of sharing data at scale. My background is actually in policy, not technology. I’ve had a career working on open data and science policy, including building the open data program at AWS before Radiant Earth.

Daniel: Why does shared understanding of the world matter so much?

Jed: We can’t cooperate on a pandemic if we don’t understand its contours. We can’t cooperate on climate change if we don’t share an understanding of the observable state of the world. That’s where we operate: can we improve access to data so that we have a shared understanding of things, and from that, cooperate more effectively as a species? The geospatial community has something rare and valuable to offer the rest of the world — decades of experience producing truly global data products. That’s exceptional, and it’s worth protecting.

Data vs. Data Products — Why Language Matters

Daniel: You’ve written about the difference between data and data products. Why does that distinction matter?

Jed: For too long we’ve been writing policies and having conversations about “data” and “the value of data” — as if data is intrinsically useful regardless of how it’s packaged. What we end up with is researchers getting grants to “produce data” and government agencies setting up “open data portals” that look like someone just threw up a bunch of files. Technically open, practically useless. “Thou shalt open your data” policies often just produce compliance exercises — the data goes on a portal, the job is declared done, no one thinks about whether anyone can actually use it.

Jed: Using the word “product” forces you to ask concrete questions: What is this for? Who is it for? How much does it cost to produce and maintain? Will it be here tomorrow? A product has documentation, clear licensing, consistent formatting, and someone accountable for it. I learned this from working with people at NOAA — they’ve been talking about “data products” for years, and it’s a useful mental trick that changes how you think about what you’re building.

What Great Data Products Look Like

Daniel: Can you give us examples of great data products?

Jed: Landsat is extraordinary — the audacity of launching it 50 years ago and committing to imaging the entire planet continuously. The decision to open up the data, championed by Barbara Ryan, transformed what was possible. Google Earth Engine made it accessible to a whole generation of researchers. And the fact that it’s still running, still producing data after 50 years, is itself remarkable. Common Crawl is another — one nonprofit, effectively started by one person, that scrapes the entire web and shares it openly. Common Crawl is the training foundation for almost every large language model ever built. Two people at a nonprofit changed the world.

Daniel: What makes those great beyond just the data itself?

Jed: Consistent availability, clear licensing, known formats, documented provenance, and a credible answer to the question “will this be here tomorrow?” A data dump doesn’t answer those questions. A product does — or at least tries to. And when you add in streamability, good metadata, and discoverability, you have something people can actually build on.

Source Cooperative: Object Storage for Everyone

Daniel: Tell me about Source Cooperative. How does it connect to all of this?

Jed: Source Cooperative is object storage for people who should never log into a cloud console. A lot of people don’t need to create an AWS account and navigate hundreds of services. We provide a clear path to publish data to S3 — or any object store — with a drag-and-drop interface. Everything is organized as products: user accounts or organizations publish products, which are collections of objects with readme files, very similar to how GitHub works for code. NASA has products on Source Cooperative. The idea is that if you have data and want it on the web, there’s a simple, stable path to get it there.

Daniel: Does putting data into Source Cooperative automatically make it a good product?

Jed: Absolutely not — and that’s the honest answer. Some publishers produce excellent, well-documented, perfectly structured data products. Others just dump files. We don’t police data quality. What we’re trying to create is a culture of people striving for excellence. GitHub is a good analogy: a repo is just a repo, it could be terribly documented. But if you’re thoughtful about how you present it, it gains traction and becomes something people rely on.

Cloud-Native Architecture and Why It Matters

Daniel: Can you explain cloud-native and object storage for people who aren’t familiar?

Jed: Object storage is the most basic cloud service: you put a file in the cloud and it lives there. Amazon S3 is the most familiar, but “S3-compatible” has become a standard term. Cloud-native goes further — it’s about designing data so it can be streamed, queried in parts, or accessed at scale without downloading everything first. A cloud-optimized GeoTIFF lets you fetch just a specific tile. Zarr lets you query multidimensional array subsets. STAC provides the catalog layer so you can search across millions of objects. Together, these let you work with very large datasets without moving them around.

Daniel: It sounds like product thinking applied to data format — who will read this file? What will they do with it? Choose the format for the reader, not the writer.

Jed: Exactly right. Cloud-native forces you to think that way. That’s one of the reasons I find these format discussions so interesting — they’re really product discussions in disguise.

Long-Term Sustainability and the Endowment Concept

Daniel: How do you answer the question “will this be here tomorrow?” What’s the business model?

Jed: Right now we’re a nonprofit funded by grants, with cloud hosting covered by agreements with AWS and Azure. That gets us far, but it doesn’t solve the long-term problem. I’d like to move toward a model where publishers pay to host data — a flat monthly or annual fee per terabyte, for example. We control bandwidth so individuals browsing get free access while large-scale programmatic users can be throttled. Beyond that, I’m very interested in raising an endowment — a foundational fund covering operational costs so publishers don’t have to worry whether Source will exist in five years.

Jed: There’s also a model where specific data products can be endowed by philanthropists or institutions. Imagine a data product adopted by a foundation and made available for 50 years — with their name attached. People care about legacy, and if that motivates long-term funding of critical data infrastructure, I’ll take it.

Gazelles: Designing Institutions for the 21st Century

Daniel: You wrote about “gazelles” as a new kind of institution. What are they?

Jed: The 20th century gave us big, heavy institutions — the World Bank, the United Nations — designed to go toe-to-toe with nation states. They’re important, but they’re slow and expensive. A gazelle is something different: smaller, sometimes ephemeral, moving quickly as part of a herd. Common Crawl is a perfect example — one nonprofit, a staff of two, created by essentially one person with resources, that changed the world by becoming the training corpus for most large language models. We can create these kinds of lightweight, internet-native entities deliberately, and they can have outsized impact. Lots of people should do lots of things like this.

Model Outputs as the Next Frontier in Data Products

Daniel: Do you see AI model outputs — embeddings, predictions — as the next wave of data products?

Jed: I think you’re onto something. Radiant Earth has been sharing model outputs and this coming year we’re focusing on figuring out standards for sharing embeddings. The potential is significant: instead of working with vast satellite imagery archives directly, you query a distilled set of embeddings on a laptop and get meaningful results. It dramatically lowers the cost of asking questions of the data. We’re still in early days — we don’t fully know how well these embeddings correspond to real-world ground truth — but the early results are exciting. Heard it here first: we’re going to be working on standards for sharing embeddings and I think it’s going to be an important frontier for the community.

About the Author
I'm Daniel O'Donohue, the voice and creator behind The MapScaping Podcast ( A podcast for the geospatial community ). With a professional background as a geospatial specialist, I've spent years harnessing the power of spatial to unravel the complexities of our world, one layer at a time.