46 episódios
- In this episode I sat down with Joe Redmond from the Allen Institute for AI (AI2) to discuss OlmoEarth, AI2's open geospatial foundation model for Earth observation. Joe explains how the project emerged from AI2's environmental and climate initiatives, where partners needed practical tools for analysing satellite imagery across applications such as agriculture, wildfire risk, ecosystem mapping, and conservation. We discuss the unique challenges of remote sensing data, including its temporal and multispectral nature, why geospatial machine learning differs from traditional computer vision, and AI2's philosophy of building open models and tools that can be adapted to real-world environmental problems.A major focus of the conversation is Latent MIM Lite, OlmoEarth's self-supervised pretraining approach. Joe explains how the method strikes a balance between masked autoencoders, which reconstruct pixels and train reliably but often learn weaker representations, and latent-space methods such as I-JEPA and Latent MIM, which can produce stronger features but are notoriously unstable. By replacing the target encoder with a frozen random linear projection in token space, Latent MIM Lite achieves stable training while preserving the benefits of latent-space prediction. We also discuss the broader challenges of evaluating geospatial foundation models, the trade-offs between embeddings and fine-tuning, and why practical performance on partner applications often matters more than leaderboard results.
* 📺 Video of this conversation on YouTube
* 🖥️ OlmoEarth on Github
* 🖥️ OlmoEarth Platform
* 👤 Joe’s website
Bio: Joseph Redmon is a research scientist at Ai2 building multimodal foundation models for geospatial data. As part of the OlmoEarth team he’s working to bring cutting edge AI research to non profits and NGOs working on conservation, ecological, and environmental problems.
This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit www.satellite-image-deep-learning.com - In this episode I sat down with Lakshay Sharma, a machine learning scientist at Instacart and former member of Microsoft’s geospatial AI team, to discuss self-supervised learning for remote sensing and his recent research on efficient pretraining for semantic segmentation. Lakshay explains the evolution of self-supervised learning, covering predictive, generative, and contrastive approaches, and discusses how foundation models such as DINO have transformed computer vision and geospatial machine learning. We explore the unique challenges of applying these techniques to remote sensing imagery, where assumptions that work for natural images often break down.We then dive into Lakshay’s recent paper, Sub-Image Overlap Prediction: Task-Aligned Self-Supervised Pretraining for Semantic Segmentation in Remote Sensing Imagery, presented at the Computer Vision for Earth Observation Workshop at WACV 2026. He walks through the intuition behind the method, which trains models to localize extracted sub-images within larger scenes as a proxy task for semantic segmentation. We discuss the experimental setup, comparisons against established self-supervised learning approaches, and the surprising finding that the method achieves competitive or superior results using only thousands of pretraining images rather than millions. Along the way, we explore transfer learning across datasets, the growing importance of data efficiency, and why targeted pretraining may offer a compelling alternative to increasingly resource-intensive foundation model development for niche geospatial applications.
* 📺 Video of this conversation on YouTube
* 👤 Lakshay on LinkedIn
* 🖥️ Personal website of Lakshay
* 📖 Paper
Bio: Lakshay Sharma is a Senior Machine Learning Scientist / Engineer at Instacart. His research spans Computer Vision (CV) and Vision-Language Models (VLMs) with a focus on Self-Supervised and Semi-Supervised Learning. He has previously worked at Microsoft on multi-modal representation learning, and using aerial/satellite and streetside imagery for maps and geospatial applications. He has also worked at Amazon where he was focused on representation learning for videos. Based in New York City, Lakshay is an avid fan of soccer, snowboarding, and cricket. He often daydreams of some day applying his computer vision chops to sports.
This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit www.satellite-image-deep-learning.com - In this episode I sat down with Jennifer Marcus and Isaac Corley from Taylor Geospatial to explore Fields of the World - an open initiative to create globally consistent agricultural field boundary datasets from satellite imagery using AI and cloud-native geospatial infrastructure. Taylor Geospatial, a newly formed research organization, is building openly licensed global datasets as foundational public goods. Jen and Isaac explain the motivation behind the project, the challenges of scaling machine learning beyond well-labelled regions, and why openness in datasets, tooling, and intermediate model outputs, is central to their approach.We dive into the technical details behind the first global release: assembling noisy and uneven benchmark datasets from around the world, training models that generalise across diverse agricultural systems, and releasing everything from Sentinel-2 mosaics and raw segmentation probabilities to polygonised field boundaries through Source Cooperative. Along the way, we discuss community-driven improvement loops inspired by OpenStreetMap, the limitations of 10 m imagery for smallholder agriculture, and the importance of pairing academic researchers with engineering teams to rapidly operationalise new methods. Finally, we look ahead to Taylor Geospatial’s next phase - richer agricultural datasets, “Features of the World,” and a benchmarking initiative aimed at improving evaluation standards and reproducibility across geospatial foundation models.
* 📺 Video of this conversation on YouTube
* 🖥️ Taylor Geospatial website
* 🖥️ FTW website
Bio: Jennifer Marcus is Vice President of Strategic Innovation Programs at Taylor Geospatial, where she advances partnerships and programs that translate breakthrough geospatial AI research into real-world impact. With deep experience across defence, federal government, and open-source geospatial ecosystems, Jennifer brings decades of expertise translating emerging technologies into mission-critical impact. She previously served as the inaugural Executive Director of Taylor Geospatial Engine, which in 2024, launched what would become Fields of The World, and has held leadership roles at Planet, Boundless Spatial, and Northrop Grumman.
Bio: Isaac Corley is Director of AI/ML Research at Taylor Geospatial, where he leads a team to build the models behind earth observation research and to create open data products that elevate the geospatial market and community as a whole. Isaac builds and publishes geospatial AI from research through production, including the RasterFlow platform at Wherobots, which was used to run Fields of The World. He has served as PI on the IARPA SMART program at BlackSky and maintains widely-used open-source projects, including TorchGeo and SMP. Check out his blog with Caleb Robinson at geospatialml.com.
This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit www.satellite-image-deep-learning.com BetaEarth: Open Embeddings of Sentinel-2 and Sentinel-1 with a Little Help of AlphaEarth
29/04/2026 | 26minIn this episode I sat down with Mikolaj (Miko) Czerkawski from Asterisk Labs to explore BetaEarth, an experimental open-source emulator trained on AlphaEarth Foundations' public embedding archive. AEF — released by Google and Google DeepMind as a global 10 m embedding product derived from a wide range of Earth-observation modalities — is what makes BetaEarth possible: its openness lets the community build lightweight independent emulators that approximate AEF's pixelwise outputs from standard Sentinel inputs, and use them to probe how much of a model's behaviour is captured in its public embeddings. Miko walks through BetaEarth's design — compact architectures based on SegFormer-B2 with separate per-modality encoders, and a shared DINOv3 backbone over 3-band spectral primitives — and the surprising finding that reasonably strong approximations can be achieved even from simple RGB inputs.We then dive into a live demo: generating BetaEarth embeddings for arbitrary regions and time ranges using Sentinel-1, Sentinel-2, and COP-DEM data. Along the way, we cover practical considerations such as cloud contamination, modality trade-offs, tiling artefacts, and strategies for merging multi-temporal signals. Finally, we discuss what this complementary tooling enables for the geospatial ML community — embeddings as pretraining or regularisation signals, lightweight local inference alongside AEF's global annual rasters, and what the combination of large proprietary archives and open emulator-style tools could unlock next.
* 📺 Video of this conversation & demo on YouTube
* 🖥️ BetaEarth Github page
* 🖥️ BetaEarth demo on Huggingface
Bio: Miko is a researcher specialising in AI, computer vision, signal processing and Earth observation. Before co-founding Asterisk Labs he was a postdoctoral research fellow at the European Space Agency. His research interests include data-centric analyses of large-scale Earth observation data, dataset curation, generative modelling, and restoration tasks for satellite imagery. He is a co-founder of the Major TOM community project, a platform for collaborating and reusing Earth observation datasets designed specifically for AI pipelines. He received the B.Eng. degree in electronic and electrical engineering in 2019 from the University of Strathclyde in Glasgow, United Kingdom, and the Ph.D. degree in 2023 at the same institution, specialising in applications of computer vision to Earth observation data.
This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit www.satellite-image-deep-learning.com- In this episode I sat down with Kentaro Wada, a computer vision engineer at Mujin and creator of LabelMe, to explore the evolution of image annotation workflows. We discuss how his need to label data for a robotics challenge led to building one of the most widely used open-source annotation tools, and how it has evolved alongside the shift from traditional computer vision to deep learning. Kentaro explains the impact of foundation models like Segment Anything (SAM), and how annotation is rapidly moving toward a prompt-and-verify paradigm where models do the heavy lifting and humans focus on quality control. We also dive into his recent work integrating SAM into LabelMe, the challenges of applying these models to satellite imagery, and why approaches like bounding-box prompting outperform text in that domain. Finally, we cover new support for large, multi-channel geospatial data, practical deployment considerations, and what this means for scaling annotation in real-world machine learning systems. Note that a recording of this conversation, along with a demonstration of geospatial annotation using LabelMe, is available on YouTube via the links below:
* 🖥️ LabelMe website
* 🖥️ Kentaro’s personal website
* 📺 Video of this conversation on YouTube
* 📺 Demo video on YouTube
Bio: Kentaro Wada was born in Japan in 1994. He received his B.Sc. (2016) and M.Sc. (2018) from Mechanical Engineering and Computer Science Department in The University of Tokyo (UTokyo). In his research at UTokyo, he was working on learning-based scene understanding for robotic manipulation at JSK Laboratory supervised by Prof. Masayuki Inaba and Prof. Kei Okada. He received his PhD in 2022, at Dyson Robotics Laboratory in Imperial College London supervised by Prof. Andrew Davison. During his PhD, he worked on object-level semantic scene understanding, a general scene representation useful for robotic manipulation, and showed several novel capabilities of robots. He joined Mujin, Inc. in 2022 as a computer vision engineer, and is working on advancing robots' capabilities in the real-world environment.
This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit www.satellite-image-deep-learning.com
Mais podcasts de Ciência
Podcasts em tendência em Ciência
Sobre Satellite image deep learning
Dive into the world of deep learning for satellite images with your host, Robin Cole. Robin meets with experts in the field to discuss their research, products, and careers in the space of satellite image deep learning. Stay up to date on the latest trends and advancements in the industry - whether you’re an expert in the field or just starting to learn about satellite image deep learning, this a podcast for you. Head to https://www.satellite-image-deep-learning.com/ to learn more about this fascinating domain www.satellite-image-deep-learning.com
Site de podcastOuça Satellite image deep learning, Hidden Brain e muitos outros podcasts de todo o mundo com o aplicativo o radio.net

Obtenha o aplicativo gratuito radio.net
- Guardar rádios e podcasts favoritos
- Transmissão via Wi-Fi ou Bluetooth
- Carplay & Android Audo compatìvel
- E ainda mais funções
Obtenha o aplicativo gratuito radio.net
- Guardar rádios e podcasts favoritos
- Transmissão via Wi-Fi ou Bluetooth
- Carplay & Android Audo compatìvel
- E ainda mais funções


Satellite image deep learning
Leia o código,
baixe o aplicativo,
ouça.
baixe o aplicativo,
ouça.

































