Efficient Python Workflows for Accessing, Analyzing, Visualizing, and Indexing Large Environmental Datasets in the Cloud

Join us for an immersive, hybrid, full-day Python workshop designed to equip attendees with applied skills in accessing, querying, processing, visualizing, and indexing large-scale environmental datasets stored in public cloud object storage systems. This workshop focuses on real-world climatological and meteorological use cases, featuring datasets from the Coupled Model Intercomparison Project Phase 6 (CMIP6) and NOAA’s Real-Time Mesoscale Analysis (RTMA), with access via Google Cloud Storage (GCS) and Amazon S3.

Participants will learn how to efficiently build scalable Python pipelines for exploring, filtering, and visualizing high-resolution climate datasets across time and space. Hands-on examples will demonstrate how to extract time series for point locations and areal shapes, visualize multi-dimensional model output, and apply optimization strategies for handling big cloud-hosted environmental data.

New in this 2026 edition: we introduce powerful cloud-native indexing techniques using Kerchunk, enabling virtual Zarr access to GRIB2, NetCDF, and HDF5 files without full conversion. Attendees will gain practical experience managing rate limiting, leveraging metadata-aware access patterns, and scaling workflows using tools such as Xarray, fsspec, Dask, and Kerchunk.

2026 AMS Madison Summit
Monona Terrace Community and Convention Center
2 August 2026, Madison, WI (Hybrid): 8am - 4pm Central Time

Registration

REGISTRATION RATES

  AMS Member Early Rate Non-Member Early Rate AMS Student Member Early Rate AMS Member Late Rate Non-Member Late Rate AMS Student Member Late Rate
  Through July 10 Through July 10 Through July 10 Through Aug 1 Through Aug 1 Through Aug 1
Python Workshop  $190.00 $220.00 $105.00 $230.00 $260.00 $145.00

 

Course Description:

By the end of this Python workshop, attendees will be able to:

  • Comfortably navigate and explore large CMIP6-like environmental datasets stored in public cloud repositories using Python
  • Identify and differentiate between spatiotemporal data structures, file formats, and relevant Python libraries
  • Build robust Python data pipelines (using functions and classes) that connect to cloud object storage systems (e.g., GCS, Amazon S3)
  • Analyze ECMWF historical climate simulations and NOAA-RTMA high-resolution operational meteorological data
  • Access, slice, filter, and query multi-dimensional climate data using Xarray
  • Generate time series from multidimensional arrays for both points and polygonal regions
  • Create virtual Zarr stores from GRIB2 and NetCDF files using Kerchunk for fast, cloud-native access • Index large environmental datasets stored in the cloud to optimize performance and reduce access time
  • Manage and process out-of-memory datasets using Dask and metadata-aware strategies
  • Produce stunning visualizations of big climate and meteorological data all within Python

VIEW AGENDA

If you have questions regarding the course, please contact Ryan Lafler.

Instructors:

Ryan Lafler
Ryan Paul Lafler

Founder, Principal Systems Architect, and Lead Consultant - Premier Analytics Consulting, LLC

Miguel Bravo
Miguel Angel Bravo

Data Scientist and Consultant - Premier Analytics Consulting, LLC