Skip to main content
Public Preview

RasterFlow Overview

Learn about RasterFlow’s key features and capabilities

SAM3 Notebook

The full text-prompted geometry inference walkthrough

RasterFlow Datasets

Learn about built-in datasets and how to bring your own

RasterFlow Billing

Learn how RasterFlow usage is metered in RasterFlow Spatial Units
SAM3 detects objects directly from text prompts, so there is no fixed set of detectable categories. The same function call that finds roofs can help find solar panels, shipping containers, parking lots, roads, or sidewalks. Goal of this guide: Roof detection across a property portfolio is the working example this guide follows. Nothing in the workflow is specific to roofs: change the prompt and the same steps apply to whatever you need to find.
What you will be able to do at the end of this guide:
  1. Estimate each task before you start it, from the area and resolution you chose. The preview mosaic and overviews in this walkthrough are billed separately from the recipe’s mosaic and inference.
  2. Confirm 30 cm NAIP covers your area in the window you asked for, and get back the years that would work when it does not, before anything is billed.
  3. Establish how old the imagery is from the scene index, which is the only place the flight date is available.
  4. Build a preview mosaic and inspect it for gaps and seams before the model runs.
  5. Run one inference call with several text prompts and read the polygons and confidence scores it returns.
  6. Set an area floor, persist the detections, and put them on a map, then score roof age against the staleness of the evidence.
Reading this guide: The numbered steps under Inspect your own area of interest and Augment old data with newer available datasets are the parts you run yourself. Every other section is reference: a finished run to look at, what it costs, and where the model’s limits are.

Spend your data budget on the properties that need it

Managing a portfolio of 40,000 properties usually means you only have the budget to closely inspect a few hundred. Licensing national roof-condition and hazard datasets charges you for nationwide coverage when you only need three counties, and renews every year. RasterFlow runs on demand instead: describe what you want to find in plain language, define your target area, and receive georeferenced polygons with confidence scores. Estimate the charges for each task from your area of interest before you run it. RasterFlow provides a fast triage layer, low-cost detection you run ahead of the better and more expensive options, helping you pinpoint the exact addresses that justify an expensive full inspection or condition report, instead of having to buy the whole thing.
What SAM3 returns: A polygon and a confidence score for each detection. There is no condition grade and no material class. Roof and solar detections are separate polygons; review or spatially match them before deciding whether a particular roof carries solar. The detections help target a closer inspection. Grading condition still takes a condition dataset, an aerial-inspection vendor, or an adjuster.

What an inspection run costs

RasterFlow Spatial Units come from pixels, so you can estimate each task before you start it. The recipe in step 7 bills a mosaic and SAM3 inference. This walkthrough also builds a separate preview mosaic and runs Build Multiscales in step 6. The estimates below assume 30 cm NAIP, four-band mosaics, three RGB bands for SAM3 inference, and one time period. Each task carries a 1 RasterFlow Spatial Unit minimum. These are area-based estimates; the recipe adjusts its mosaic grid and may buffer a small area of interest, so actual usage can differ. Each task has its own price per Spatial Unit, so multiply the units in each column by that task’s current price and add the costs, rather than treating the units as one combined rate. A full cache hit for an identical task within seven days is not charged, but the separate preview mosaic and the recipe mosaic use different inputs and should be budgeted as separate tasks. A WherobotsDB runtime is metered separately.
Validate the prompt and the confidence threshold on the smallest area that contains a representative sample, then scale up. Resolution and area are the only large levers, and a rerun of an identical task within the 7-day cache period is not charged. For the current rate per RasterFlow Spatial Unit, see Wherobots Pricing.

What one prompt returns

This run has already finished, so there is nothing to run or set up here. The screenshot below comes from a SAM3 run over the whole of Marion County, Oregon: 3,085 km² in one call. The red polygons are the "roofs" prompt. Two things had to be computed before this map existed, both of them by that finished run:
  • The input mosaic. RasterFlow pulled the 30 cm NAIP scenes covering the county over the date window and stitched them into one seamless Zarr mosaic. That is the imagery under the polygons, and it is what SAM3 reads.
  • The detections. SAM3 ran over that mosaic and returned a polygon and a confidence score for every object it found. For a county-sized result those polygons are tiled into a single PMTiles layer so the whole county renders in a browser.
Both layers are on the map, listed in the Layers panel as SAM3 input mosaic and SAM3 PM Tiles, so you can switch between the detections and the raw imagery SAM3 saw. When you inspect your own area further down, predict_mosaic_geometries_recipe() does both steps in one call.
The Wherobots Cloud map viewer showing several hundred red roof polygons over 30 cm aerial imagery of a residential neighborhood, with a Layers panel on the left listing eight SAM3 sublayers, one per text prompt

A finished SAM3 run over Marion County, Oregon, viewed in Wherobots Cloud. Every red polygon is a roof SAM3 detected from the prompt 'roofs'. The sublayers on the left are the text prompts submitted in the same call.

Open the Marion County run on the map

No account or runtime needed. Zoom into a neighborhood and toggle the sublayers.

Several prompts in one run

That run carried eight text prompts. The output is a single PMTiles layer with one sublayer per prompt, so you toggle whichever question you are asking.
The Layers panel from Wherobots Cloud, showing a single layer named SAM3 PM Tiles with eight sublayers listed beneath it, one for each text prompt in the run. The solar panels and roofs sublayers are visible; the rest are toggled off.

The sublayers of the Marion County output. Each one is a text prompt passed to the same predict_mosaic_geometries_recipe() call.

The published RasterFlow Spatial Unit estimate uses spatial pixels × bands × time periods; prompt count is not a term in that estimate. See RasterFlow Billing for the calculation, and check Workload History for the units actually billed on your run. For a property portfolio, one pass answers several underwriting questions:
The "solar panels" sublayer is switched on in the screenshots and returns few detections. Compare a sample with known panel locations before interpreting that count; the model can miss panels, especially small or obscured ones.

Before you start

  • RasterFlow access in your Organization. RasterFlow is in Public Preview for all paid Organizations.
  • A Wherobots notebook, a VS Code Extension workspace, or a Job Run with rasterflow_remote available.
  • The Micro runtime is enough. RasterFlow manages its own compute, so a larger runtime adds cost without making the run faster.
  • An area of interest in the continental United States, in any format GeoPandas can read (GeoJSON, GeoParquet, or Shapefile) or an in-memory GeoDataFrame.
  • 30 cm NAIP coverage for that area and date range. Both SAM3 recipes are fixed to 30 cm NAIP, and not every state has 30 cm in every year. Some states have none. Confirm 30 cm NAIP coverage for your area before you commit to one.

Inspect your own area of interest

Each step below is work you do in a notebook, against your own area of interest. If you have not opened the finished Marion County run yet, do that first: it needs no account or runtime, and it gives you a reference output to compare yours against. The code below is the notebook’s own cells. Everything you would change to point the run somewhere else is in the first cell, and the rest of the notebook reads from it.

Start a Micro runtime and open the notebook

Open the SAM3 solution notebook from the Model Hub in Wherobots Cloud, or open examples/Analyzing_Data/RasterFlow_SAM3.ipynb in a Wherobots notebook.
Starting a runtime begins a billable event in Spatial Units. RasterFlow tasks are billed separately, in RasterFlow Spatial Units.

Set what you edit in one cell

Every parameter the run takes is hoisted into a single cell, so pointing the notebook at a different area or a different set of objects is one edit in one place.
The cost estimate uses area, resolution, bands, and time periods, so these four prompts do not change that estimate. Keep "roofs" and "solar panels" for the inspection; "parking lots" and "roads" show what else a multi-prompt pass returns and can be removed.

Point the run at your own area of interest

wkls resolves a city or county name to a boundary, so there is no boundary file to find. The recipe reads the area of interest from storage, so it is written out to your own S3 path first.The screenshots above are from Keizer, Oregon, a city of 18.8 km² inside the Marion County run. Starting there lets you compare your own output against the reference map.
For your own area, wkls.us.oregon.cities() and wkls.us.oregon.counties() list what is available at each level. A wrong name raises an error that suggests the closest matches. County accessors carry a county suffix, so Marion County is wkls.us.oregon.marioncounty.
Other areas to try. Each of these resolves through wkls. The two estimates below are for the recipe’s mosaic and SAM3 inference only; the separate preview tasks add charges. The last column is the acquisition year that carries complete 30 cm NAIP coverage for that area, so set START and END to it.

Confirm 30 cm NAIP coverage for your area

NAIP collects state-level imagery at mixed resolutions, so SAM3’s required 30 cm data may not exist for your area or timeframe. Start with everything NAIP has flown over your area, at any resolution. The scene index RasterFlow queries is a public object, read here straight from S3.
For Keizer this returns three scenes in each of 2011, 2012 and 2014 at 1.0 m, three in 2020 at 0.6 m, and three in 2022 at 0.3 m. Only the 2022 scenes are usable, which is why the window in step 2 is 2022.Two details decide whether a year is usable:
  • The area has to be fully contained, not merely overlapped. A set of scenes covering most of your area still fails.
  • The filter is on time, not year. A flight on 2022-07-14 falls outside a window of 2023-01-01 to 2024-01-01, so an area flown in mid-2022 needs a 2022 window even though the calendar years sit next to each other.
The cell below applies the same test the recipe applies, and raises before anything is billed.
The USDA’s NAIP coverage map shows the same picture geographically.
Some areas have no 30 cm NAIP in any year. Where a state was flown at 60 cm throughout, no date window will satisfy the recipe. California is the clearest case: roughly 100 of the ~68,000 NAIP scenes over the state are 30 cm, so most Californian areas of interest cannot run these recipes at all. Bring your own rasters through a STAC catalog for those areas.

Read how old the imagery is

covering now holds the exact scenes the mosaic will be built from, and their time values are the flight dates. This is the age of the evidence behind every detection you are about to produce, so read it here, before the run.
Keizer’s three 2022 scenes were all flown on 2022-07-14. Keep that date: it is the vintage of your evidence, and the scoring below needs it.
This is the only place the flight date is available. The time column on the detections is the mosaic’s own time coordinate, not the date the imagery was flown, so it cannot be used for this. See Data freshness.

Build a preview mosaic and look at it

The recipe in the next step builds its own mosaic internally and returns the detections. Building a separate preview mosaic here, over the same area and window, gives you imagery to inspect before inference. The recipe can adjust its mosaic grid for inference, so this preview is a coverage and visual check, not an exact copy of the pixels SAM3 will read. The preview mosaic and Build Multiscales are separate billable tasks.
build_mosaics() stitches the 30 cm NAIP scenes verified above into one Zarr store. resolution is left at its default, the native resolution of the dataset: 30 cm for NAIP_30CM, which is what SAM3 needs.
build_zarr_multiscales() reads that store and writes a second one carrying overview levels and per-band histograms. A mosaic is written at a single native resolution, so without overviews a map has to stream full-resolution pixels for every pan and zoom.
Look at the mosaic before the model run:
  • Gaps. A hole means this preview mosaic has no imagery there. Investigate it before inference; the recipe builds a separate mosaic and may have different edges.
  • Seams. A tone shift across a straight line is where two flight dates meet.
  • The season. NAIP is flown leaf-on, so canopy sitting over a roof here is canopy the model sees too.
  • Coarse zoom levels. Blank or washed out means nodata pixels are being averaged into the overviews. Rebuild with an explicit fill value, rf_client.build_zarr_multiscales(source_store=mosaic_store, nodata=0) for 8-bit NAIP.
The widget loads the optimized preview store, not the source mosaic: the overview levels are what let it redraw as you zoom.

Run the inference recipe

One call ingests the imagery, builds the mosaic, and runs text-prompted geometry inference with SAM3. Every argument comes from the cell in step 2.
Budget roughly 22 minutes on a first run, most of it fixed setup rather than a function of your area. Nothing needs your attention while it runs, and Workload History shows progress and the RasterFlow Spatial Units each task consumed. A rerun of an identical task within the 7-day cache period is not charged.

Read the detections

Each row is one detected object:
  • geometry: the georeferenced polygon, in EPSG:4326 lon/lat
  • label: the text prompt that matched, exactly as you wrote it
  • bbox_score: confidence score for the detection
  • time: the time coordinate of the mosaic the pixel came from, not the flight date. Read the vintage off the scene dates printed in step 5 instead.
  • source_store: the mosaic the detection came from
  • local_x_offset, local_y_offset, global_x_offset, global_y_offset: the patch the detection was found in
The last two lines print the two dates side by side, so the gap between the mosaic’s time coordinate and the actual flight date is visible in your own run.
The file also carries a bbox struct (xmin, ymin, xmax, ymax): the GeoParquet covering for each polygon. GeoPandas consumes it as spatial metadata, so it is absent from detections_gdf.columns while sedona.read exposes it as a column.

Set area floors and save to the catalog

Persist the detections to the catalog so the SQL below has something to join to.ST_AreaSpheroid measures on the WGS 84 spheroid and reads coordinates as lon/lat. The recipe returns EPSG:4326, so it can be applied directly. If you ever hand it projected geometries, every area comes back near zero and the filter silently empties your table, so the cell below derives the CRS from the output rather than assuming it.
SAM3 emits a long tail of slivers and partial detections. Look at what the prompt-specific floors would drop before applying them.
The write keeps every prompt’s detections in one table, with the matched prompt in a layer column, so a multi-prompt run stays queryable and mappable as one layer set.
Without an area filter the roof layer carries slivers and partial detections that inflate counts. The 30 m² starting floor applies to roofs and other prompts; solar panels use a 0 m² floor so small arrays remain available for review. Choose each floor from the summary above and inspect the retained polygons on the map. A roof-sized floor will also discard many small "roads" segments.

Map your own detections

The row count is not the check. Put the polygons over the imagery for your own area and zoom in.
For a large detection set, generate PMTiles instead. The layer column becomes one sublayer per prompt, so a multi-prompt run stays togglable on the map, the way the Marion County run is.
Two stores go on the map together: the preview mosaic from step 6, and the PMTiles of the filtered detections. In Map, paste a URI into Layer URL, select Add Layer, and repeat for the second one.

How to approach output triage

The map from the last step is where triage happens. The rest of this section is what to expect when you zoom in, not more to run.
Close-up of 30 cm aerial imagery over a residential street, with red polygons outlining each house roof. The polygons follow complex L-shaped and T-shaped rooflines closely. Several small outbuildings and a detached garage carry no polygon.

Roof detections at full zoom in Keizer, Oregon. An attached garage comes through inside the house polygon. At a 0.3 confidence threshold, detached garages and sheds get no polygon at all.

  • Primary structures. SAM3 traces the roofline closely, L-shaped and T-shaped plans included. An attached garage lands inside the same polygon as the house.
  • Detached garages, sheds, and small outbuildings. Most get no polygon at confidence_threshold=0.3. To pick them up, lower the threshold and take on the extra false positives, or prompt for them in a separate pass.
  • Tree canopy. NAIP is flown during the agricultural growing season, so every scene is leaf-on. A roof under mature canopy comes back partial or missing, and the built-in datasets carry no leaf-off alternative.
  • A detection is not a building record. RasterFlow does not deduplicate these polygons against a parcel or a policy. One structure can produce two polygons, and the model can return a row of attached houses as one. Match to a footprint layer, then review unmatched and multi-footprint detections before counting buildings: see the Advanced Roof Inspection Guide.

Data freshness

Age drives an important component of a roof’s risk profile. As such, the age of the imagery matters considerably.
The most recent year for your state may not be 30 cm. Because 2025 delivered only about half the states at 30 cm, the newest imagery over your area of interest may be 60 cm, which SAM3’s recipes cannot use. Satisfying the 30 cm requirement can mean going back to an older acquisition year. Check the NAIP coverage map for both resolution and year before you pick a date window.
Read the vintage off the scene index, not off the detections. The start and end arguments are a window RasterFlow searches, not the date it found, and the time column on the detections is the mosaic’s own time coordinate rather than the date the imagery was flown. Over Keizer that column reads 2022-01-01 while the three scenes behind the mosaic were flown on 2022-07-14. The flight dates in step 5 are the age of your evidence, and the input to the scoring below.

Caveats and limitations

Both SAM3 recipes are configured for 30 cm NAIP, which covers the continental United States. For areas outside that footprint, or for imagery newer than the last NAIP flight, bring your own rasters through a STAC catalog and a GDAL Raster Tile Index.
The output is geometry plus a confidence score. There is no condition grade, material class, or damage flag. You can prompt for visually distinct roof types such as "metal roofs" or "tarped roofs", but SAM3’s accuracy on those categories is not benchmarked. Validate any such prompt against a sample you have ground truth for before it feeds a pricing decision.
confidence_threshold is required and has no safe default. At 0.3, primary structures detect reliably and small accessory structures are missed. Lower it to catch more and expect more false positives: driveway aprons, patios, and pale ground surfaces read as roofs. Tune it on a small area first, and check the result on a map instead of a row count.
"roofs", "roof", and "rooftops" are three different prompts and can return different detections. Fix the wording once you have validated it, and record it alongside the results.
The SAM3 recipes always resize the configured patch_size to 1008x1008. You can use it to control how much spatial context the model sees per pass, but a patch larger than 1008x1008 upsamples the imagery instead of adding detail.
RasterFlow is in Public Preview. Interfaces and recipe behavior can change. Keep a copy of any result you make decisions from; a rerun may not reproduce it exactly.

Augment old data with newer available datasets

If your imagery is from 2023 and you are underwriting in 2026, the roof in the picture is three years older than it looks. Add those years back before you score it. Divide the roof’s age by the expected service life of the covering. That service life depends on the material, roughly 20 years for three-tab asphalt shingle and 25 to 30 for architectural laminate. The number in the query below is a stand-in for whatever your own prediction curve says.

Register your own policy records

The scoring query joins the detections to your own policy records, so register those as my_portfolio first. It needs one row per insured property, with policy_id, a geometry to match against the detections, and roof_installed_year:

Score the roofs against imagery age

The imagery date is bound in from step 5 rather than read off a column, because the detections do not carry the flight date. The table holds every prompt from the run, so the join filters to the roofs layer.
roof_installed_year has to come from your own records: policy data, permit data, or a prior inspection. No open dataset carries it, and RasterFlow does not infer it. Without it, imagery_age_years on its own still works as a staleness flag, telling you which parts of your inspection rest on the oldest evidence.

Next steps

Advanced Roof Inspection Guide

Match detections to building footprints and measure nearby canopy height.

SAM3 notebook

The full walkthrough, including PMTiles generation for large detection sets.

Insurance & Risk solutions

Catastrophe exposure, underwriting enrichment, and portfolio concentration patterns.

Visualize RasterFlow outputs

Check a mosaic before inference and detections after, on the same map.

Run as a Job

Put the inspection on a schedule once the prompt and threshold are settled.

Bring your own rasters

Use newer or non-US imagery through a STAC catalog when NAIP is too old.

RasterFlow Billing

The full RasterFlow Spatial Unit calculation and the cost controls that matter.