Skip to main content
Spatial connected components identify groups of geometries that are transitively connected through a spatial predicate. Unlike density-based clustering (DBSCAN), they do not require a minimum density: any chain of pairwise-related geometries forms a single component. The spatialConnectedComponents method adds a component ID to each row in a DataFrame.

Parameters

DataFrame
DataFrame containing at least one geometry column. All rows must be unique.
String
Spatial predicate to use for connectivity. One of: "intersects", "touches", "overlaps", "contains", "covers", "within", "coveredby", "crosses", or "dwithin".
String
Name of the geometry column.If not provided, the algorithm automatically selects the geometry column:
  • If only one GeometryType column exists, it is used.
  • If multiple GeometryType columns exist and one is named geometry, it is used.
  • If multiple GeometryType columns exist and none are named geometry, the column name must be explicitly provided.
String
Name of the output column for the component ID. Default is "component_id".
Double
Distance threshold for the "dwithin" predicate. Ignored for other predicates. Default is 0.0.
Boolean
Whether to use spheroidal distance for the "dwithin" predicate. Default is false.
String
Optional column name for partitioning. When non-null, only rows with the same value in this column are considered for edge construction. Default is null.

Returns

The input DataFrame with a component ID column (BIGINT) added. Every row receives a component ID, including singletons.

Usage Examples

Column Function API

The ST_*CC functions are also available as Spark column functions via st_functions: