spatialConnectedComponents method adds a component ID to each row in a DataFrame.
Parameters
DataFrame
DataFrame containing at least one geometry column. All rows must be unique.
String
Spatial predicate to use for connectivity. One of:
"intersects", "touches", "overlaps", "contains", "covers", "within", "coveredby", "crosses", or "dwithin".String
Name of the geometry column.If not provided, the algorithm automatically selects the geometry column:
- If only one
GeometryTypecolumn exists, it is used. - If multiple
GeometryTypecolumns exist and one is namedgeometry, it is used. - If multiple
GeometryTypecolumns exist and none are namedgeometry, the column name must be explicitly provided.
String
Name of the output column for the component ID. Default is
"component_id".Double
Distance threshold for the
"dwithin" predicate. Ignored for other predicates. Default is 0.0.Boolean
Whether to use spheroidal distance for the
"dwithin" predicate. Default is false.String
Optional column name for partitioning. When non-null, only rows with the same value in this column are considered for edge construction. Default is
null.Returns
The input DataFrame with a component ID column (BIGINT) added. Every row receives a component ID, including singletons.
Usage Examples
Column Function API
TheST_*CC functions are also available as Spark column functions via st_functions:

