> ## Documentation Index
> Fetch the complete documentation index at: https://docs.wherobots.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Spatial Connected Components Scala Object

> Identify connected groups of geometries in Scala using spatial predicates, including intersection, touching, overlap, containment, and distance.

Spatial connected components identify groups of geometries that are transitively connected through a spatial predicate. Unlike density-based clustering (DBSCAN), they do not require a minimum density: any chain of pairwise-related geometries forms a single component.

The `spatialConnectedComponents` method adds a component ID to each row in a DataFrame.

```scala theme={"system"}
def spatialConnectedComponents(
      dataframe: DataFrame,
      predicate: String,
      geometry: String = null,
      componentColumnName: String = "component_id",
      distance: Double = 0.0,
      useSpheroid: Boolean = false,
      partitionColumn: String = null): DataFrame
```

## Parameters

<ParamField path="dataframe" type="DataFrame">
  DataFrame containing at least one geometry column. All rows must be unique.
</ParamField>

<ParamField path="predicate" type="String">
  Spatial predicate to use for connectivity. One of: `"intersects"`, `"touches"`, `"overlaps"`, `"contains"`, `"covers"`, `"within"`, `"coveredby"`, `"crosses"`, or `"dwithin"`.
</ParamField>

<ParamField path="geometry" type="String">
  Name of the geometry column.

  If not provided, the algorithm automatically selects the geometry column:

  * If only one `GeometryType` column exists, it is used.
  * If multiple `GeometryType` columns exist and one is named `geometry`, it is used.
  * If multiple `GeometryType` columns exist and none are named `geometry`, the column name must be explicitly provided.
</ParamField>

<ParamField path="componentColumnName" type="String">
  Name of the output column for the component ID. Default is `"component_id"`.
</ParamField>

<ParamField path="distance" type="Double">
  Distance threshold for the `"dwithin"` predicate. Ignored for other predicates. Default is `0.0`.
</ParamField>

<ParamField path="useSpheroid" type="Boolean">
  Whether to use spheroidal distance for the `"dwithin"` predicate. Default is `false`.
</ParamField>

<ParamField path="partitionColumn" type="String">
  Optional column name for partitioning. When non-null, only rows with the same value in this column are considered for edge construction. Default is `null`.
</ParamField>

## Returns

The input DataFrame with a component ID column (`BIGINT`) added. Every row receives a component ID, including singletons.

## Usage Examples

```scala theme={"system"}
import org.apache.sedona.stats.clustering.SpatialConnectedComponents.spatialConnectedComponents

// Intersection-based connected components
val result = spatialConnectedComponents(dataframe, "intersects")

// Distance-based connected components
val result = spatialConnectedComponents(dataframe, "dwithin", distance = 1.0)

// With partition column for performance
val result = spatialConnectedComponents(dataframe, "intersects", partitionColumn = "state")
```

### Column Function API

The `ST_*CC` functions are also available as Spark column functions via `st_functions`:

```scala theme={"system"}
import org.apache.spark.sql.sedona_sql.expressions.st_functions._

// Intersection-based predicate
val result = df.select(col("*"), ST_IntersectsCC(col("geometry")).as("component_id"))

// With partition_by
val result = df.select(
  col("*"),
  ST_IntersectsCC(col("geometry"), col("category")).as("component_id")
)

// Distance-based
val result = df.select(
  col("*"),
  ST_DWithinCC(col("geometry"), lit(1.0), lit(false)).as("component_id")
)
```


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.