# Unable to run ‘ExpectColumnPairValuesToBeEqual’ with spark on Databricks

**URL:** https://discourse.greatexpectations.io/t/unable-to-run-expectcolumnpairvaluestobeequal-with-spark-on-databricks/2062
**Category:** GX Core Support
**Tags:** databricks, types-of-expectation
**Created:** [January 21, 2025, 9:11am UTC](https://discourse.greatexpectations.io/t/unable-to-run-expectcolumnpairvaluestobeequal-with-spark-on-databricks/2062 "2025-01-21T09:11:54Z")
**Posts on this page:** 10
**Page:** 1

<div class="post-metadata">

### Author: ![motaslimi](https://avatars.discourse-cdn.com/v4/letter/m/94ad74/32.png) [@motaslimi](https://discourse.greatexpectations.io/u/motaslimi)
#### Post date: [January 21, 2025, 9:11am UTC](https://discourse.greatexpectations.io/t/unable-to-run-expectcolumnpairvaluestobeequal-with-spark-on-databricks/2062/1 "2025-01-21T09:11:54Z")

</div>

When I’m trying to validate any expectation with parameters column\_A, column\_B, I get the following error:

[CANNOT\_RESOLVE\_DATAFRAME\_COLUMN] Cannot resolve dataframe column "col\_1". It’s probably because of illegal references like `df1.select(df2.col(\"a\"))`. SQLSTATE: 42704

My code:

df = spark.sql(f"SELECT \* FROM samples.sales\_schema.sales")  
context = gx.get\_context()  
data\_source\_name = sales  
data\_source = context.data\_sources.add\_spark(name=data\_source\_name)  
data\_asset\_name = sales\_data\_asset  
data\_asset = data\_source.add\_dataframe\_asset(name=data\_asset\_name)  
batch\_parameters = {“dataframe”: df}  
batch\_definition\_name = f"sales\_batch\_definition"  
batch\_definition = data\_asset.add\_batch\_definition\_whole\_dataframe(batch\_definition\_name)  
batch = batch\_definition.get\_batch(batch\_parameters=batch\_parameters)  
Exp = gxe.ExpectColumnPairValuesToBeEqual(column\_A = col\_1, column\_B = col\_2, mostly=0.5)  
batch.validate(Exp)

My environment:  
Databricks Runtime: 15.4 LTS (includes Apache Spark 3.5.0, Scala 2.12)  
great\_expectations[spark]: 1.3.2

The same issue happens also for: ExpectColumnPairValuesAToBeGreaterThanB, ExpectColumnPairValuesToBeInSet.

---

<div class="post-metadata">

### Author: ![bidek56](https://avatars.discourse-cdn.com/v4/letter/b/b77776/32.png) [@bidek56](https://discourse.greatexpectations.io/u/bidek56)
#### Post date: [January 21, 2025, 10:30pm UTC](https://discourse.greatexpectations.io/t/unable-to-run-expectcolumnpairvaluestobeequal-with-spark-on-databricks/2062/2 "2025-01-21T22:30:21Z")

</div>

What does your `df` look like? Does it have `col_1` and `col_2`?

---

<div class="post-metadata">

### Author: ![Han](https://avatars.discourse-cdn.com/v4/letter/h/977dab/32.png) [@Han](https://discourse.greatexpectations.io/u/Han)
#### Post date: [January 22, 2025, 1:59am UTC](https://discourse.greatexpectations.io/t/unable-to-run-expectcolumnpairvaluestobeequal-with-spark-on-databricks/2062/3 "2025-01-22T01:59:03Z")

</div>

> [@motaslimi](#):
>
> ExpectColumnPairValuesToBeEqual

This works for me

Possible to print out your dataframe? Also, check if there is missing quotation

 ![image](https://canada1.discourse-cdn.com/flex031/uploads/greatexpectations/original/1X/7d30d74fb5d72e38c5d56fa2a26be61088f2ca1c.png)

---

<div class="post-metadata">

### Author: ![motaslimi](https://avatars.discourse-cdn.com/v4/letter/m/94ad74/32.png) [@motaslimi](https://discourse.greatexpectations.io/u/motaslimi)
#### Post date: [January 22, 2025, 12:16pm UTC](https://discourse.greatexpectations.io/t/unable-to-run-expectcolumnpairvaluestobeequal-with-spark-on-databricks/2062/4 "2025-01-22T12:16:33Z")

</div>

yes it does, all other expectations with only one column work for me.

---

<div class="post-metadata">

### Author: ![motaslimi](https://avatars.discourse-cdn.com/v4/letter/m/94ad74/32.png) [@motaslimi](https://discourse.greatexpectations.io/u/motaslimi)
#### Post date: [January 22, 2025, 12:18pm UTC](https://discourse.greatexpectations.io/t/unable-to-run-expectcolumnpairvaluestobeequal-with-spark-on-databricks/2062/5 "2025-01-22T12:18:06Z")

</div>

I checked with quotations but still gives me the same error. It works for Expectations with one column input.

---

<div class="post-metadata">

### Author: ![bidek56](https://avatars.discourse-cdn.com/v4/letter/b/b77776/32.png) [@bidek56](https://discourse.greatexpectations.io/u/bidek56)
#### Post date: [January 22, 2025, 3:47pm UTC](https://discourse.greatexpectations.io/t/unable-to-run-expectcolumnpairvaluestobeequal-with-spark-on-databricks/2062/7 "2025-01-22T15:47:36Z")

</div>

This local pyspark example works fine for me.

```Python
import great_expectations as gx
from pyspark.sql.functions import col
from pyspark.sql import SparkSession, DataFrame

import os
os.environ["GX_ANALYTICS_ENABLED"] = "false"

def check_frame():

    spark = SparkSession.builder.master("local").getOrCreate()

    df: DataFrame = spark.createDataFrame([(1, 1.0), (1, 2.0), (2, 3.0), (2, 5.0), (2, 10.0)],("id", "col_1")) \
            .withColumn("col_2", col("col_1") )

    df.createOrReplaceTempView("tbl2")
    df2 = spark.sql(f"SELECT * FROM tbl2")
    print(df2.show())

    context = gx.get_context()
    data_source = context.data_sources.add_spark(name="data_source_name")
    data_asset = data_source.add_dataframe_asset(name="data_asset_name")
    batch_parameters = {"dataframe": df} 
    batch_definition_name = f"sales_batch_definition" 
    batch_definition = data_asset.add_batch_definition_whole_dataframe(batch_definition_name)
    batch = batch_definition.get_batch(batch_parameters=batch_parameters)
    Exp = gx.expectations.ExpectColumnPairValuesToBeEqual(column_A = "col_1", column_B = "col_2", mostly=0.5)
    batch.validate(Exp)

if __name__ == ' __main__':
    check_frame()

```

---

<div class="post-metadata">

### Author: ![OhadE](https://avatars.discourse-cdn.com/v4/letter/o/3be4f8/32.png) [@OhadE](https://discourse.greatexpectations.io/u/OhadE)
#### Post date: [March 25, 2025, 9:45pm UTC](https://discourse.greatexpectations.io/t/unable-to-run-expectcolumnpairvaluestobeequal-with-spark-on-databricks/2062/8 "2025-03-25T21:45:30Z")

</div>

Having the exact same issue.  
My environment:

- Databricks Runtime: 15.4 LTS (Apache Spark 3.5.0, Scala 2.12)
- `great_expectations[spark]`: 1.3.9

It always points to `column_A`. I tried switching the values of A and B, but it still reports the issue with `column_A`.

Also tried running the example shared by @bidek56, but got an error there as well.  
Really frustrating.

---

<div class="post-metadata">

### Author: ![OhadE](https://avatars.discourse-cdn.com/v4/letter/o/3be4f8/32.png) [@OhadE](https://discourse.greatexpectations.io/u/OhadE)
#### Post date: [March 26, 2025, 12:59pm UTC](https://discourse.greatexpectations.io/t/unable-to-run-expectcolumnpairvaluestobeequal-with-spark-on-databricks/2062/9 "2025-03-26T12:59:10Z")

</div>

I found the root cause and identified a workaround.

The issue is with the logic used to filter rows, controlled by the `ignore_row_if` parameter. It can take one of the following values:

- `both_values_are_missing` (default)
- `either_value_is_missing`
- `neither`

What happens is that with the default value, some rows are filtered out, which modifies the DataFrame. This occurs in `sparkdf_execution_engine.py`, inside the `get_domain_records` method.

The modified DataFrame then causes errors later, such as the one we saw:  
`df1.select(df2.col("a"))`  
This likely results from operations being applied to DataFrames of different sizes. I suspect the issue comes from this part of the code:

```python
data = df.withColumn("__unexpected", unexpected_condition)
filtered = data.filter(F.col("__unexpected") == True).drop(F.col("__unexpected"))

```

**Workaround:**  
Set `ignore_row_if` to `neither`. This prevents the DataFrame from being modified and allows the code and tests to run as expected.

Example:

```python
suite.add_expectation(gxe.ExpectColumnPairValuesToBeEqual(
    column_A="my_column_a",
    column_B="my_column_b",
    ignore_row_if="neither"
))

```

---

<div class="post-metadata">

### Author: ![motaslimi](https://avatars.discourse-cdn.com/v4/letter/m/94ad74/32.png) [@motaslimi](https://discourse.greatexpectations.io/u/motaslimi)
#### Post date: [March 26, 2025, 1:54pm UTC](https://discourse.greatexpectations.io/t/unable-to-run-expectcolumnpairvaluestobeequal-with-spark-on-databricks/2062/10 "2025-03-26T13:54:58Z")

</div>

Thanks a lot OhadE! Your workaround works perfectly!

---

<div class="post-metadata">

### Author: ![brett.koblinger](https://avatars.discourse-cdn.com/v4/letter/b/8edcca/32.png) [@brett.koblinger](https://discourse.greatexpectations.io/u/brett.koblinger)
#### Post date: [July 28, 2025, 12:12pm UTC](https://discourse.greatexpectations.io/t/unable-to-run-expectcolumnpairvaluestobeequal-with-spark-on-databricks/2062/11 "2025-07-28T12:12:38Z")

</div>

I think the issue also depends on the access mode of the Databricks cluster. I can successfully run ExpectColumnPairValuesAToBeGreaterThanB expectations when the access mode is “Dedicated (formerly: Single user)”. If the access mode is changed to “Standard (formerly: Shared)”, I get exceptions of the type " **`[CANNOT_RESOLVE_DATAFRAME_COLUMN] Cannot resolve dataframe column`**"

I am currently using Databricks runtime 17.0, but the same behavior existed with previous runtime’s too.
