# GE with Databricks Delta

**URL:** <https://discourse.greatexpectations.io/t/ge-with-databricks-delta/82>\
**Category:** Archive\
**Created:** [March 17, 2020, 4:45am UTC](https://discourse.greatexpectations.io/t/ge-with-databricks-delta/82 "2020-03-17T04:45:14Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![routdeepak](https://avatars.discourse-cdn.com/v4/letter/r/8491ac/32.png) [@routdeepak](https://discourse.greatexpectations.io/u/routdeepak)\
**Post date:** [March 17, 2020, 4:45am UTC](https://discourse.greatexpectations.io/t/ge-with-databricks-delta/82/1 "2020-03-17T04:45:14Z")

</div>

I want to understand how GE integrates with data bricks delta tables? If you have any sample notebook it will be helpful to understand.

---

<div class="post-metadata">

**Author:** ![eugene.mandel](https://yyz1.discourse-cdn.com/flex031/user_avatar/discourse.greatexpectations.io/eugene.mandel/32/22_2.png) [@eugene.mandel](https://discourse.greatexpectations.io/u/eugene.mandel)\
**Post date:** [March 17, 2020, 3:34pm UTC](https://discourse.greatexpectations.io/t/ge-with-databricks-delta/82/2 "2020-03-17T15:34:20Z")

</div>

Rich Louden wrote up a very detailed blog post on using GE with Databricks: [https://www.unsupervised-learnings.co.uk/post/setting-your-data-expectations-data-profiling-and-testing-with-the-great-expectations-library/](https://www.unsupervised-learnings.co.uk/post/setting-your-data-expectations-data-profiling-and-testing-with-the-great-expectations-library/)

If you want GE to read data from delta, go into your project’s great\_expectations.yml config file and add an [Batch Kwargs Generator](https://docs.greatexpectations.io/en/latest/features/batch_kwargs_generator.html) of this class to your datasource: [S3GlobReaderBatchKwargsGenerator](https://docs.greatexpectations.io/en/latest/module_docs/generator_module.html#s3globreaderbatchkwargsgenerator).

The reference for this class has a sample yml config that you can copy and modify.

Replace `reader_method: parquet` in the generator’s config with `reader_method: delta`

**NOTE:**

Some API’s have changed since that blog post was published. If you are using GE 0.9.0 (or higher),  
please replace the build\_expectations method from the post with this:

```
def build_expectations(database, asset_name, context):
  
  exp_suite_name = database + "." + asset_name + "." + str(datetime.today().strftime("%d-%m-%Y")) + "_expectations"
  
  data = spark.table(database + "." + asset_name)
  
  spark_data = SparkDFDataset(data)
  
  profile = spark_data.profile(BasicDatasetProfiler)
  
  context.save_expectation_suite(profile[0], exp_suite_name)
  
  sqlContext.uncacheTable(database + "." + asset_name)

```

The profile method returns a tuple (expectation suite, validation results), so you need to pass the first member of that tuple to save\_expectation\_suite.

Also, you don’t have to create an empty expectation suite before profiling.

---

<div class="post-metadata">

**Author:** ![ychebaro](https://avatars.discourse-cdn.com/v4/letter/y/43a26b/32.png) [@ychebaro](https://discourse.greatexpectations.io/u/ychebaro)\
**Post date:** [May 4, 2020, 10:51pm UTC](https://discourse.greatexpectations.io/t/ge-with-databricks-delta/82/3 "2020-05-04T22:51:53Z")

</div>

Great post!  
Is it correct to assume that

`from great_expectations.datasource.generator.databricks_generator import DatabricksTableBatchKwargsGenerator`

becomes

`from great_expectations.datasource.batch_kwargs_generator import DatabricksTableBatchKwargsGenerator`

in the latest release of GE?

Thanks a lot!

---

<div class="post-metadata">

**Author:** ![aylr](https://avatars.discourse-cdn.com/v4/letter/a/8797f3/32.png) [@aylr](https://discourse.greatexpectations.io/u/aylr)\
**Post date:** [May 4, 2020, 10:54pm UTC](https://discourse.greatexpectations.io/t/ge-with-databricks-delta/82/4 "2020-05-04T22:54:42Z")

</div>

Almost! Here’s the full path to import.

from great\_expectations.datasource.batch\_kwargs\_generator import DatabricksTableBatchKwargsGenerator

> [@ychebaro](#):
>
> DatabricksTableBatchKwargsGenerator
