# Slow validations

**URL:** <https://discourse.greatexpectations.io/t/slow-validations/1345>\
**Category:** GX Core Support\
**Created:** [August 24, 2023, 12:57pm UTC](https://discourse.greatexpectations.io/t/slow-validations/1345 "2023-08-24T12:57:16Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![erman](https://avatars.discourse-cdn.com/v4/letter/e/94ad74/32.png) [@erman](https://discourse.greatexpectations.io/u/erman)\
**Post date:** [August 24, 2023, 12:57pm UTC](https://discourse.greatexpectations.io/t/slow-validations/1345/1 "2023-08-24T12:57:16Z")

</div>

Hello everyone,

I am running airflow 2.6.1 in docker. Within a DAG I have imported the GX python library and am running around 40 validations on Dataframes.

```auto
config_data_docs_sites = {
        "s3_site": {
            "class_name": "SiteBuilder",
            "store_backend": {
                "class_name": "TupleS3StoreBackend",
                "bucket": "great-expectations",
                "prefix": "data_docs",
                "boto3_options": BOTO3_OPTIONS
            },
        },
    }
    data_context_config = DataContextConfig(
        store_backend_defaults=S3StoreBackendDefaults(default_bucket_name=GX_BUCKET_NAME),
        data_docs_sites=config_data_docs_sites
    )
    context = BaseDataContext(project_config=data_context_config)
    asset_names, data_source = gx_preparation(s3_client, context, latest_version, latest_file_name, normalised_file_list)
    failed = False

    for asset_name in asset_names:
        data_asset = data_source.get_asset(asset_name)
        my_batch_request = data_asset.build_batch_request()

        if asset_name.endswith('df1'):
            expectation = 'Exp_Abteilung'
            column_names = ["Hauptabteilung", "Nebenabteilung"]
        elif asset_name.endswith('df2'):
            expectation = 'Exp_Person'
            column_names = ["Person", "PID"]
        elif asset_name.endswith('df3'):
            expectation = 'Exp_Ausruestung'
            column_names = ["Ausrüstung", "KID"]
        else:
            expectation = 'Exp_Combined'
            column_names = []

        checkpoint = gx.checkpoint.SimpleCheckpoint(
            name=f"{asset_name.replace('/', '_')}_test",
            data_context=context,
            validations=[
                {
                    "batch_request": my_batch_request,
                    "expectation_suite_name": expectation,
                },
            ],
            run_name_template=latest_file_name + "_" + latest_version,
            runtime_configuration={
                    "result_format": {
                        "result_format": "COMPLETE",
                        "unexpected_index_column_names": column_names,
                        "return_unexpected_index_query": True,
                        "include_unexpected_rows": True
                    },
                },
        )

        result = checkpoint.run()
        if not result["success"]:
            failed = True

```

All the stores are inside a MinIO bucket. The issue is the speed of these validations. They take about 4 min to complete. Is there a better way to do this and improve the speed?

---

<div class="post-metadata">

**Author:** ![austin\_gx](https://yyz1.discourse-cdn.com/flex031/user_avatar/discourse.greatexpectations.io/austin_gx/32/193_2.png) [@austin\_gx](https://discourse.greatexpectations.io/u/austin_gx)\
**Post date:** [August 31, 2023, 4:46pm UTC](https://discourse.greatexpectations.io/t/slow-validations/1345/2 "2023-08-31T16:46:40Z")

</div>

Hey @erman ! Thanks for reaching out. If you build a list of your `batch_request` / `expectation_suite_name` pairs as you iterate, you can pull the actual validation step out of your loop & pass all of your validations in at once.

```auto
validations = []

for asset_name in asset_names:
        data_asset = data_source.get_asset(asset_name)
        my_batch_request = data_asset.build_batch_request()

        if asset_name.endswith('df1'):
            expectation = 'Exp_Abteilung'
            column_names = ["Hauptabteilung", "Nebenabteilung"]
        elif asset_name.endswith('df2'):
            expectation = 'Exp_Person'
            column_names = ["Person", "PID"]
        elif asset_name.endswith('df3'):
            expectation = 'Exp_Ausruestung'
            column_names = ["Ausrüstung", "KID"]
        else:
            expectation = 'Exp_Combined'
            column_names = []

        validations.append(
                {"batch_request": my_batch_request, "expectation_suite_name": expectation}
        )

...

checkpoint = gx.checkpoint.SimpleCheckpoint(
            name="checkpoint_name",
            data_context=context,
            validations=validations,
            runtime_configuration={
                    "result_format": {
                        "result_format": "COMPLETE",
                        "unexpected_index_column_names": column_names,
                        "return_unexpected_index_query": True,
                        "include_unexpected_rows": True
                    },
                },
        )

```

This will prevent you from being able to individually name checkpoints & run\_names in the same way, but will also prevent you from having to create a checkpoint for every iteration & should improve overall speed of the validation process.
