# How to configure a PySpark datasource for accessing the data from AWS S3?

**URL:** <https://discourse.greatexpectations.io/t/how-to-configure-a-pyspark-datasource-for-accessing-the-data-from-aws-s3/78>\
**Category:** Archive\
**Created:** [March 12, 2020, 3:30pm UTC](https://discourse.greatexpectations.io/t/how-to-configure-a-pyspark-datasource-for-accessing-the-data-from-aws-s3/78 "2020-03-12T15:30:40Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![Evgeny\_Nikolin](https://avatars.discourse-cdn.com/v4/letter/e/e47c2d/32.png) [@Evgeny\_Nikolin](https://discourse.greatexpectations.io/u/Evgeny_Nikolin)\
**Post date:** [March 12, 2020, 3:30pm UTC](https://discourse.greatexpectations.io/t/how-to-configure-a-pyspark-datasource-for-accessing-the-data-from-aws-s3/78/1 "2020-03-12T15:30:40Z")

</div>

Current version of Great Expectation framework documentation (0.9.4) does not contain any samples of how to configure a PySpark datasource in order to access the AWS S3 files. It would be really helpful if there is any example of it’s configuration.

---

<div class="post-metadata">

**Author:** ![jpcampbell42](https://avatars.discourse-cdn.com/v4/letter/j/838e76/32.png) [@jpcampbell42](https://discourse.greatexpectations.io/u/jpcampbell42)\
**Post date:** [March 28, 2020, 12:59pm UTC](https://discourse.greatexpectations.io/t/how-to-configure-a-pyspark-datasource-for-accessing-the-data-from-aws-s3/78/2 "2020-03-28T12:59:11Z")

</div>

You’re right! In fact, it’s really similar to the example for pandas, since spark’s reader methods also know how to process s3 paths:

```auto
datasources:
  nyc_taxi:
    class_name: SparkDFDatasource
    generators:
      s3:
        class_name: S3GlobReaderBatchKwargsGenerator
        bucket: nyc-tlc
        delimiter: '/'
        reader_options:
          sep: ','
          engine: python
        assets:
          taxi-green:
            prefix: trip data/
            regex_filter: 'trip data/green.*\.csv'
          taxi-fhv:
            prefix: trip data/
            regex_filter: 'trip data/fhv.*\.csv'
    data_asset_type:
      class_name: SparkDFDataset

```
