# Configure datasource for json files

**URL:** <https://discourse.greatexpectations.io/t/configure-datasource-for-json-files/121>\
**Category:** Archive\
**Created:** [May 13, 2020, 4:44pm UTC](https://discourse.greatexpectations.io/t/configure-datasource-for-json-files/121 "2020-05-13T16:44:17Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![ramon.oliveira](https://avatars.discourse-cdn.com/v4/letter/r/898d66/32.png) [@ramon.oliveira](https://discourse.greatexpectations.io/u/ramon.oliveira)\
**Post date:** [May 13, 2020, 4:44pm UTC](https://discourse.greatexpectations.io/t/configure-datasource-for-json-files/121/1 "2020-05-13T16:44:17Z")

</div>

Hey There,

Anyone knows how to properly configure the datasource for json files composed with lines (one json per line)?  
Here’s my datasource configuration at the moment:

```auto
datasources:
  raw:
    class_name: PandasDatasource
    data_asset_type:
      class_name: PandasDataset
      module_name: great_expectations.dataset
    batch_kwargs_generators:
      subdir_reader:
        class_name: SubdirReaderBatchKwargsGenerator
        base_directory: ../data/raw
    module_name: great_expectations.datasource

```

Using pandas I can read my file with pd.read\_json(‘data/raw/file.json’, lines=True). How can I configure a file like this in the datasources?

I tried configuring `batch_kwargs_generators` like this with no luck:

```auto
batch_kwargs_generators:
  subdir_reader:
    class_name: SubdirReaderBatchKwargsGenerator
    base_directory: ../data/raw
    reader_method: read_json
    reader_options:
      lines: true 

```

---

<div class="post-metadata">

**Author:** ![eugene.mandel](https://yyz1.discourse-cdn.com/flex031/user_avatar/discourse.greatexpectations.io/eugene.mandel/32/22_2.png) [@eugene.mandel](https://discourse.greatexpectations.io/u/eugene.mandel)\
**Post date:** [May 13, 2020, 10:04pm UTC](https://discourse.greatexpectations.io/t/configure-datasource-for-json-files/121/2 "2020-05-13T22:04:45Z")

</div>

Your second configuration snippet is correct - if you add “line: true” under reader\_options, this option will be passed to pandas.read\_json and will read a file with one JSON object per line correctly.

I suspect the only reason this is not working yet is the relative path to the base\_directory. If you are using a relative path to data, make sure that it is relative to the “great\_expectations” directory in your project (this is the directory where great\_expectations.yml is located in).

---

<div class="post-metadata">

**Author:** ![ramon.oliveira](https://avatars.discourse-cdn.com/v4/letter/r/898d66/32.png) [@ramon.oliveira](https://discourse.greatexpectations.io/u/ramon.oliveira)\
**Post date:** [May 18, 2020, 3:27pm UTC](https://discourse.greatexpectations.io/t/configure-datasource-for-json-files/121/3 "2020-05-18T15:27:05Z")

</div>

I managed to fix with your suggestion. Thanks!

---

<div class="post-metadata">

**Author:** ![kuberketes](https://avatars.discourse-cdn.com/v4/letter/k/e5b9ba/32.png) [@kuberketes](https://discourse.greatexpectations.io/u/kuberketes)\
**Post date:** [December 12, 2020, 10:23pm UTC](https://discourse.greatexpectations.io/t/configure-datasource-for-json-files/121/4 "2020-12-12T22:23:06Z")

</div>

I have the same issue but my raw jsons are in a azure datalake filesystem.  
Would it be possible to use ge to validate them?
