Generating Parquet Files for the Amazon S3 Integration Using Python
Learn how to generate Parquet files for the Amazon S3 integration using Python.
Overview
Use the following Python script to convert an NDJson file containing events into a Parquet file for use with the Amazon S3 integration with Split.
Prerequisites
The following environments:
Python 3.7
Pandas 1.2.2
Pyarrow 3.0.0
ndjson 0.3.1
Prepare your event file
The NDJSON file must contain a valid structure for Split events.
For example:
{
"environmentId": "029bd160-7e36-11e8-9a1c-0acd31e5aef0",
"trafficTypeId": "e6910420-5c85-11e9-bbc9-12a5cc2af8fe",
"eventTypeId": "agus_event",
"key": "gDxEVkUsd3",
"timestamp": 1625088419634,
"value": 86.5588664346,
"properties": {
"age": "51",
"country": "argentina",
"tier": "basic"
}
}Replace the file names and paths in the
input_fileandoutput_filevariables in the script.Replace the values for:
key(Customer key)value(Event value)Property names and values under
properties
Make sure the timestamp field uses Epoch time in milliseconds.
To add more events, repeat the append line:
df = df.append(...).The resulting Parquet file can be copied to the Amazon S3 bucket used for Split’s event ingestion.
Run the Python script
Last updated
Was this helpful?