AWS Glue MySQL-S3 ETL

Cliente Freelancer · Remoto · Remoto · freelance · mid · 8–15 USD

Publicada el 2026-07-20

Descripción de la oferta

My Terraform code already spins up the MySQL source, VPC, subnets and security groups; Glue can reach the database without additional networking work. What I still need is the Glue side itself—fully defined in Terraform—so the data moves from that MySQL backend into the raw tier of my data lake. Here is what has to happen: • The Glue job (PySpark preferred) connects to MySQL via JDBC, performs the required data cleansing, then writes the output as Parquet files into a specified S3 prefix. • All Glue resources—Job, Connection, IAM role, catalog tables, optional crawler—must be described in Terraform so a single terraform apply recreates the stack in any account. • Partitioning or dynamic frame optimisations that make downstream analytics easier are welcome, as long as the raw tier remains untouched beyond the Parquet conversion. Acceptance criteria 1. Running the Terraform plan from a clean account provisions every Glue asset with no manual edits. 2. A test run copies at least one representative table, shows the cleansing logic executed, and lands Parquet files in the configured S3 path. 3. Clear README or inline comments explaining variables, how to extend to more tables, and how the cleansing rules can be customised. I will handle bucket creation and the network; your deliverable is the Terraform module plus the Glue script(s) and any wrapper code required to execute them. If you have delivered similar MySQL-to-Parquet pipelines with AWS Glue and Terraform, please point me to them when you bid.

Skills

Fuente original: freelancer

Análisis JobHunter