Custom Data Pipeline Engineering
Publicada el 2026-07-27
Descripción de la oferta
I have an in-house project that needs a fully custom-built data pipeline. The goal is to ingest Parquet files arriving in our file system, apply the required transformations, and move the cleaned data on to our analytics environment. Off-the-shelf ETL tools do not meet our requirements, so I am looking for a bespoke solution designed from the ground up. Scope of work • Design the end-to-end pipeline architecture, choosing the most suitable framework (for example Python, Spark, or similar) while keeping future scalability in mind. • Build robust code that can automatically detect new Parquet files, validate their schema, transform the data as specified, and deliver the output to the target location I will provide once we start. • Add logging, error handling, and simple configuration files so the pipeline can be tweaked without code changes. • Package everything in a repo with clear setup instructions and a lightweight README so I can deploy it internally. Acceptance criteria 1. End-to-end execution succeeds on a supplied Parquet sample set. 2. Any corrupted or schema-drifted file is logged and quarantined without halting the run. 3. All configurable paths, credentials, and parameters sit outside the core code. 4. A short hand-off call and walkthrough of the repository. If you have built similar customer-made pipelines that process Parquet data straight from file systems, I’d like to see an example commit or short demo clip. Let me know your proposed tech stack and estimated timeline, and we can get started right away.
Skills
Fuente original: freelancer