Senior Data Engineer for Batch Pipelines
Publicada el 2026-07-21
Descripción de la oferta
I need a robust, production-ready batch processing pipeline that can reliably pull data from several relational databases, perform the required transformations, and load the cleansed results into our analytics environment on a scheduled basis. Scope • Design and build the end-to-end pipeline architecture—ingestion, transformation, validation, and load. • Optimise for scalability and fault tolerance; orchestration should handle retries and alerting gracefully. • Source systems will be only relational databases for now (PostgreSQL and MySQL are the main ones), but the design should allow new sources to be plugged in later without major refactoring. • All code must be version-controlled and deployable through our existing CI/CD setup (GitHub Actions). Tech context Our stack already uses Python 3.10, Apache Airflow, and dbt; feel free to propose alternative open-source tools if they make scheduling, monitoring, or performance easier to manage. Acceptance criteria 1. A parameterised Airflow DAG (or equivalent) that picks up incremental changes from the relational sources. 2. Data quality checks that fail fast and log detailed diagnostics. 3. Automated unit tests covering the critical transformation logic. 4. Clear README documenting setup, configuration, and typical run times. Once the solution runs successfully against a QA database and passes the tests, I will promote it to production.
Skills
- PostgreSQL
- MySQL
- Python
- CI/CD
- Linux
- Airflow
- ETL
- QA Testing
- GitHub Actions
- GitHub
- dbt
- SQL
- Apache Spark
- Data Engineering
Fuente original: freelancer