Data Engineer (AWS, PySpark)

  •  Job Reference: 160674
  •  Industry: Information and Communications Technology
  •  Bonus Package: R1434374
  •  Salary Description: 02C3423

Key Responsibilities

  • Design, develop, and maintain scalable and high-performance data pipelines using Python and PySpark.
  • Build, deploy, and manage ETL/ELT workflows using AWS Glue and AWS Step Functions.
  • Develop serverless applications and automation solutions using AWS Lambda.
  • Write clean, efficient, maintainable, and reusable code following software engineering best practices.
  • Provision and manage cloud infrastructure using Terraform and Infrastructure as Code (IaC) principles.
  • Design, implement, and maintain CI/CD pipelines to automate code deployment, testing, and infrastructure changes.
  • Monitor, troubleshoot, and optimize data pipelines to ensure performance, reliability, and cost efficiency.
  • Implement logging, monitoring, and alerting solutions using AWS CloudWatch and other observability tools.
  • Collaborate with business stakeholders, solution architects, and technology teams to understand requirements and deliver effective data solutions.
  • Support production deployments, incident resolution, root cause analysis, and continuous improvement initiatives.
  • Ensure all solutions adhere to security, compliance, DevOps, and coding standards.

 

Requirements

  • Bachelor's Degree in Computer Science, Information Technology, Software Engineering, Data Science, or a related discipline.
  • Minimum 5 years of experience in Data Engineering, with at least 3 years of hands-on experience working on AWS cloud platforms.
  • Hands-on experience with Python and PySpark for large-scale data processing and transformation.
  • Proven experience developing and managing ETL/ELT pipelines using AWS Glue.
  • Proficiency in workflow orchestration using AWS Step Functions.
  • Experience building serverless applications and integrations using AWS Lambda.
  • Hands-on experience with Terraform for Infrastructure as Code (IaC).
  • Experience implementing and managing CI/CD pipelines using tools such as GitHub Actions, GitLab CI/CD, Jenkins, Azure DevOps, or similar platforms.
  • Good understanding of AWS services, including S3, IAM, CloudWatch, and data lake architectures.
  • Experience in monitoring, troubleshooting, and optimizing cloud-based data solutions.
  • Experience working with large-scale data platforms and distributed data processing environments.
  • Excellent communication and stakeholder management skills, with the ability to work effectively across technical and business teams.
  • AWS certifications such as AWS Certified Data Engineer, AWS Certified Solutions Architect, or AWS Certified Developer are good to have.