Data Engineer (AWS, PySpark)
reference-number: 160674
industry: Information and Communications Technology
brand-id: R1434374
brand-name: 02C3423
Key Responsibilities
- Design, develop, and maintain scalable and high-performance data pipelines using Python and PySpark.
- Build, deploy, and manage ETL/ELT workflows using AWS Glue and AWS Step Functions.
- Develop serverless applications and automation solutions using AWS Lambda.
- Write clean, efficient, maintainable, and reusable code following software engineering best practices.
- Provision and manage cloud infrastructure using Terraform and Infrastructure as Code (IaC) principles.
- Design, implement, and maintain CI/CD pipelines to automate code deployment, testing, and infrastructure changes.
- Monitor, troubleshoot, and optimize data pipelines to ensure performance, reliability, and cost efficiency.
- Implement logging, monitoring, and alerting solutions using AWS CloudWatch and other observability tools.
- Collaborate with business stakeholders, solution architects, and technology teams to understand requirements and deliver effective data solutions.
- Support production deployments, incident resolution, root cause analysis, and continuous improvement initiatives.
- Ensure all solutions adhere to security, compliance, DevOps, and coding standards.
Requirements
- Bachelor's Degree in Computer Science, Information Technology, Software Engineering, Data Science, or a related discipline.
- Minimum 5 years of experience in Data Engineering, with at least 3 years of hands-on experience working on AWS cloud platforms.
- Hands-on experience with Python and PySpark for large-scale data processing and transformation.
- Proven experience developing and managing ETL/ELT pipelines using AWS Glue.
- Proficiency in workflow orchestration using AWS Step Functions.
- Experience building serverless applications and integrations using AWS Lambda.
- Hands-on experience with Terraform for Infrastructure as Code (IaC).
- Experience implementing and managing CI/CD pipelines using tools such as GitHub Actions, GitLab CI/CD, Jenkins, Azure DevOps, or similar platforms.
- Good understanding of AWS services, including S3, IAM, CloudWatch, and data lake architectures.
- Experience in monitoring, troubleshooting, and optimizing cloud-based data solutions.
- Experience working with large-scale data platforms and distributed data processing environments.
- Excellent communication and stakeholder management skills, with the ability to work effectively across technical and business teams.
- AWS certifications such as AWS Certified Data Engineer, AWS Certified Solutions Architect, or AWS Certified Developer are good to have.
