Site Reliability Engineer
Job Reference: 161328
Industry: Information and Communications Technology
Bonus Package: R22108699
Salary Description: 02C3423
We are looking for a Site Reliability Engineer (SRE) with experience in platform engineering, DevOps, and production operations. The role involves building reliable systems, automating infrastructure, and ensuring observability across mission-critical applications. You will work closely with development, infrastructure, and operations teams to design scalable solutions, improve system resilience, and define operational best practices.
Key Responsibilities
-
System reliability: Ensure high availability, performance, and resilience of production systems.
-
Infrastructure automation: Build and maintain CI/CD pipelines, automate deployments, and manage containerized workloads.
-
Observability & monitoring: Implement logging, metrics, tracing, alerting, and dashboards using modern monitoring tools.
-
Incident response: Integrate systems with enterprise monitoring, SIEM, and incident management workflows; participate in on-call rotations.
-
Operational excellence: Define and implement runbooks, escalation paths, and production support models.
-
Collaboration: Work with cross-functional teams in Agile environments to deliver reliable and secure systems.
-
Continuous improvement: Conduct root cause analysis, implement permanent fixes, and drive automation initiatives to reduce manual effort.
Skills Required
-
5+ years of experience as an SRE, DevOps Engineer, Platform Engineer, Infrastructure Engineer, or Production Engineer.
-
Solid knowledge of Linux administration, Shell scripting, Git, CI/CD, Docker, and Kubernetes.
-
Hands-on experience with observability platforms (Grafana, Prometheus, ELK, CloudWatch, Azure Monitor, etc.).
-
Experience integrating systems with enterprise monitoring, alerting, SIEM, or incident response workflows.
-
Proven ability to define and implement runbooks, operational procedures, escalation paths, and production support models.
-
Familiarity with cloud platforms (AWS, Azure, GCP) and infrastructure-as-code tools (Terraform, Ansible).
-
Excellent problem-solving skills, ability to troubleshoot complex systems, and experience in Agile/DevOps practices.
-
Excellent communication skills and ability to collaborate with cross-functional stakeholders.
