Site Reliability Engineer

Experian · Cyberjaya, my

onsitefull-time3-6 years

posted 15h

Role Summary Our Product Development team is looking for a Site Reliability Engineer to help build reliable, scalable, and high-performing products. As we transform into a product-led organization, you will play a key role in improving system reliability, operational efficiency, and cost effectiveness through automation, monitoring, and data-driven insights. You will collaborate with technical experts, contribute to our future technology strategy, and help deliver greater value to our clients. The ideal candidate is a engineer with strong experience designing and implementing solutions on AWS, together with a passion for continuous improvement, and operational excellence. What You'll Be Doin Delivery of high-quality infrastructure as code solutions and CI/CD pipelines Implementation of monitoring solutions for client-facing products and internal data pipelines, including intelligent alarming for quicker incident detection and resolution Track and manage infrastructure & application vulnerabilities You will be working closely with software engineers, data engineers, data scientists and product managers to ensure smooth deployment and operation of systems Troubleshoot issues across production and non-production environments, conduct root cause analysis, and implement preventive solutions to strengthen platform reliability. You will be identifying the resource inefficiencies, optimize platform performance and costs, and progress, risks, and recommendations to stakeholders. Development of and adherence to policies and procedures to ensure all work is executed to a high standard You will be reporting to Software Engineering Manager Qualifications 5 years+ of experience in Site Reliability Engineering, DevOps, Cloud Engineering, or a related role. Proven experience managing AWS workloads, particularly ECS and EKS clusters . Strong understanding of high availability, disaster recovery, platform and performance monitoring. Hands-on experience with DevOps and Infrastructure-as-Code tools, including Terraform and Jenkins . 3+ years of experience using data and monitoring insights to improve AWS performance, reliability, and cost efficiency. Complex technical matters conveyed clearly and drive alignment across technical and product teams. Strong mindset with the initiative to innovate, automate processes, and continuously improve platform performance.

Site Reliability Engineer at Experian — TalentDesi