Senior Staff Engineer - Site Reliability

Freshworks · Hyderabad, TS, in

onsitefull-time6-10 years

posted 1d

● Design, write, and deliver software to improve the availability, latency, and efficiency of Freshworks’ Products & Platforms.   ● Develop scalable, cloud-native architectures that support business growth.  ● Design and implement self-healing and auto-scaling mechanisms.  ● Manage availability, latency and performance of mission critical services and build automation to prevent problem recurrence.   ● Independently determine and develop architectural approaches and Infrastructure solutions.   ● Define strategy, vision, and roadmap adhering to well architected principles of performance, availability, scalability, performance and resilience  ● Experience with AI-driven system optimization or predictive analytics for IT operations.  ● Drive blameless postmortems for large scale incidents.   ● Define and drive automation and orchestration strategies.   ● Strategize cost optimization across Freshworks Cloud environment.   10-15 years of experience in SRE handling performance, architecture and design of applications ● Strong understanding of cloud computing, networking, Linux systems administration, containerization (e.g., Docker, Kubernetes), and infrastructure as code (e.g., Terraform, Ansible) ● Understanding of SRE principles, including SLOs, SLIs, SLAs, and error budgets. ● Experience in managing incident and retrospectives ● Experience in cloud cost management, cloud architecture ● In-depth knowledge of cloud computing platforms (e.g., AWS) ● Experience with infrastructure as code (IaC) tools and practices ● Experience with monitoring, logging & telemetry tools like New Relic, Splunk, ELK, Nagios, SolarWinds, Prometheus, AWS Cloudwatch, Datadog, Opentelemetry ● Expert in designing, creating and supporting Automation and Identify opportunities for self-healing systems, automated deployments, and other scalable solutions. ● Experience in performance engineering and identify opportunities for performance tuning and profiling ● Experience in prioritizing and managing technical roadmaps. ● Strong skills in stakeholder communication, requirements gathering, and documentation. ● Ability to work with cross-functional teams and build consensus around reliability goals ● Improve operational processes and team practices ● Problem-solving: Ability to analyze complex systems, troubleshoot issues, and devise effective solutions ● Excellent communication skills, with the ability to inspire and motivate cross-functional teams. ● Experience in dealing with the intricacies of large-scale distributed systems and ensuring their reliability and performance.