Site Reliability Engineer

HighRadius · Hyderabad, Telangana, India

onsitefull-time3-6 years

posted 1d

Sign in to apply

<p><strong>About Us&nbsp;</strong></p> <p>HighRadius, a renowned provider of cloud-based Autonomous Software for the Office of the CFO, has transformed critical financial processes for over 800+ leading companies worldwide. Trusted by prestigious organizations like 3M, Unilever, Anheuser-Busch InBev, Sanofi, Kellogg Company, Danone, Hershey's, and many others, HighRadius optimizes order-to-cash, treasury, and record-to-report processes, earning us back-to-back recognition in Gartner's Magic Quadrant and a prestigious spot in Forbes Cloud 100 List for three consecutive years.&nbsp;</p> <p>With a remarkable valuation of $3.1B and an impressive annual recurring revenue exceeding $100M, we experience a robust year-over-year growth of 24%. With a global presence spanning 8+ locations and a recent addition in Poland, we're in the pre-IPO stage, poised for rapid growth. We invite passionate and diverse individuals to join us on this exciting path to becoming a publicly traded company and shape our promising future.&nbsp;</p> <p><strong>Job Summary:</strong><br>We are looking for a highly skilled and adaptable Site Reliability Engineer (7+ Years) to become a key member of our Cloud Engineering team. In this crucial role, you will be instrumental in designing and refining our cloud infrastructure with a strong focus on reliability, security, and scalability. As an SRE, you'll apply software engineering principles to solve operational challenges, ensuring the overall operational resilience and continuous stability of our systems. This position requires a blend of managing live production environments and contributing to engineering efforts such as automation and system improvements.<br><br><strong>Key Responsibilities:</strong><br>● Cloud Infrastructure Architecture and Management: Design, build, and maintain resilient cloud&nbsp;infrastructure solutions to support the development and deployment of scalable and reliable&nbsp;applications. This includes managing and optimizing cloud platforms for high availability,&nbsp;performance, and cost efficiency.<br>● Enhancing Service Reliability: Lead reliability best practices by establishing and managing&nbsp;monitoring and alerting systems to proactively detect and respond to anomalies and&nbsp;performance issues. Utilize SLI, SLO, and SLA concepts to measure and improve reliability. Identify&nbsp;and resolve potential bottlenecks and areas for enhancement.<br>● Driving Automation and Efficiency: Contribute to the automation, provisioning, and&nbsp;standardization of infrastructure resources and system configurations. Identify and implement&nbsp;automation for repetitive tasks to significantly reduce operational overhead. Develop Standard&nbsp;Operating Procedures (SOPs) and automate workflows using tools like Rundeck or Jenkins.<br>● Incident Response and Resolution: Participate in and help resolve major incidents, conduct&nbsp;thorough root cause analyses, and implement permanent solutions. Effectively manage&nbsp;incidents within the production environment using a systematic problem-solving approach.<br>● Collaboration and Innovation: Work closely with diverse stakeholders and cross-functional&nbsp;teams, including software engineers, to integrate cloud solutions, gather requirements, and&nbsp;execute Proof of Concepts (POCs). Foster strong collaboration and communication. Guide&nbsp;designs and processes with a focus on resilience and minimizing manual effort. Promote the&nbsp;adoption of common tooling and components, and implement software and tools to enhance&nbsp;resilience and automate operations. Be open to adopting new tools and approaches as&nbsp;needed.</p> <p><strong>Required Skills and Experience:</strong></p> <p>● Experience: We are looking for 7+ Years of industry experience .<br>● Cloud Platforms: Demonstrated expertise in at least one major cloud platform (AWS, Azure, or&nbsp;GCP). Extensive experience with containerization (Docker) and orchestration (Kubernetes)&nbsp;techn