Job Req ID: 61834
Posting Date: 03 Sep 2026
Location: Bengaluru
Salary: Competitive
About the role
We are looking for an experienced Site Reliability Engineer (SRE) to ensure the reliability, performance, and operational excellence of critical customer-facing applications across Siebel, Java, and AWS platforms. The role is responsible for understanding end-to-end business processes, application architecture, system integrations, and customer journeys to effectively manage service reliability and drive continuous improvement. Working with support partners, engineering teams, and business stakeholders, the role uses observability, operational data, and technical analysis to identify issues, improve application resilience, increase automation, and enhance customer and business outcomes. The role also supports the adoption of AIOps and self-healing capabilities to improve operational efficiency and service quality.
What you’ll be doing
• Develop a strong understanding of customer journeys, business processes, application architecture, and integrations across CRM platform, AWS, and supporting platforms to effectively support service outcomes.
• Monitor service health, performance, and availability using observability platforms and operational metrics, proactively identifying risks and improvement opportunities.
• Investigate complex production issues through analysis of application logs, database queries, transaction flows, and monitoring data to support rapid resolution and root cause identification.
• Provide technical oversight on investigations, validating root causes, challenging assumptions, and ensuring high-quality incident resolution.
• Driving incident, problem, and change management activities, leveraging root cause analysis, preventive actions, and continuous service improvement to enhance reliability and prevent recurring issues.
• Partner with engineering, business, and supplier teams to deliver application, operational, and customer experience improvements aligned to business objectives.
• Develop and maintain service dashboards, operational reporting, and performance insights to support governance, decision-making, and continuous improvement.
• Identify and implement opportunities for automation, self-service, and AI-driven operational capabilities to improve efficiency, resilience, and customer outcomes
Essential Skills / Experience
• Experience supporting or operating enterprise applications within complex production environments.
• Strong understanding of application architecture, system integrations, data flows and customer-facing business processes.
• Ability to analyse application, middleware and infrastructure logs to identify issues and support root cause investigations.
• Knowledge of SQL/Oracle databases with experience troubleshooting production application issues.
• Hands-on experience with monitoring and observability platforms such as Dynatrace, including dashboard creation, alert analysis and trend identification.
• Understanding of Java-based applications, APIs, integration services and AWS-hosted platforms.
• Experience working with incident, problem, change and release management processes.
• Strong analytical and stakeholder management skills with the ability to challenge technical investigations, drive service improvements and influence engineering teams.
• Ability to identify automation, resilience and customer experience improvement opportunities through operational insights.
• Exposure to scripting or automation technologies such as Shell script, Python etc.
• Experience leveraging AI/ML/GenAI capabilities to improve service reliability, operational efficiency, observability, incident response, root cause analysis, and self-healing automation.
• Understanding of application and platform performance management, capacity planning, trend analysis, demand forecasting, and resilience engineering for business-critical production services.
Desirable Skills / Experience
• Experience with Site Reliability Engineering (SRE) principles, including service reliability, availability, and operational excellence.
• Knowledge of AWS Cloud services and Java application support in production environments.
• Strong understanding of SQL and Oracle database analysis, troubleshooting, and performance tuning.
• Experience with observability and monitoring tools such as Dynatrace and AWS CloudWatch.
• Skilled in incident management, problem management, root cause analysis (RCA), and using ServiceNow for service operations.
• Experience in automation, AIOps, and AI/GenAI-driven operations, with Python scripting being highly desirable.
BT is the UK’s leading communications group and the holding company behind some of the country’s most recognised brands – including BT, EE, Openreach and Plusnet. Our purpose is as simple as it is ambitious: we connect for good. Our customers include consumers, small, medium and large businesses, public sector organisations and other communications providers.
Having come through the most capital-intensive phase of our fibre investment, our focus now is on what comes next – simplifying how we operate, using technology and AI to work smarter, and organising ourselves to serve customers better and grow sustainably.
We have a singular culture that unites all our people: we are customer-first challengers, who are committed, clear and connected. These behaviours unite us as one team to deliver for our colleagues, our customers, our stakeholders and the country. Joining BT means working at the heart of a business that matters to the UK, with the opportunity to shape decisions, influence outcomes and help set the future course of one of the country’s most important companies.