Job Req ID: 61824
Posting Date: 20 Aug 2026
Location: Bengaluru
Salary: Competitive
About the role
As part of a team of skilled engineers, this role contributes to platform engineering, cloud infrastructure delivery, and operational support for the Domain Services integration platform. Working under the guidance of senior specialists, the role is responsible for implementing infrastructure solutions, maintaining CI/CD pipelines, and supporting platform resilience. The position collaborates closely with DevOps, Security, and Service Assurance teams to deliver reliable cloud services and participate in incident response activities. Key responsibilities include implementing Infrastructure as Code, supporting containerised deployments, maintaining observability tooling, and contributing to DevSecOps practices throughout the platform lifecycle.
Key goals:
- Implement and maintain cloud infrastructure for Domain Services, supporting 80+ microservices across AWS and on-prem environments.
- Execute automation, monitoring, and deployment tasks that support platform reliability and operational efficiency.
- Follow security-by-design practices and contribute to DevSecOps pipeline maintenance.
- Support observability and alerting improvements using platform telemetry and log analytics.
What you’ll be doing
- Implement and maintain cloud infrastructure using IaC tools (Terraform, CloudFormation, Helm) following established patterns and standards.
- Deploy and manage containerised workloads on Kubernetes (EKS/ECS) — supporting scaling, upgrades, and troubleshooting.
- Maintain CI/CD pipelines — build configurations, automated testing integrations, and deployment workflows.
- Support resiliency practices — implementing health checks, configuring auto-scaling, and supporting failover testing.
- Contribute to DevSecOps pipeline maintenance — running SAST/SCA scans, managing secrets, and triaging vulnerability findings.
- Monitor platform health using observability tools (CloudWatch, Prometheus, Grafana) — responding to alerts and escalating as needed.
- Perform log analysis and assist with root cause investigation for production incidents across distributed services.
- Write automation scripts (Python, Bash) for operational tasks — deployment, cleanup, reporting, and monitoring.
- Maintain documentation for infrastructure configurations, runbooks, and operational procedures.
- Participate in on-call rotation and incident response following established playbooks
Essential Skills / Experience
- Experience in cloud infrastructure engineering, including hands-on support and deployment of containerised applications in cloud-native environments.
- Working knowledge of AWS services, including EC2, ECS/EKS, S3, RDS, CloudWatch, IAM, and VPC.
- Understanding of Kubernetes fundamentals, including deployments, services, pods, ConfigMaps, and secrets.
- Experience with Infrastructure as Code (IaC) tools such as Terraform or CloudFormation, including the ability to read, modify, and deploy infrastructure configurations.
- Proficiency in Linux administration and shell scripting (Bash) for operational and automation activities.
- Experience building and maintaining CI/CD pipelines using tools such as GitLab CI, Jenkins, or GitHub Actions.
- Knowledge of container technologies, including Docker and Helm.
- Familiarity with monitoring, observability, and alerting tools such as CloudWatch, Prometheus, and Grafana.
- Programming or scripting experience in Python or Java to support automation and platform operations.
- Experience using Git version control and standard branching and collaboration workflows.
- Ability to troubleshoot deployment, infrastructure, and application runtime issues across multiple service layers.
- Strong analytical, problem-solving, and collaboration skills, with the ability to work effectively within cross-functional engineering teams.
Desirable Skills / Experience
- Knowledge of DevSecOps practices and tools, including SonarQube, Black Duck, Trivy, and secret-scanning solutions.
- Familiarity with log analytics and troubleshooting using tools such as ELK Stack and CloudWatch Logs Insights.
- Understanding of networking concepts, including VPCs, subnets, security groups, load balancers, and DNS.
- Exposure to performance and load testing tools such as JMeter or k6.
- Experience with GitOps methodologies and deployment tools such as ArgoCD.
- Familiarity with event-driven architectures and AWS messaging services, including SQS, SNS, and EventBridge.
- Understanding of microservices architecture principles and distributed application design patterns.
- Exposure to cloud security, compliance, and operational best practices in enterprise environments.
- Experience working within Agile, DevOps, or site reliability engineering (SRE) teams.
- Awareness of observability practices, including logging, metrics, tracing, and proactive service monitoring.
BT is the UK’s leading communications group and the holding company behind some of the country’s most recognised brands – including BT, EE, Openreach and Plusnet. Our purpose is as simple as it is ambitious: we connect for good. Our customers include consumers, small, medium and large businesses, public sector organisations and other communications providers.
Having come through the most capital-intensive phase of our fibre investment, our focus now is on what comes next – simplifying how we operate, using technology and AI to work smarter, and organising ourselves to serve customers better and grow sustainably.
We have a singular culture that unites all our people: we are customer-first challengers, who are committed, clear and connected. These behaviours unite us as one team to deliver for our colleagues, our customers, our stakeholders and the country. Joining BT means working at the heart of a business that matters to the UK, with the opportunity to shape decisions, influence outcomes and help set the future course of one of the country’s most important companies.