06 Aug
|
BT Group
|
Bengaluru
06 Aug
BT Group
Bengaluru
Job Title: Software Engineering Specialist Req ID: 61185 Job Function: Software Engineering Posting Start Date: 04/08/2026 Posting End Date: 07/08/2026 Division: Networks Job Location: IND-Bengaluru-Pritech Advertised Salary: Competitve Recruiter: Nishita Jena Hiring Manager: Hari Annamalai Career Grade: D About The Role We are looking for a System Architect for Observability & Monitoring Platforms with 15+ years of experience who will own and drive the end-to-end architecture of a large-scale infrastructure monitoring and cloud orchestration platform. This role focuses on system architecture, technical strategy, and long-term platform evolution, enabling the platform to help metrics, logs, distributed tracing, and application performance monitoring at scale. This role combines judicious architectural ownership with hands-on validation through design reviews, code reviews, and targeted proof-of-concepts (POCs).
What You’ll Be Doing Architecture Ownership & Vision Own the overall system architecture of the observability platform across ingestion, processing, storage, and query layers.
Work closely with Enterprise Architect to align system architecture with long-term architectural vision and technical roadmap aligned with business and platform goals.
Create and review high-level system architecture, data flows, integration patterns, and core technology choices.
Act as the technical authority for complex architectural decisions, trade-offs, and create reviews. Platform & Domain Leadership Architect systems for infrastructure monitoring, metrics, logs, and distributed tracing at scale.
Guide the evolution from infrastructure monitoring to a full observability platform including APM.
Define architectural patterns for high-throughput telemetry ingestion, real-time processing, and query-at-scale.
Ensure architectural consistency for multi-tenant, cloud-native distributed platforms. Standards, Governance & Enablement Establish and evolve architecture standards, design-principles, and best practices across teams.
Identify architectural risks and proactively drive mitigation strategies.
Enable teams through reference architectures, design-frameworks, and technical guidance.
Support modernization initiatives including scalability, performance optimization, resilience, and cost efficiency. Hands-On Architectural Validation Perform code reviews of critical and performance-aware components.
Build and guide proof-of-concepts (POCs) to formalize architectural decisions and reduce risk new technologies.
Develop reference implementations to demonstrate architectural intent.
Collaborate closely with senior engineers to troubleshoot complex system-level issues Delivery Collaboration & Execution Enablement Work closely with Software Engineering Managers to align architectural decisions with delivery and release plans.
Assist in breaking down large architectural initiatives into phased, incremental deliverables.
Identify architectural and technical risks early and proactively surface them to influence release planning.
Support release readiness by formalizing that architecture, scalability, and non‑functional requirements are addressed ahead of key milestones.
Essential
Skills / Experience Strong foundation in system design, distributed systems, and application scalability
Experience creating microservice based, event driven architectures
Ability to make architectural trade offs involving scalability, reliability, performance, and cost
Demonstrate experience creating large scale, cloud based distributed platforms
Solid backend experience with Java (Spring Boot–based microservices)
Working knowledge of Python and/or Go for scripting, automation, and collectors
Ability to review and reason about performance critical backend code
Strong understanding of infrastructure monitoring and observability concepts with hands on experience on follow:
Metrics, logs, and distributed tracing
Agent based and agentless monitoring models
Push / pull / subscription based data collection
Hands on or deep working knowledge of:
Metric data models and storage (e.g., VictoriaMetrics or equivalent)
Log aggregation and search platforms (Elasticsearch / OpenSearch)
Familiarity with OpenTelemetry, exporters, and custom collectors
Understanding of common monitoring tools and protocols (e.g., SNMP, Prometheus style systems)
Robust experience with
Time series databases for metrics
Relational databases (PostgreSQL / MySQL) for metadata and control plane services
NoSQL / in memory stores (e.g., Redis) for high volume or low latency workloads
Working knowledge of search and analytics engines for logs and traces
Understanding of data modeling, query patterns, and data lifecycle management
Experience with Kafka class distributed messaging systems
Designing and operating event driven telemetry ingestion pipelines
Understanding of throughput, backpressure, and reliability concerns in streaming systems
Strong hands on exposure to Docker and Kubernetes based platforms
Infrastructure automation using
Terraform (IaC)
Ansible (deployment and configuration automation)
CI/CD pipeline understanding and usage (e.g., GitLab CI or equivalent)
Familiarity with Grafana, Kibana, or similar visualization platforms
Awareness of frontend technologies such as Angular (architecture level understanding)
Experience with OpenStack, MAAS, Juju, or similar cloud/service orchestration platforms
Exposure to Canonical Observability Stack (COS) or equivalent platforms
Ability to perform code reviews for critical and performance responsive components
Experience building or guiding proof of concepts (POCs) to formalize architecture
Strong collaboration with Engineering Managers and senior engineers
Clear technical communication and mentoring capability Desirable Skills / Experience Experience with columnar / methodical databases (e.g., ClickHouse or equivalent)
Exposure to graph databases (Neo4j or similar) for relationship heavy domains
Experience with cloud native deployments in AWS, Azure, or GCP
Experience with secrets management (Vault, cert manager, etc.)
Understanding of secure ingestion pipelines, RBAC, and tenant isolation
Awareness of compliance and security considerations in telemetry platforms
Exposure to Telecom OSS / large scale infrastructure observability platforms
Familiarity with monitoring agents and collectors (e.g., NRPE, Filebeat, Telegraf) Our Package BT Group is the UK’s leading communications group and the holding company behind some of the country’s most recognised brands – including BT, EE, Openreach and Plusnet.
Our purpose is as simple as it is ambitious: we connect for good. Our customers include consumers, small, medium and large businesses, public sector organisations and other communications providers. BT Group’s role is about setting direction, unlocking value and creating the conditions for our brands and businesses to thrive.
Having come through the most capital-intensive phase of our fibre investment, our focus now is on what comes next – simplifying how we operate, using technology and AI to work smarter, and organising ourselves to serve customers better and grow sustainably. Group teams shape strategy, policy, brand, capital allocation and transformation, helping the whole organisation perform at its best. We have a singular culture that unites all our people: we are customer-first challengers, who are committed, clear and connected.
These behaviours unite us as one team to deliver for our colleagues, our customers, our stakeholders and the country. Joining BT Group means working at the heart of a business that matters to the UK, with the opportunity to shape decisions, influence outcomes and help set the future course of one of the country’s most important companies.
📌 Software Engineering Specialist (Bengaluru)
🏢 BT Group
📍 Bengaluru