01 Sep
|
BT Group
|
Bengaluru
01 Sep
BT Group
Bengaluru
About the role
We are looking for a System Architect for Observability & Monitoring Platforms with 15+ years of experience who will own and drive the end-to-end architecture of a large-scale infrastructure monitoring and cloud orchestration platform. This role focuses on system architecture, technical strategy, and long-term platform evolution, enabling the platform to support metrics, logs, distributed tracing, and application performance monitoring at scale.
This role combines strategic architectural ownership with hands-on validation through design reviews, code reviews, and targeted proof-of-concepts (POCs).
What you will be doing (Role Accountabilities)
Architecture Ownership & Vision
Own the overall system architecture of the observability platform across ingestion, processing, storage, and query layers.
Work closely with Enterprise Architect to align system architecture with long-term architectural vision and technical roadmap aligned with business and platform goals.
Design and review high-level system designs, data flows, integration patterns, and core technology choices.
Act as the technical authority for complex architectural decisions, trade-offs, and design reviews.
Platform & Domain Leadership
Architect systems for infrastructure monitoring, metrics, logs, and distributed tracing at scale.
Guide the evolution from infrastructure monitoring to a full observability platform including APM.
Define architectural patterns for high-throughput telemetry ingestion, real-time processing, and query-at-scale.
Ensure architectural consistency for multi-tenant, cloud-native distributed platforms.
Standards, Governance & Enablement
Establish and evolve architecture standards, design principles, and best practices across teams.
Identify architectural risks and proactively drive mitigation strategies.
Enable teams through reference architectures, design frameworks, and technical guidance.
Support modernization initiatives including scalability, performance optimization, resilience, and cost efficiency.
Hands-On Architectural Validation
Perform code reviews of critical and performance-sensitive components.
Build and guide proof-of-concepts (POCs)
to validate architectural decisions and de-risk new technologies.
Develop reference implementations to demonstrate architectural intent.
Collaborate closely with senior engineers to troubleshoot complex system-level issues
Delivery Collaboration & Execution Enablement
Work closely with Software Engineering Managers to align architectural decisions with delivery and release plans.
Assist in breaking down large architectural initiatives into phased, incremental deliverables.
Identify architectural and technical risks early and proactively surface them to influence release planning.
Support release readiness by validating that architecture, scalability, and nonfunctional requirements are addressed ahead of key milestones.
What you’ll need to succeed (Skills & Experience)
Robust foundation in system design, distributed systems, and application scalability
Experience designing microservicebased, eventdriven architectures
Ability to make architectural tradeoffs involving scalability, reliability, performance, and cost
Proven experience designing largescale, cloudbased distributed platforms
Strong backend experience with Java (Spring Boot–based microservices)
Working knowledge of Python and/or Go for scripting, automation, and collectors
Ability to review and reason about performancecritical backend code
Strong understanding of infrastructure monitoring and observability concepts with hands on experience on follow:
o Metrics, logs, and distributed tracing
o Agentbased and agentless monitoring models
o Push / pull / subscriptionbased data collection
Handson or deep working knowledge of:
o Metric data models and storage (e.g., VictoriaMetrics or equivalent)
o Log aggregation and search platforms (Elasticsearch / OpenSearch)
Familiarity with OpenTelemetry, exporters, and custom collectors
Understanding of common monitoring tools and protocols (e.g., SNMP, Prometheusstyle systems)
Strong experience with:
o Timeseries databases for metrics
o Relational databases (PostgreSQL / MySQL) for metadata and controlplane services
o NoSQL / inmemory stores (e.g., Redis) for highvolume or lowlatency workloads
Working knowledge of search and analytics engines for logs and traces
Understanding of data modeling, query patterns, and data lifecycle management
Experience with Kafkaclass distributed messaging systems
Designing and operating eventdriven telemetry ingestion pipelines
Understanding of throughput, backpressure, and reliability concerns in streaming systems
Strong handson exposure to Docker and Kubernetesbased platforms
Infrastructure automation using:
o Terraform (IaC)
o Ansible (deployment and configuration automation)
CI/CD pipeline understanding and usage (e.g., GitLab CI or equivalent)
Familiarity with Grafana, Kibana, or similar visualization platforms
Awareness of frontend technologies such as Angular (architecturelevel understanding)
Experience with OpenStack, MAAS, Juju, or similar cloud/service orchestration platforms
Exposure to Canonical Observability Stack (COS) or equivalent platforms
Ability to perform code reviews for critical and performancesensitive components
Experience building or guiding proofofconcepts (POCs) to validate architecture
Strong collaboration with Engineering Managers and senior engineers
Clear technical communication and mentoring capability
GoodtoHave Skills
Experience with columnar / analytical databases (e.g., ClickHouse or equivalent)
Exposure to graph databases (Neo4j or similar) for relationshipheavy domains
Experience with cloudnative deployments in AWS, Azure, or GCP
Experience with secrets management (Vault, certmanager, etc.)
Understanding of secure ingestion pipelines, RBAC, and tenant isolation
Awareness of compliance and security considerations in telemetry platforms
Exposure to Telecom OSS / largescale infrastructure observability platforms
Familiarity with monitoring agents and collectors (e.g., NRPE, Filebeat, Telegraf)
📌 Software Engineering Specialist (Bengaluru)
🏢 BT Group
📍 Bengaluru