04 Aug
|
HCL Comnet
|
Noida
- : Development Lead – Network & Infrastructure Observability
- Location: [Noida/Bangalore] |
- Experience: 10 years (5 years in a technical leadership capacity)
- Send resumes to: [Confidential Information] with below details:
- Name:
- Exp:
- CTC:
- ECTC:
- Notice period:
- Current location:
About the Role
We are looking for a Development Lead to own the engineering of a core module within our observability platform, focused on discovery, monitoring, and performance management of networking devices and infrastructure (routers, switches, firewalls, load balancers, wireless controllers, and related network/IoT endpoints). This is a hands-on technical leadership role: you will design and build high-throughput data collection and processing pipelines, lead a team of engineers, and set the technical direction for how the platform monitors, correlates, and alerts on network health at scale.
The ideal candidate has deep experience building monitoring/NMS (Network Management System) style products or modules — comparable to the engineering work behind commercial network and infrastructure monitoring platforms — and is equally comfortable writing production Python code and guiding architectural decisions.
Key Responsibilities
- Lead design, development, and delivery of network device discovery, polling, and monitoring capabilities within the observability platform.
- Design and develop scalable data collection pipelines (polling, streaming, and event-driven) capable of handling large device inventories and high-frequency metric ingestion.
- Drive engineering decisions around polling engines, protocol handlers, MIB/OID management, topology mapping, and alerting/thresholding logic.
- Lead, mentor, and grow a team of Python developers; conduct code reviews, design reviews, and enforce engineering best practices.
- Ensure the platform's collectors and agents are resilient, secure, and performant across heterogeneous vendor devices and firmware versions.
- Collaborate with QA, DevOps, and Support teams to ensure reliability, observability of the observability stack itself, and fast root-cause turnaround for field issues.
- Evaluate and integrate new protocols, data sources,
and third-party MIBs/device packs as network technology evolves (SD-WAN, wireless, cloud-networking, IoT/OT).
- Participate in sprint planning, estimation, and stakeholder communication; report on progress, risks, and technical trade-offs to senior leadership.
- Maintain a strong security posture across credential handling, encrypted transport, and role-based access for device communication.
Domain & Protocol Scope
Candidates should have hands-on experience developing against or integrating with the following network management and telemetry protocols/standards:
- SNMP (v1/v2c/v3): GET/GETNEXT/GETBULK polling, trap/inform handling, MIB parsing and compilation, custom OID mapping.
- Flow-based telemetry: NetFlow, sFlow, IPFIX for traffic analytics and bandwidth monitoring.
- Syslog: RFC 3164/5424 ingestion, parsing, and correlation for event and log-based monitoring.
- ICMP: Ping/traceroute-based reachability and latency monitoring.
- CLI-based collection: SSH/Telnet automation and screen-scraping for devices lacking API/SNMP support.
- NETCONF/YANG and RESTCONF: Structured configuration and state retrieval from modern network devices.
- WMI/WinRM: Windows-based server and infrastructure monitoring.
- Topology discovery protocols: LLDP, CDP for physical/logical network mapping.
- REST/SOAP APIs and gRPC: Vendor and cloud-provider integrations (firewalls, wireless controllers, SD-WAN controllers, public cloud networking).
- Streaming telemetry & modern observability standards: OpenTelemetry, Prometheus exposition format/remote-write, StatsD, gNMI (desirable).
- IoT/OT telemetry (desirable): MQTT, TR-069/CWMP for CPE/customer premise device management.
- Time-series and event pipelines: Kafka, RabbitMQ, or similar for high-volume metric/event streaming.
Familiarity with routing/switching fundamentals (BGP, OSPF, VLANs, STP) and general network engineering concepts is expected,
even though the role is development-focused rather than network operations.
Required Skills & Experience
- 10 years of professional software development experience, with at least 5 years in a technical or engineering lead capacity.
- Solid, production-grade Python development skills: asyncio/concurrency patterns, performance optimization, packaging, and testing (pytest or equivalent).
- Demonstrated experience building or extending network monitoring, NMS, APM, or observability products/modules.
- Solid understanding of SNMP and at least two additional protocols listed above, with real implementation experience (not just conceptual knowledge).
- Experience with time-series ,columnar and graph databases (e.g., Prometheus, InfluxDB, OpenTSDB, Clickhouse, Neo4J) and metric storage/query patterns at scale.
- Practical experience with containerization and orchestration (Docker, Kubernetes) and CI/CD pipelines.
- Strong Linux systems knowledge; comfortable with networking fundamentals (TCP/IP stack, DNS, routing).
- Experience with relational and NoSQL data stores, and designing schemas for large-scale telemetry data.
- Track record of leading small-to-mid-sized engineering teams (4-10 engineers), including hiring, mentoring, and performance management.
- Strong written and verbal communication skills; ability to translate between deep technical detail and business/product priorities.
Preferred / Nice-to-Have
- Experience with additional languages relevant to agent/collector development (Go, Java, or C/C ).
- Exposure to public cloud networking services (AWS VPC/Transit Gateway, Azure VNet, GCP networking) and their native monitoring APIs.
- Familiarity with anomaly detection, baselining, or ML-assisted alerting techniques applied to network/infrastructure metrics.
- Experience with ITSM/ITOM integrations (ticketing, CMDB sync, event correlation with tools like ServiceNow).
- Prior work on agent-based and agentless monitoring architectures, and plugin/extensibility frameworks for third-party device support.
- Security certifications or hands-on experience with secure credential vaulting, TLS/mTLS, and role-based access control for device communication.
📌 Development Lead Network & Infrastructure Observability (Noida)
🏢 HCL Comnet
📍 Noida