31 Aug
|
NexGen Tech Solutions
|
Kolkata
31 Aug
NexGen Tech Solutions
Kolkata
Position: Automation & Tools Engineer (OSS Engineering)
Location: Remote (India)
Work Type: Full Time
Experience: 5-10+ Years
JOB DESCRIPTION:
The Automation & Tools Engineer will design, build and operate the Operations Support Systems
(OSS), dashboards, data products and workflow automations that enable Network Operations to support
cloud solutions and large fleets of customer-premises equipment (CPE), including broadband gateways,
routers, ONTs and connected-home devices. This is a hands-on software and systems role for an
engineer who understands how NOCs, service assurance, fault and performance management, device
operations and service-provider support work in practice.
Reporting to the Network Operations Manager, you will turn repetitive work and fragmented operational
data into secure, maintainable capabilities used by the NOC, Support and Operational Excellence teams.
You will partner with cloud, DevOps, firmware, QA, Database Engineering, Product and Security, while
remaining close to the operators who use the tools every day.
Role mandate
Own the engineering lifecycle for internal OSS tools, operational dashboards, workflow automations
and integrations assigned to the role.
Translate the team’s pain points, controls and support needs into reusable, product-like
capabilities—not disconnected one-off scripts.
Unify cloud, application, network, customer and CPE data into trusted operational views and
actionable workflows.
Automate high-volume diagnosis, enrichment, notification, remediation, reporting and evidence-
collection activities with safe controls and auditability.
Increase the capacity and consistency of Network Operations so customer and managed-device
growth does not require equivalent growth in manual effort.
What you will own
OSS platform and tool engineering
Design, develop, test, deploy and maintain internal web applications, APIs, command-line tools and
services for fault, performance, configuration, inventory and service-assurance use cases.
Build tool capabilities around service-provider tenants, cloud services, CPE inventory, topology and
dependencies, device reachability, firmware versions, alarms, incidents and customer impact.
Apply sound software engineering practices: modular design, code review, automated tests, version
control, CI/CD, release notes, rollback, telemetry and lifecycle ownership.
Assess whether to build, buy or integrate a capability; integrate commercial and open-source OSS
components without creating avoidable operational complexity.
Operational dashboards and data products
NETWORK OPERATIONS
Create role-based dashboards for executives, the NOC, Support and service owners covering service
health, incident performance, alert quality, SLA/SLO attainment, customer impact and managed-CPE
fleet health.
Correlate metrics, logs, traces, events, tickets, changes, customer/tenant data and device telemetry so
operators can move from alarm to scope,
likely cause and next action quickly.
Workflow automation and orchestration
Automate alert enrichment, deduplication, correlation, ticket creation, routing, escalation, stakeholder
notification, evidence collection and operational reporting.
Measure adoption, success rate, time saved, error reduction and operator outcomes; retire fragile
manual steps and low-value automations.
Cloud, CPE and service-provider integrations
Integrate with cloud APIs, Kubernetes, observability platforms, ITSM systems, collaboration channels,
CI/CD systems, databases, messaging platforms and customer-facing operational systems.
Use APIs, webhooks, event streams and device-management data to support onboarding,
provisioning, inventory, telemetry, command execution and firmware operations across managed CPE
fleets.
Understand and model the cloud-to-device service path sufficiently to distinguish application,
infrastructure, network, protocol, firmware, device and customer-environment signals.
Partner with Engineering on supported interfaces and data contracts; avoid unsupported direct
changes to customer-facing product systems.
Operational intelligence and continuous improvement
Convert incident reviews, ticket trends, alert noise, escalations and operator feedback into a prioritized
automation and tooling backlog managed with the Operations Manager.
Develop analytics for repeat incidents, chronic customers or device cohorts, change-related failures,
capacity trends, knowledge gaps and automation opportunities.
Prototype anomaly detection, correlation and assisted-diagnosis approaches where they provide
measurable value, while preserving explainability and human control.
Demonstrate improvements in mean time to detect, acknowledge, diagnose and recover; first-contact
resolution; escalation quality; SLA performance; and hours of toil removed.
Reliability, security and supportability
Operate internal OSS capabilities as production systems with availability targets, monitoring, alerting,
capacity planning, backup, recovery and dependency documentation.
Implement authentication, role-based access control, least privilege, secrets management, certificate
handling, audit logging, secure coding and vulnerability remediation.
Create runbooks, user guidance, support models and ownership records; train NOC and Support
users and incorporate their feedback into the roadmap.
Provide escalation support for critical tool failures and automation defects, including occasional
approved maintenance or incident work outside normal hours.
Required qualifications
5+ years of experience in OSS engineering, network automation, NOC tooling, software engineering
for operations or a closely related role.
Experience building OSS capabilities for service providers, broadband operators, managed-network
providers or telecom equipment/software vendors.
Demonstrated experience with telecom Operations Support Systems—fault, performance,
configuration, inventory, topology, service assurance or network/service orchestration—not solely
open-source software.
Robust programming skills in Python, Go, Java, TypeScript/JavaScript or a comparable language, with
experience delivering maintainable production applications, services or automation frameworks.
Experience building APIs, event-driven integrations, data pipelines, databases and user-facing
operational dashboards.
Hands-on experience with Linux, containers, Kubernetes and a public cloud platform
Experience integrating observability, ITSM, collaboration and operational platforms using REST APIs,
webhooks, message queues or event streams.
Strong understanding of software delivery practices including Git, code review, automated testing,
CI/CD, release management and rollback.
Working networking knowledge including TCP/IP, DNS, DHCP, TLS, routing, APIs and systematic
troubleshooting across distributed systems.
Ability to work directly with NOC, support and operations engineers, convert ambiguous operational
problems into clear requirements, and measure whether the delivered capability improved the
outcome.
Bachelor’s degree in computer science, engineering or equivalent practical experience.
Preferred qualifications
Experience with cloud-managed CPE such as broadband gateways, routers, ONTs, Wi-Fi/mesh
systems or similar edge devices.
Familiarity with TR-069/CWMP, TR-369/USP, TR-181 data models, ACS/USP controllers, device
telemetry and remote lifecycle management.
Experience with Google Cloud Platform, Kubernetes, Terraform, Helm and Git-based CI/CD.
Experience with Grafana, Prometheus, OpenTelemetry, Elasticsearch or comparable observability and
visualization platforms.
Experience with Apache Pulsar, Kafka or another messaging/streaming platform and relational or
NoSQL operational data stores.
Experience integrating Jira Service Management or comparable ITSM tooling and collaboration
platforms through APIs and webhooks.
Familiarity with TM Forum Open APIs, YANG, NETCONF, RESTCONF, gNMI, SNMP or other
telecom/network-management interfaces.
Experience with workflow engines, low-code orchestration or event correlation, combined with the
judgment to know when custom software is the better choice.
Experience applying RBAC, secrets management, certificate handling, audit logging and data-
protection controls to operational tooling.
Front-end development or UX experience creating efficient interfaces for operators under time
pressure.
📌 Automation & Tools Engineer (OSS) (Kolkata)
🏢 NexGen Tech Solutions
📍 Kolkata