Role description
Who we are:
At UST, we help the world’s best organizations grow and succeed through transformation. Bringing together the right talent, tools, and ideas, we work with our client to co-create lasting change. Together, with over 30,000 employees in 30+ countries, we build for boundless impact—touching billions of lives in the process. Visit us at Summary:
UST is looking for a Monitoring Platform Engineer who will build, operate, and continuously improve enterprise monitoring and observability platforms that support reliable service delivery. You will design scalable telemetry pipelines for metrics, logs, and traces; improve quality; automate platform operations; and help teams detect, diagnose, and resolve service issues quickly.
The Opportunity:
Key Roles and Responsibilities:
- Engineering and operating enterprise observability platforms such as Dynatrace, Datadog, New Relic, Splunk, Elastic, Prometheus, and Grafana across cloud and on-premises environments.
- Designing and maintaining telemetry collection for servers, networks, storage, databases, containers, applications, and cloud services.
- Managing the capacity, performance, availability, security, and lifecycle upgrades of monitoring platforms.
- Establishing ing standards covering severity, thresholds, anomaly detection, suppression, correlation, deduplication, and routing.
- Reducing fatigue and creating actionable s supported by relevant dashboards, runbooks, and diagnostic context.
- Building standardized dashboards for service health, performance, capacity, availability, and SLO/SLA reporting.
- Partnering with service owners to define SLIs, SLOs, and error budgets where applicable.
- Automating monitoring deployment, configuration, and onboarding through infrastructure-as-code, configuration-management, and CI/CD practices.
- Integrating monitoring platforms with Jira, CMDB systems, Microsoft Teams, Slack, PagerDuty, Opsgenie, and other operational tools.
- Creating reusable templates, policies, s, and dashboards that enable consistent self-service adoption.
- Defining monitoring standards, telemetry requirements, tagging conventions, governance controls, documentation, and operational runbooks.
- Supporting incident response and post-incident reviews to improve detection, reduce recurrence, and strengthen monitoring coverage.
- Participating in a scheduled 24×7 on-call rotation supporting monitoring platforms and services.
- Collaborating with Security and Compliance teams to meet access-control, retention, privacy, audit, and regulatory requirements.
What you need:
- Three to seven or more years of experience in monitoring and observability, platform engineering, site reliability engineering, or IT operations.
- Hands-on experience with at least one enterprise monitoring or observability platform.
Required Skills:
- A strong understanding of metrics, logs, traces, distributed-system troubleshooting, and OpenTelemetry practices.
- Experience administering and troubleshooting Linux and Windows systems in large-scale environments.
- Working knowledge of networking concepts, including DNS, TCP/IP, load balancers, and firewalls.
- Experience with cloud-monitoring services such as Azure Monitor and Log Analytics, AWS CloudWatch, or Google Cloud Operations.
- Experience monitoring containers and Kubernetes using technologies such as Prometheus, OpenTelemetry, Grafana, Tempo, or Loki.
- Automation and scripting experience with Python, PowerShell, or Bash, along with familiarity with CI/CD pipelines.
- Experience with infrastructure-as-code or configuration-management tools such as Terraform, Bicep, CloudFormation, or Ansible.
- Familiarity with ITSM, CMDB, event-management workflows, and SRE practices such as SLOs and error budgets.
- Strong systems thinking and the ability to troubleshoot across application, platform, infrastructure, and network layers.
Desired Skills:
- Excellent judgment in separating actionable signals from noise and tuning s accordingly.
- Clear communication skills and the ability to collaborate with technical and non-technical stakeholders, especially during incidents.
- A customer-focused approach to enablement, self-service, documentation, and platform adoption.
- The ability to prioritize effectively, work in a fast-paced operational environment, and manage escalations calmly.
- Relevant monitoring-platform, cloud-platform,
or ITIL Foundation certifications would be advantageous.
Qualification:
- A bachelor’s degree in Computer Science, Information Systems, or a related discipline, or equivalent qualified experience.
What we believe:
We’re proud to embrace the same values that have shaped UST since the beginning. Since day one, we’ve been building enduring relationships and a culture of integrity. And today, it's those same values that are inspiring us to encourage innovation from everyone, to champion diversity and inclusion and to place people at the centre of everything we do.
Humility:
We will listen, learn, be empathetic and help selflessly in our interactions with everyone.
Humanity:
Through business, we will better the lives of those less fortunate than ourselves
Integrity:
We honour our commitments and act with responsibility in all our relationships.
Equal Employment Opportunity Statement
UST is an Equal Opportunity Employer. We believe that no one should be discriminated against because of their differences, such as age, disability, ethnicity, gender, gender identity and expression, religion, or sexual orientation.
All employment decisions shall be made without regard to age, race, creed, colour, religion, sex, national origin, ancestry, disability status, veteran status, sexual orientation, gender identity or expression, genetic information, marital status, citizenship status or any other basis as protected by federal, state, or local law.
UST reserves the right to periodically redefine your roles and responsibilities based on the requirements of the organization and/or your performance.
- To support and promote the values of UST.
- Comply with all Company policies and procedures
Skills
Observability, OpenTelemetry, Cloud Infrastructure, CI/CD
About UST
UST is a global digital transformation solutions provider. For more than 20 years, UST has worked side by side with the world’s best companies to make a real impact through transformation. Powered by technology, inspired by people and led by purpose, UST partners with their clients from design to operation. With deep domain expertise and a future-proof philosophy, UST embeds innovation and agility into their clients’ organizations. With over 30,000 employees in 30 countries, UST builds for boundless impact—touching billions of lives in the process.
📌 Lead I - DevOps Engineering (India)
🏢 UST
📍 India