02 Oct
|
TriluxTech
|
India
Title - Senior DevOps Engineer / IT Operations
Shift : Night shift: 12:30 AM to 8:30 AM IST
Type - 6 months contract to hire
Location- Remote
Key Responsibilities
Pillar 1: Platform Monitoring & Reliability (SRE / DevOps)
- Own production health across our AWS environment: ECS Fargate, Lambda, Aurora, RDS, and ElastiCache. You decide what gets watched and what is worth waking someone for.
- Own the alerting design and dashboards in CloudWatch, Grafana, OpenSearch, and our on-call platform: severity, thresholds, grouping, and routing, with noise fixed at the source, never muted. An alert nobody acts on is yours to fix, not to forward.
- Run platform incidents in your window: declare severity, direct the investigation, escalate when the call belongs to a service owner or the US on-call, and write the postmortem with named owners. You drive the incident; you don't just document it.
- Own release verification and rollback: error rates, latency, and task health after every deploy.
- Automate toil in Python or Bash, and build the checks that catch a missed batch window, an expiring certificate, or a failed backup.
- Own the Terraform behind your monitoring, alarms, and IAM, reviewed like application code.
Pillar 2: IT Operations & Employee Support
- Run employee IT support against published response and resolution SLAs, routing only what needs a vendor or the US team.
- Own the identity lifecycle in Okta, Google Workspace, and Entra ID: joiner, mover, and leaver flows, access reviews you run and data owners approve, and provisioning automated through their APIs, including Microsoft Graph.
- Own the Windows and macOS endpoint baseline in Intune and Kandji,
plus the engineering Linux fleet: enrollment, patch and compliance policy, and asset and license accuracy with disposal evidence.
- Escalate suspected phishing, malware, and account compromise to Information Security immediately, per the incident response plan. Speed matters more than certainty here.
- Handle consumer and partner data only in approved systems, and escalate any suspected exposure immediately. This role does not make or automate credit, lending, or eligibility decisions.
Shared Foundation
Documentation & Shift Handoff
- Write handoff notes covering open alerts, in-flight work, and any call the next region needs.
- Keep runbooks and SOPs current and executable by a responder who hasn't seen the system.
Ticket & Change Discipline
- Record work in Jira in enough detail to audit months later. Ticket and change records are SOC 2 evidence.
- Follow change management for every production-affecting action. Where no runbook exists, act on the evidence, then write one; repeat work is a defect, not a workload.
Required Skills and Qualifications
- 5 to 7 years in DevOps, SRE, systems administration, or IT operations, with production ownership. Depth on one pillar and hands-on experience on the other.
- Production AWS depth: ECS or EKS, Lambda,
Aurora or RDS, ElastiCache, S3, IAM, VPC, CloudWatch, Docker, and Linux fundamentals across systemd, networking, TLS, and log analysis.
- Monitoring you built: alerts, dashboards, and severity models in CloudWatch, Grafana, Prometheus, Datadog, or OpenSearch, plus on-call platform administration.
- Terraform modules you maintained and plans you read closely enough to catch a destructive one, CI/CD pipelines you built and debugged in Buildkite, GitHub Actions, or Jenkins, and Python or Bash automation.
- Identity administration in Okta, Google Workspace, or Entra ID, plus endpoint management in Intune, Kandji, or a comparable MDM: SSO, MFA, lifecycle automation, and access reviews.
- Incident response as the responder in an ITSM practice you helped shape: you made the severity call and closed the postmortem actions.
- Experience under SOC 2, ISO 27001, or PCI: control evidence, access reviews, and backup verification.
- Excellent written English. Bachelor's degree in Computer Science or IT, or a self-taught equivalent.
Preferred Experience
- ECS Fargate depth (task sizing, capacity providers, autoscaling, circuit breakers) and PagerDuty administration (escalation policies, event rules).
- Running all three identity platforms together, with lifecycle automation in Okta Workflows or the equivalent.
- Fintech, banking, or another regulated industry, or standing up an operational practice in a current region.
- Certifications such as AWS Certified DevOps Engineer Skilled, Certified Kubernetes Administrator, or HashiCorp Terraform Associate.
📌 Senior DevOps Engineer / IT Operations (India)
🏢 TriluxTech
📍 India