09 Oct
|
TMUS Global Solutions
|
Hyderabad
09 Oct
TMUS Global Solutions
Hyderabad
About the Role:
This role is a hands-on technical leadership position that owns the architecture, automation, and reliability of a core infrastructure domain across on-premise compute and other bare-metal, virtualized and containerized platforms at both the edge and the datacenter. The Principal Engineer is the accountable technical owner for that space: setting direction, defining standards, making and defending design decisions, and answering for outcomes such as availability, delivery speed, and operational quality. The role involves building scalable, self-service platforms, leveraging tools such as Ansible Automation Platform, agentic AI solutions, and open source bare-metal automation tools, enabling teams to provision, patch, and manage infrastructure on demand. The Principal Engineer is a mentor and technical anchor for the team, raising the capability of other engineers through design reviews, hands-on guidance, incident leadership, and shared standards. The engineer will collaborate with infrastructure, cybersecurity, and platform teams, including counterparts across time zones, to streamline operations, automate security and compliance workflows, and drive the migration from legacy tools to modern automation frameworks. This role will champion operational excellence and developer enablement, ensuring our automation ecosystems are secure, efficient, and continuously evolving to support traditional workloads, AI workloads (training, inference), and agent-driven operations.
What Youll Do:
- Own the technical strategy, roadmap, and health of a core infrastructure domain, and be the accountable technical lead for its design, reliability, and lifecycle.
- Lead the design, development, and deployment of end-to-end automation solutions that meet business and technical requirements.
- Mentor engineers and senior engineers through design reviews, code and automation reviews, pairing, and onboarding, and actively develop their technical depth and independence.
- Provide technical leadership during critical P1/P2 production incidents, drive root cause analysis, and ensure corrective actions are completed and shared across the team.
- Define infrastructure standards, reference architectures, and operational documentation, and evaluate designs for availability, resiliency, scalability, security, and supportability.
- Drive innovation by evaluating, recommending, and adopting new technologies and tools, including AI-enabled automation capabilities.
- Enable development and operational teams through robust self-service platforms and targeted support, reducing friction while accelerating delivery.
- Design automation with a self-service-first mindset, abstracting complexity behind APIs, workflows, portals, and MCP servers that expose capabilities to AI agents, while enforcing guardrails.
- Collaborate with platform, SRE, cybersecurity, and application teams, including US-based counterparts, to translate manual processes into scalable automation patterns.
- Support and evolve on-premise AI compute and GPU-as-a-Service platforms, enabling teams to provision, manage, and consume accelerated compute through standardized, automated workflows.
- Deliver self-service capabilities for deploying and serving LLM models on GPU enabled infrastructure, with standardized, automated workflows for model rollout, scaling, and observability.
- Build and maintain Ansible Automation Platform-driven workflows for bare-metal configuration, OS patching, and rapid deployment of environments at both the edge and the datacenter.
- Use, build, and manage AI agents as a core part of day-to-day engineering, leveraging them to accelerate automation work and helping define how the team operates, governs, and scales agent-driven workflows safely and reliably.
- Ensure automation solutions are observable, supportable, and resilient, with clear logging, error handling, and operational documentation.
What Youll Bring:
- Bachelors degree in computer science, information systems, or related field.
- Proven experience owning an infrastructure platform or domain end to end and delivering scalable, reliable, and secure automation solutions.
- Demonstrated track record of mentoring and developing engineers, such as leading design reviews, coaching through incidents, or growing team members into larger roles.
- Proficiency in at least one modern programming language (e.g., Java, Python, Go, JavaScript) applied to automation, tooling, or platform integrations.
- Familiarity with on-premise AI compute, bare-metal, and storage solutions across edge and datacenter deployments.
- Strong analytical thinking and collaborative problem-solving skills, with the judgment to make and defend architecture decisions.
- Excellent communication and technical documentation abilities.
Must Have Skills:
- 10+ years technical engineering experience, preferably in multiple technology focus areas.
- Deep expertise in at least one infrastructure domain (for example Linux platforms, virtualization, OpenShift/Kubernetes, or enterprise storage), with broad working knowledge across the others.
- Demonstrated ability to mentor and develop engineers and to act as the technical authority and escalation point for a team.
- Ansible Automation Platform, applied to on-premise AI workloads, bare-metal configurations, OS patching, and rapid deployment of environments at the edge and the datacenter.
- Linux system administration across enterprise distributions (RHEL, Ubuntu, DGX or equivalent).
- Kubernetes (cluster operations, workload orchestration, and platform integration), including OpenShift and OpenShift Virtualization.
- Working knowledge of enterprise server hardware, virtualization (VMware, OpenShift Virtualization, or similar), SAN/NAS storage, and TCP/IP networking.
- Proven ability to design and implement automation across diverse infrastructure platforms, including bare-metal and virtual compute, GPU-accelerated high-performance computing, and enterprise storage platforms.
- Strong understanding of Infrastructure as Code principles, including modular design, version control, testing, and environment promotion.
- Experience delivering automation that supports self-service consumption, balancing developer experience, guardrails, and operational reliability.
- Demonstrated ability to troubleshoot complex infrastructure and automation issues across hardware, system, network, and platform layers, and to lead technical response during major incidents.
- Hands-on experience using and managing AI agents in an engineering or operations context, including prompt design, agent orchestration, guardrails, and integrating agents into automation workflows. This is a critical capability for the role
Nice To Have:
- Experience integrating AI/ML capabilities into automation workflows for predictive insights or intelligent orchestration.
- Experience with bare-metal automation and provisioning using Canonical MAAS, PXE-based workflows, and Redfish APIs for hardware management.
- Experience with configuring and using monitoring tools (e.g., Prometheus, Grafana, Dynatrace, Datadog, Current Relic).
- Experience designing and maintaining self-service automation platforms or developer enablement portals.
- Exposure to policy-as-code or automated compliance frameworks.
- Experience with infrastructure security hardening, vulnerability remediation, and patch management at scale.
- Experience with large-scale hardware refreshes, platform migrations, or datacenter transformation
📌 Principal Engineer, Systems Architecture (Hyderabad)
🏢 TMUS Global Solutions
📍 Hyderabad