Engineer, Systems Architecture (Hyderabad)

Engineer, Systems Architecture (Hyderabad)

30 Jul
|
TMUS Global Solutions
|
Hyderabad

30 Jul

TMUS Global Solutions

Hyderabad

About the Role:

This role will be responsible for designing, developing, and maintaining automation solutions that accelerate infrastructure delivery, configuration enforcement, and reliability across on-premise AI compute and other bare-metal, virtualized and containerized platforms at both the edge and the datacenter. The role involves building scalable, self-service platforms, leveraging tools such as Ansible Automation Platform, agentic AI solutions, and open source bare-metal automation tools, enabling teams to provision, patch, and manage infrastructure on demand.

The engineer will collaborate with infrastructure, cybersecurity, and platform teams to streamline operations, automate security and compliance workflows, and drive the migration from legacy tools to modern automation frameworks. This role will champion operational excellence and developer enablement, ensuring our automation ecosystems are secure, productive, and continuously evolving to support traditional workloads, AI workloads (training, inference), and agent-driven operations.

What Youll Do:

- Assist in design, development, and deployment of end-to-end automation solutions that meets business and technical requirements.
- Drive innovation by evaluating, recommending, and adopting current technologies and tools, including AI-enabled automation capabilities.
- Enable development and operational teams through robust self-service platforms and targeted support, reducing friction while accelerating delivery.
- Design automation with a self-service-first mindset, abstracting complexity behind APIs, workflows, portals, and MCP servers that expose capabilities to AI agents, while enforcing guardrails.
- Collaborate with platform, SRE, and application teams to translate manual processes into scalable automation patterns.
- Support and evolve on-premise AI compute and GPU-as-a-Service platforms, enabling teams to provision, manage,



and consume accelerated compute through standardized, automated workflows.
- Deliver self-service capabilities for deploying and serving LLM models on GPU enabled infrastructure, with standardized, automated workflows for model rollout, scaling, and observability.
- Build and maintain Ansible Automation Platform-driven workflows for bare-metal configuration, OS patching, and rapid deployment of environments at both the edge and the datacenter.
- Use, build, and manage AI agents as a core part of day-to-day engineeringleveraging them to accelerate automation work, and helping define how the team operates, governs, and scales agent-driven workflows safely and reliably.
- Ensure automation solutions are observable, supportable, and resilient, with transparent logging, error handling, and operational documentation.

What Youll Bring:

- Bachelors degree in computer science, information systems, or related field.
- 3-6 years of hands-on experience in software development and system design.
- Proven experience delivering scalable, reliable, and secure automation solutions.
- Proficiency in at least one modern programming language (e.g., Java, Python, Go, JavaScript) applied to automation, tooling, or platform integrations.
- Familiarity with on-premise AI compute, bare-metal, and storage solutions across edge and datacenter deployments.
- Strong analytical thinking and collaborative problem-solving skills.
- Excellent communication and technical documentation abilities.

Must Have Skills:





- 5+ years technical engineering experience, preferably in multiple technology focus areas.
- Ansible Automation Platformapplied to on-premise AI workloads, bare-metal configurations, OS patching, and rapid deployment of environments at the edge and the datacenter
- Linux system administration across enterprise distributions (RHEL, Ubuntu, DGX or equivalent)
- Kubernetes (cluster operations, workload orchestration, and platform integration)
- Proven ability to design and implement automation across diverse infrastructure platforms, including bare-metal and virtual compute, GPU-accelerated high-performance computing, and enterprise storage platforms.
- Strong understanding of Infrastructure as Code principles, including modular design, version control, testing, and environment promotion.
- Experience delivering automation that supports self-service consumption, balancing developer experience, guardrails, and operational reliability.
- Demonstrated ability to troubleshoot complex infrastructure and automation issues across system, network, and platform layers.
- Hands-on experience using and managing AI agents in an engineering or operations context, including prompt design, agent orchestration, guardrails, and integrating agents into automation workflows. This is a critical capability for the role.

Nice To Have:

- Experience integrating AI/ML capabilities into automation workflows for predictive insights or intelligent orchestration.
- Experience with bare-metal automation and provisioning using Canonical MAAS, PXE-based workflows, and Redfish APIs for hardware management.
- Experience with configuring and using monitoring tools (e.g., Prometheus, Grafana).
- Experience designing and maintaining self-service automation platforms or developer enablement portals.
- Exposure to policy-as-code or automated compliance frameworks.

📌 Engineer, Systems Architecture (Hyderabad)
🏢 TMUS Global Solutions
📍 Hyderabad

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: engineer, systems architecture (hyderabad) / hyderabad

Subscribe to this job alert:

Get the latest job offers by email for: engineer, systems architecture (hyderabad) / hyderabad