Principal Engineer, Site Reliability (Hyderabad)

Principal Engineer, Site Reliability (Hyderabad)

19 Sep
|
TMUS Global Solutions
|
Hyderabad

19 Sep

TMUS Global Solutions

Hyderabad

About the Role

nThe Principal Engineer, Site Reliability (SRE) will play a critical role in ensuring the stability, scalability, and operational excellence of Accounting and Finance platforms. This role is focused on leading the operational health of these platforms, ensuring the delivery of highly reliable financial applications and data services that meet the demanding requirements of accuracy, compliance, and availability to support business operations.

nAs a Principal SRE, you will build automation, implement monitoring, improve incident response, and champion DevOps practices that enable Finance and Accounting systems to operate with consistency and trustworthiness, while also coaching and mentoring junior SREs to ensure overall operational excellence.

nWhat Youll Do

nOperational Oversight: Own day-to-day operations for Accounting and Finance applications and data platforms, ensuring they run smoothly and meet business expectations.

nReliability & Availability: Ensure Accounting and Finance platforms meet defined SLAs, SLOs, and SLIs for performance, reliability, and uptime.

nAutomation & Efficiency: Build automation for deployments, monitoring, scaling, and self-healing capabilities to reduce manual effort and operational risk.

nObservability & Monitoring: Implement and maintain comprehensive monitoring, alerting, and logging for accounting applications and data pipelines (e.g., Snowflake, dbt workflows, ERP integrations).

nIncident Response: Lead and participate in on-call rotations, perform root cause analysis, and drive improvements to prevent recurrence of production issues.

nOperational Excellence: Establish and enforce best practices for capacity planning, performance tuning, disaster recovery, and compliance controls in financial systems.

nCollaboration with Engineering & Finance: Partner with software engineers, data engineers, and Finance/Accounting teams to ensure operational needs are met from development through production.

nTeam Coordination: Manage workload, priorities,



and escalations for operations staff and partner teams, ensuring alignment with SLAs and compliance requirements.

nSecurity & Compliance: Ensure financial applications and data pipelines meet audit, compliance, and security requirements.

nContinuous Improvement: Drive post-incident reviews, implement lessons learned, and proactively identify opportunities to improve system resilience.

nAudit & Compliance Support: Ensure operational practices meet internal controls, audit requirements, and financial compliance standards.

nWhat Youll Bring

nBachelors in Computer Science, Engineering, Information Technology, or related field (or equivalent experience).

n7-12 years of experience in Site Reliability Engineering, DevOps, or Production Engineering, ideally supporting financial or mission-critical applications.

nStrong experience with monitoring/observability tools (Datadog, Prometheus, Grafana, Splunk, or equivalent).

nHands-on expertise with CI/CD pipelines, automation frameworks, and IaC tools (Terraform, Ansible, GitHub Actions, Azure DevOps, etc.).

nFamiliarity with Snowflake, dbt, and financial system integrations from an operational support perspective.

nStrong scripting/programming experience (Python, Bash, Go, or similar) for automation and tooling.

nProven ability to manage incident response and conduct blameless postmortems.

nExperience ensuring compliance, security, and audit-readiness in enterprise applications.

nMust Have Skills

nSRE

nSQL

nSnowflake OR Databricks

nDevOps OR CICD OR Github Actions

nmonitoring/observability tools (Datadog, Prometheus, Grafana, Splunk, or equivalent)

nAutomation

nNice To Have

nExperience supporting financial applications (ERP, revenue recognition systems, accounting platforms).

nExposure to FinOps practices for optimizing cloud spend in finance-related platforms.

nFamiliarity with containers and orchestration (Docker, Kubernetes).

nExperience building resilience into data pipelines and ensuring auditability for accounting data.

nStrong communication skills to articulate operational issues and risks to both technical and non-technical stakeholders.

📌 Principal Engineer, Site Reliability (Hyderabad)
🏢 TMUS Global Solutions
📍 Hyderabad

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: principal engineer, site reliability (hyderabad) / hyderabad

Subscribe to this job alert:

Get the latest job offers by email for: principal engineer, site reliability (hyderabad) / hyderabad