Site Reliability Engineer
Are you ready to join an exciting early-stage start-up that detects active data breaches and protects businesses? Be part of a team that's revolutionizing data security through AI-driven solutions.
The Site Reliability Engineer owns the reliability, performance, and operability of Theom's multi-tenant SaaS platform in the cloud. You will deep-dive into production systems across Linux, networking, AWS, Kubernetes, and infrastructure-as-code, and support the data plane that connects enterprise customers' environments (including Snowflake and modern data platforms) to Theom's security and analytics services.
This is a hands-on senior IC role: you debug hard production problems, design for scale, harden automation with Terraform, and raise the operational bar for the engineering organization. You partner with Product and Engineering on architecture and incident response, and you help customer-facing teams when platform reliability or complex environment issues block adoption.
If you thrive on root-cause analysis under pressure, care about SLOs and sustainable on-call, and want to shape reliability for a fast-growing cloud data security platform, this role is for you.
What You’ll Do
- Improve the reliability, performance, and scalability of Theom’s multi-tenant SaaS platform through SLOs, observability, automation, capacity planning, and effective incident response.
- Diagnose complex Linux performance issues involving CPU, memory, I/O, scheduling, and resource contention
- Troubleshoot and design networks across TCP/IP, DNS, TLS, load balancing, NAT, routing, AWS/Azure VPC connectivity, peering.
- Design, operate, secure, and optimize AWS/Azure services, including but not limited to, EKS, EC2, VPC, IAM, S3, RDS/Aurora, load balancers, CloudWatch, Entra, Azure App Services, Azure SQL Database, and Azure Storage etc
- Operate production Kubernetes at scale, including upgrades, autoscaling, CNI, ingress, scheduling, RBAC/IRSA, resource governance, and protected deployments.
- Build and maintain advanced Terraform modules, remote state, CI/CD workflows, drift remediation, provider upgrades, and reproducible environments.
- Improve the reliability and security of data lakes like Snowflake/Databricks and data-engineering integrations, including ingestion, connectivity, governance, observability, and performance.
- Lead incident investigations, partner with Product, Engineering, and customer-facing teams on architecture and production readiness.
What We Need To See
- Degree in Computer Science or a related technical field involving coding, or equivalent experience
- 3+ years of experience in SRE, Platform Engineering, DevOps, or a similar production reliability role
- Deep Linux systems experience with proven performance and reliability troubleshooting in production
- Strong networking fundamentals: TCP/IP, DNS, HTTP/TLS, load balancers, and cloud VPC networking with hands-on debugging skill
- Strong programming or scripting skills (Python, Go, or similar)
- Advanced production experience with AWS/Azure (compute, networking, IAM, containers/orchestration, observability)
- Advanced Kubernetes experience operating production clusters (preferably EKS): networking, security, scaling, and failure diagnosis
- Advanced Terraform / CloudFormation / infrastructure-as-code: modules, state, CI/CD for infra, Gitlab
- Experience with observability stacks (prometheus, grafana, ELK, metrics, logs, traces) and turning signals into actionable SLOs and alerts
- Ability to solve complex problems in distributed systems and multi-tenant SaaS at scale
- Working knowledge of SQL, technical logs, APIs, and architecture diagrams
- Excellent communicator; clear under incident pressure; able to explain root cause to engineers and stakeholders
- Comfortable and productive in a remote / hybrid collaboration model with global time zones
Ways To Stand Out From The Crowd
- Hands-on experience with Snowflake and/or Databricks (connectivity, performance, governance, cost)
- Broader data-platform experience: data lakes/warehouses, streaming (Kafka/Kinesis), ETL pipelines
- Multi-cloud familiarity (GCP and/or Azure) in addition to deep AWS strength
- Experience with enterprise data security, governance, and compliance-minded operations
- Advanced Linux performance tooling in production
- Prior ownership of on-call rotations, incident command, and reliability roadmaps.
About Theom
Theom was founded in 2020 by an experienced founding team with extensive backgrounds at Google, Yahoo, and Cisco. These security practitioners, who had previously created Tetration Analytics, one of the first Zerotrust cloud security platforms,
united their expertise and experience to build cutting-edge solutions for protecting and governing enterprise data in the cloud.
Theom is backed by leading VC firms, enterprise data platforms, and leading cloud provider Microsoft, CISOs, CTOs, CDOs, angel investors, serial entrepreneurs, and passionate technologists who understand the complexity of securing data in the cloud. They believe in Theom's innovative approach to solving the security challenge and the team's track record in execution. Our investors:
https://www.microsoft.com/en-in
Microsoft is the world's largest vendor of computer software and a leading provider of cloud computing services, video games, computer and gaming hardware, search, and other online services.
https://www.databricks.com/blog/data-governance-and-security-ai-era-databricks-ventures-invests-theom
Databricks is the data and AI company, helping data teams solve the world’s toughest problems by providing a unified platform for data, analytics, and AI.
https://www.snowflake.com/en/why-snowflake/partners/all-partners/theom/
Snowflake Inc. is an 80 billion dollar American cloud-based data storage and analytics company.
https://www.wing.vc/
Wing is an early-stage investor behind some of the successful unicorns, such as Cohesity, Pinecone, Gong …
https://ridge.vc/
Ridge Ventures is an early-stage (seed and Series A) venture capital fund with successful investments behind Discord, F5, Fastly …
https://www.sentinelone.com/s-ventures/blog/reimagining-data-security-why-sentinelone-is-investing-in-theom-ai/
SentinelOne is a leader in AI-powered cybersecurity.
Trusted by Leading Enterprises
Since emerging from stealth, Theom has gained rapid traction with Fortune 500 companies and high-growth enterprises such as FiServ, Grammarly, TradeWeb, and JetBlue. Customers rely on Theom to ensure continuous compliance, prevent insider threats, enforce least-privilege access, and securely activate their data for AI — even as data is shared across Gen AI tools, third-party apps, and distributed teams.
Location & Work Arrangement
- Location: Hyderabad
- Work Hours: Flexible hours with some overlap with the US Pacific Time Zone for team collaboration
Ready to Join Our Global Mission?
If you're excited about building the next generation of data security solutions and want to be part of a Silicon Valley startup's journey from India, we want to hear from you. Join us in protecting the world's most sensitive data and making a global impact from India.
Website:https://www.theom.ai/https://theom.ai
📌 Site Reliability Engineer (India)
🏢 theom
📍 India