10 Aug
|
theom
|
Hyderabad
Site Reliability Engineer
Are you ready to join an exciting early-stage start-up that detects active data breaches and protects businesses? Be part of a team that's revolutionizing data security through AI-driven solutions.
The Site Reliability Engineer owns the reliability, performance, and operability of Theom's multi-tenant SaaS platform in the cloud. You will deep-dive into production systems across Linux, networking, AWS, Kubernetes, and infrastructure-as-code, and support the data plane that connects enterprise customers' environments (including Snowflake and modern data platforms) to Theom's security and analytics services.
This is a hands-on senior IC role: you debug hard production problems, design for scale, harden automation with Terraform, and raise the operational bar for the engineering organization. You partner with Product and Engineering on architecture and incident response, and you help customer-facing teams when platform reliability or complex environment issues block adoption.
If you thrive on root-cause analysis under pressure, care about SLOs and sustainable on-call, and want to shape reliability for a fast-growing cloud data security platform, this role is for you.
What Youll Do
- Improve the reliability, performance, and scalability of Theom’s multi-tenant SaaS platform through SLOs, observability, automation, capacity planning, and effective incident response.
- Diagnose complex Linux performance issues involving CPU, memory, I/O, scheduling, and resource contention
- Troubleshoot and design networks across TCP/IP, DNS, TLS, load balancing, NAT, routing, AWS/Azure VPC connectivity, peering.
- Design, operate, secure, and optimize AWS/Azure services, including but not limited to, EKS, EC2, VPC, IAM, S3, RDS/Aurora, load balancers, CloudWatch, Entra, Azure App Services, Azure SQL Database, and Azure Storage etc
- Operate production Kubernetes at scale, including upgrades, autoscaling, CNI, ingress, scheduling, RBAC/IRSA, resource governance, and safe deployments.
- Build and maintain advanced Terraform modules,
remote state, CI/CD workflows, drift remediation, provider upgrades, and reproducible environments.
- Improve the reliability and security of data lakes like Snowflake/Databricks and data-engineering integrations, including ingestion, connectivity, governance, observability, and performance.
- Lead incident investigations, partner with Product, Engineering, and customer-facing teams on architecture and production readiness.
What We Need To See
- Degree in Computer Science or a related technical field involving coding, or equivalent experience
- 3+ years of experience in SRE, Platform Engineering, DevOps, or a similar production reliability role
- Deep Linux systems experience with proven performance and reliability troubleshooting in production
- Strong networking fundamentals: TCP/IP, DNS, HTTP/TLS, load balancers, and cloud VPC networking with hands-on debugging skill
- Strong programming or scripting skills (Python, Go, or similar)
- Advanced production experience with AWS/Azure (compute, networking, IAM, containers/orchestration, observability)
- Advanced Kubernetes experience operating production clusters (preferably EKS): networking, security, scaling, and failure diagnosis
- Advanced Terraform / CloudFormation / infrastructure-as-code: modules, state, CI/CD for infra, Gitlab
- Experience with observability stacks (prometheus, grafana, ELK, metrics, logs, traces) and turning signals into actionable SLOs and alerts
- Ability to solve complex problems in distributed systems and multi-tenant SaaS at scale
- Working knowledge of SQL, technical logs, APIs, and architecture diagrams
- Excellent communicator; clear under incident pressure; able to explain root cause to engineers and stakeholders
- Comfortable and productive in a remote / hybrid collaboration model with global time zones
Ways To Stand Out From The Crowd
- Hands-on experience with Snowflake and/or Databricks (connectivity, performance, governance, cost)
- Broader data-platform experience: data lakes/warehouses, streaming (Kafka/Kinesis), ETL pipelines
- Multi-cloud familiarity (GCP and/or Azure) in addition to deep AWS strength
- Experience with enterprise data security, governance, and compliance-minded operations
- Advanced Linux performance tooling in production
- Prior ownership of on-call rotations, incident command, and reliability roadmaps.
About Theom Theom was founded in 2020 by an experienced founding team with extensive backgrounds at Google, Yahoo, and Cisco. These security practitioners, who had previously created Tetration Analytics, one of the first Zerotrust cloud security platforms, united their expertise and experience to build cutting-edge solutions for protecting and governing enterprise data in the cloud.
Trusted by Leading Enterprises
Since emerging from stealth, Theom has gained rapid traction with Fortune 500 companies and high-growth enterprises such as FiServ, Grammarly, TradeWeb, and JetBlue. Customers rely on Theom to ensure continuous compliance, prevent insider threats, enforce least-privilege access, and securely activate their data for AI — even as data is shared across Gen AI tools, third-party apps, and distributed teams.
Location & Work Arrangement
- Location: Hyderabad
- Work Hours: Flexible schedule with some overlap with the US Pacific Time Zone for team collaboration
Ready to Join Our Global Mission? If you're excited about building the next generation of data security solutions and want to be part of a Silicon Valley startup's journey from India, we want to hear from you. Join us in protecting the world's most sensitive data and making a global impact from India.
📌 Site Reliability Engineer (Hyderabad)
🏢 theom
📍 Hyderabad