22 Aug
|
TrueFan AI
|
Gurugram
22 Aug
TrueFan AI
Gurugram
Site Reliability Engineer (SRE)
Company: TrueFan AI
Location: Gurugram, Haryana, India (On-site)
Employment Type: Full time
Experience: 4-8 Years
About TrueFan AI
TrueFan AI is a Gurugram-based generative AI company building India's leading enterprise video platform. From a single five-minute shoot, our technology generates hyper-personalized, ultra-realistic video in 175 languages. We work with 100 enterprises across BFSI, healthcare, travel, and consumer internet industries where accuracy, compliance, and brand safety are non-negotiable.
About the Role
We are looking for a Site Reliability Engineer (SRE) to build, maintain, and improve highly available, scalable, secure, and reliable infrastructure and applications. The ideal candidate should have strong hands-on experience with cloud infrastructure, Kubernetes, Docker, CI/CD, observability, automation, and infrastructure as code.
Key Responsibilities
- Design, deploy, and manage scalable and highly available cloud infrastructure.
- Manage AWS services including EC2, CloudWatch, IAM, Auto Scaling Groups, Route 53, and related services.
- Deploy and manage containerized applications using Docker and Kubernetes.
- Build, maintain, and improve CI/CD pipelines using tools such as Jenkins, GitHub Actions, Devtron, or similar technologies.
- Implement and maintain monitoring, logging, and observability solutions.
- Troubleshoot production issues related to performance, availability, scalability, and reliability.
- Develop automation to reduce repetitive operational tasks and manual intervention.
- Manage infrastructure using Terraform and follow Infrastructure as Code practices.
- Implement and maintain appropriate security controls, including IAM, RBAC, secrets management, least-privilege access, and SSL/TLS.
- Implement and optimize Auto Scaling, right-sizing, Spot Instances, Reserved Instances, and other cloud cost-optimization strategies.
- Manage DNS configuration and troubleshooting using services such as Route 53 and CoreDNS.
- Participate in incident management, root-cause analysis,
and preventive action planning.
- Work closely with development and infrastructure teams to improve system reliability and deployment processes.
Required Skills
- Strong hands-on experience with AWS.
- Good understanding of EC2, CloudWatch, IAM, Auto Scaling, and Route 53.
- Strong experience with Kubernetes and Docker.
- Hands-on experience with CI/CD pipelines.
- Strong Linux administration and troubleshooting skills.
- Experience with Terraform / Infrastructure as Code.
- Experience with observability and monitoring tools such as Grafana, Kibana, New Relic, or equivalent.
- Programming/scripting experience, preferably Python, for automation.
- Understanding of networking and DNS concepts.
- Good understanding of cloud security, IAM, RBAC, secrets management, and least-privilege principles.
- Understanding of cloud cost optimization and resource management.
Preferred Qualifications
- Experience managing production-grade Kubernetes environments.
- Experience with ECS, Karpenter, Graviton, multi-architecture deployments, or similar technologies.
- Experience implementing custom Auto Scaling policies.
- Experience with production incident management and troubleshooting.
- Experience working in environments with high availability and scalability requirements.
Key Competencies
- Cloud Infrastructure & Reliability
- Kubernetes & Containerization
- CI/CD & Release Automation
- Observability & Monitoring
- Infrastructure as Code
- Cloud Security
- Automation & Scripting
- Performance & Cost Optimization
- Incident Management & Troubleshooting
Why Join Us
You will operate the infrastructure behind real-time AI products used by some of India's largest enterprises, working closely with senior backend, AI/ML, and DevOps engineers in a high-ownership team where your work is visible from day one.
Pay: ₹1,800,000.00 - ₹3,000,000.00 per year
Benefits
- Health insurance
Ability to commute/relocate:
- Gurugram, Haryana: Reliably commute or planning to relocate before starting work (Required)
Application Question(s):
- Will you be able to join us within or before 30 days?
Work Location: In person
📌 Senior Site Reliability Engineer (Gurugram)
🏢 TrueFan AI
📍 Gurugram