07 Aug
|
Verificient
|
Pune
Role & responsibilities
Job Title: Reliability Engineer (Full Stack + Cloud Ops)
Location: Remote / Hybrid -Pune
Company: Proctortrack
About the Role
We are looking for a Reliability Engineer who can bridge application development and cloud operations to ensure the Proctortrack platform is secure, scalable, and highly reliable.
This role is critical in delivering a consistent, high-quality experience for enterprise customers operating at scale. You will work across the stackfrom Django/MySQL application layers to cloud infrastructuredriving system reliability, performance, and operational excellence.
Key Responsibilities
Own end-to-end reliability of the Proctortrack platform across application and infrastructure layers
Monitor, troubleshoot, and resolve production issues with strong ownership and urgency
Improve system performance, uptime, and scalability of Django-based services and MySQL databases
Design and implement observability (logging, metrics, alerting) across services
Collaborate with engineering teams to identify reliability gaps and proactively drive solutions
Optimize database performance, query efficiency, and data handling at scale
Build and enhance deployment, rollback, and incident response processes
Strengthen cloud infrastructure reliability (GCP/AWS) across compute, networking, and storage
Drive automation to reduce manual effort and minimize human error
Lead post-incident analysis and translate learnings into system improvements
Preferred candidate profile
Required Skills & Experience
48 years of experience in Backend Engineering, SRE, or DevOps roles
Solid hands-on experience with Django (Python) in production environments
Solid understanding of MySQL, including performance tuning and query optimization
Experience with cloud platforms (GCP preferred, AWS acceptable)
Hands-on experience with monitoring & observability tools (e.g., Prometheus, ELK, Grafana)
Robust debugging and problem-solving skills in distributed systems
Familiarity with CI/CD pipelines and deployment strategies
Positive understanding of system design, scalability, and fault tolerance
Nice to Have
Experience with high-scale, real-time systems
Exposure to Redis, caching strategies, and asynchronous processing
Understanding of security and compliance in enterprise environments
Prior experience in video streaming or proctoring systems
📌 Remote Work Hybrid Reliability Engineer Full Stack + Cloud Ops (Pune)
🏢 Verificient
📍 Pune