14 Sep
|
Angel One
|
Bengaluru
14 Sep
Angel One
Bengaluru
Job DescriptionAbout Angel One
NAngel One is one of India's fastest growing fin-techs, on a bold mission to make investing simple, smart, and inclusive for every Indian. With over 3+ crore clients, we're building at scale – and building for impact.
NOur Super App helps clients manage their investments, trade seamlessly, and access financial tools tailored to their goals. We are working to build personalized financial journeys for our clients, powered by new-age tech, AI, Machine Learning and Data Science.
NWe're a builder company at heart. You'll have the space to experiment, the freedom to move with velocity, and the mandate to make bold, user-first decisions – every single day.
NThe vibe? Think less hierarchy, more momentum. Everyone has a seat at the table and a shot to build something that lasts.
NBe part of a team that's scaling sustainably, thinking big, and building for the next billion.
NWhy You'll Love Working at Angel One!
N
n
- Tech Systems thatrun at Scale: From AI to real-time data infra, you'll work on tech that's ahead of the curve and solve problems that truly matter.
N
- Build one of India's Leading Fintech Platform: We're not just disrupting finance – we're shaping how billion Indians access wealth.
N
- Own It.
Drive
It.
ScaleIt: You'll have the freedom to lead, the resources to build, and the opportunity to leave your mark.
N
- Empowered Growth: We invest in your growth and empower you to explore your full potential.
N
- Exceptional Benefits: Our comprehensive benefits package includes health insurance, wellness programs, learning & development opportunities, and more.
N
nJob Title: Site Reliability Engineer 2
NLocation: Bengaluru, Karnataka
NWe are seeking an experienced Site Reliability Engineer (SRE) with deep expertise in both AWS and on-premises environments to support, scale, and optimize our hybrid cloud infrastructure. In this role, you will partner closely with engineering,
data, and platform teams to ensure the reliability, performance, and operational excellence of our containerized services, data pipelines, observability platforms, and security systems.
NWhat you will do:
NDrive reliability strategy to ensure service uptime, availability, latency targets, and overall system health across critical platforms.
NLead the definition and governance of SLIs, SLOs, and SLAs across multiple systems and services.
NLead major incident management and drive cross-team coordination for critical production issues.
NDrive organization-wide RCA processes and implement systemic reliability improvements.
NDrive strategic operational excellence programs to improve platform reliability, performance, and scalability.
NEnhance observability frameworks and implement monitoring strategies using Grafana, Prometheus, CloudWatch, and log aggregation tools across services.
NParticipate in on-call rotation;
troubleshoot incidents across the stack (network, compute, storage, data pipelines, applications).
NDrive large-scale automation initiatives to eliminate manual processes and improve platform reliability.
NDesign platform-level tools and frameworks to improve reliability and developer productivity across teams.
NParticipate in a 24×7 shift rotation, including nights, weekends, and holidays.
NDrive adoption of AI-driven operations, intelligent monitoring, and auto-healing capabilities.
NMentor engineers and drive SRE best practices across teams.
NEnforce best practices for access management,
network security, secrets management, patching, and vulnerability remediation.
NCollaborate with securityteams to ensure compliance with organizational and regulatory standards.
NWho you are:
N4–6 years of experience in an infrastructure and systems workplace delivering operational excellence to highly complex distributed systems.
NBachelor's degree in Computer Science or a related field, or equivalent work experience.
NDeep expertise in Linux, observability platforms, AWS hybrid environments, container orchestration, and automation frameworks in large-scale production environments.
NStrong expertise in Incident Management & Problem Management, leading major incident resolution and driving long-term reliability improvements.
NExtensive experience working in a 24/7 operations support environment and managing critical production systems.
NStrong experience with enterprise monitoring and observability tools such as Grafana, Prometheus, Current Relic, and Dynatrace.
NExtensive experience working with hybrid environments (AWS and on-premises infrastructure).
NAWS and CKA certifications and advanced cloud architecture knowledge are highly desirable.
NStrong experience workingwith containerization and orchestration platforms.
NExperience driving automation, infrastructure-as-code practices, and platform reliability improvements across engineering teams.
NAt Angel One, our thriving culture is rooted in Diversity, Equity, and Inclusion (DEI).
NAs an Equal prospect employer, we wholeheartedly welcome people from all backgrounds irrespective of caste, religion, gender, marital status, sexuality, disability, class or age to be part of our team. We believe that everyone's unique experiences and viewpoints make us stronger together. Come and be a part of #OneSpace*, where your individuality is celebrated and embraced.
📌 Hiring: Senior Network Operations Center Engineer (Bengaluru)
🏢 Angel One
📍 Bengaluru