Software Engineer II - SRE / DevOps / Observability (Pune)

Software Engineer II - SRE / DevOps / Observability (Pune)

20 Sep
|
Align Knowledge Centre
|
Pune

20 Sep

Align Knowledge Centre

Pune

Software Engineer II SRE / DevOps / Observability

Location: Pune

Work Model: Hybrid

Openings: 3

About the Role

We are looking for a Software Engineer II with strong SRE, observability and production engineering experience to work on large-scale, business-critical authentication and fraud decisioning platforms.

This is an engineering-focused SRE/DevOps role where you will work across CI/CD, application observability, cloud platforms, Kubernetes and production reliability.

What makes this role different?

We are specifically looking for engineers who can go beyond using existing DevOps tools.

You should be able to:

Build the pipeline deploy the application define what needs to be monitored create metrics/dashboards/alerts troubleshoot production issues improve reliability automate the solution.

Strong application-level observability experience is particularly key.

What You'll Work On

- Build and maintain end-to-end CI/CD pipelines using Jenkins, GitHub Actions and scripting.
- Design reliable deployment workflows covering build, test, deploy, validation and rollback.
- Work with deployment strategies such as rolling, blue-green and canary deployments.
- Deploy and operate applications across AWS, Kubernetes/EKS, Docker and hybrid environments.
- Define meaningful application and service-level metrics rather than relying only on infrastructure metrics.
- Build monitoring dashboards and alerts from scratch based on application/service requirements.
- Work with OpenTelemetry, Dynatrace, Grafana, Splunk, Prometheus or Datadog.
- Use metrics, logs and traces to investigate production issues and identify root causes.
- Troubleshoot unfamiliar production scenarios and determine which telemetry is relevant to the problem.
- Improve availability, performance, reliability and operational efficiency of production services.
- Develop automation using Python and Shell scripting.
- Implement Infrastructure as Code using Terraform, CloudFormation or equivalent tools.
- Work with Docker, Kubernetes and container image lifecycle management.
- Support Dev, Test, Staging and Production environments with appropriate governance and isolation.
- Participate in incident management,



root-cause analysis and reliability improvement initiatives.
- Build reusable deployment/platform capabilities that improve developer experience.
- Explore AI-driven engineering automation for areas such as incident analysis, anomaly detection and intelligent observability.

Must-Have Skills CI/CD & Automation

- Strong hands-on experience with Jenkins and/or GitHub Actions
- End-to-end deployment pipeline ownership
- Python and/or Shell scripting
- Understanding of deployment and rollback strategies

Observability Critical

- Hands-on experience with Grafana, Prometheus, Dynatrace, OpenTelemetry, Splunk or Datadog
- Experience defining application/service metrics
- Experience creating dashboards and alerts
- Understanding of metrics, logs and traces
- Ability to use telemetry for production troubleshooting

Cloud & Platform

- AWS or Azure
- Kubernetes/EKS
- Docker/containerization
- Terraform or equivalent IaC
- Linux production environments

SRE / Production Engineering

- Production incident troubleshooting
- Root-cause analysis
- Reliability and availability engineering
- Monitoring and alerting
- SLI/SLO or similar reliability practices

The Profile We're Looking For

You could be an:

SRE | Site Reliability Engineer | Platform Engineer | Production Engineer | DevOps Engineer | Software Engineer

But your background should demonstrate real engineering ownership, particularly around observability and production reliability.

Strong fit if you have experience with:

- Building CI/CD pipelines rather than only using them
- Creating Grafana/Dynatrace/Splunk/Prometheus dashboards yourself
- Defining application-level metrics
- Creating meaningful alerts
- Troubleshooting production incidents using metrics/logs/traces
- Implementing Kubernetes/EKS deployments
- Handling deployment failures and rollbacks




- Automating operational tasks using Python/Shell
- Improving MTTD/MTTR and service reliability

This role is NOT primarily for:
- Infrastructure-only administrators
- Cloud provisioning-only engineers
- Candidates who only monitor pre-built dashboards
- L1/L2 production support profiles
- Kubernetes administrators without application/CI-CD ownership
- Terraform-only infrastructure profiles
- Candidates whose experience is limited to executing existing Jenkins pipelines

Interview Focus Candidates should be prepared for practical questions around:

- Application observability and metric selection
- Dashboard and alert creation
- Metrics vs logs vs traces
- Production troubleshooting scenarios
- CI/CD architecture
- Rolling / blue-green / canary deployments
- Rollback strategies
- Kubernetes and cloud operations
- Python/Shell automation
- Basic coding/problem solving

A short CodeSignal coding exercise may also be included as part of the technical evaluation.

Experience

5+ years of relevant experience,

Technical depth and hands-on ownership are more important than simply the number of years.

Preferred Technologies

Jenkins | GitHub Actions | Python | Shell | OpenTelemetry | Dynatrace | Grafana | Splunk | Prometheus | Datadog | AWS | EKS | Kubernetes | Docker | Terraform | Linux

Why Join?

- Work on high-scale, business-critical authentication and fraud platforms.
- Solve real production reliability and observability challenges.
- Gain exposure to modern SRE, cloud-native and platform engineering practices.
- Work with Kubernetes/EKS, CI/CD, observability and automation technologies.
- Collaborate with experienced engineering and product teams in a global Agile environment.
- Build solutions that directly improve deployment reliability, monitoring and production operations.

Interested? If you are a hands-on SRE / Platform / DevOps / Software Engineer who has actually built observability, engineered CI/CD pipelines and troubleshot production systems, we'd like to hear from you.

Please apply with details of your hands-on experience in application observability, dashboard/alert creation, CI/CD and production troubleshooting.

📌 Software Engineer II - SRE / DevOps / Observability (Pune)
🏢 Align Knowledge Centre
📍 Pune

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: software engineer ii - sre / devops / observability (pune) / pune

Subscribe to this job alert:

Get the latest job offers by email for: software engineer ii - sre / devops / observability (pune) / pune