29 Aug
|
Pitney Bowes (pbi)
|
Pune
29 Aug
Pitney Bowes (pbi)
Pune
Job Summary
As a Senior Advisory Software Engineer, you will operate at the intersection of platform engineering and AI-driven automation. This is not a traditional SRE role. You will design, build, and supervise agentic systems that detect anomalies, diagnose failures, execute remediation runbooks, and escalate intelligently-with minimal human intervention. You will architect the feedback loops that make our platform progressively self-healing. You will collaborate across engineering, product, and architecture to ensure our observability and incident response capabilities stay ahead of system complexity. Being a Senior Advisory SRE here means you think in systems, build in agents, and measure success in mean-time-to-no-action.
Join Pitney Bowes as Senior Advisory Software Engineer
Years of experience: 9-12 years
Job Location: Pune
Responsibilities
- Architect and operate agentic systems for autonomous monitoring, anomaly detection, and self-healing - reducing mean-time-to-remediation without human-in-the-loop dependency for routine failure patterns.
- Build software and agentic pipelines that manage platform infrastructure autonomously - from drift detection and remediation to capacity adjustment and incident triage.
- Drive reliability engineering outcomes - SLO attainment, error budget governance, and deployment velocity - by embedding intelligence into the platform rather than adding human process overhead.
- Measure and continuously optimize system performance using agent-driven telemetry analysis - identifying degradation patterns before they manifest as customer-impacting incidents.
- Own CI/CD reliability across the SDLC - integrating agentic checks, automated rollback triggers, and intelligent deployment gates that act on signal, not on schedule.
- Scope includes:
- Agentic observability - context-aware monitoring with LLM-assisted signal interpretation
- Intelligent alert design - dynamic thresholds, noise suppression, and automated triage routing
- Runbook automation and agentic remediation - codifying institutional knowledge into executable, supervised agent workflows
- Autonomous incident response - agent-led detection, diagnosis, and escalation with human override at defined severity thresholds
- Infrastructure lifecycle management - provisioning, drift remediation, and cost optimization driven by policy-as-code and agent execution
- End-to-end configuration, deployment, and patching - governed by automated validation pipelines, not manual checklists
- Creating and maintaining GIT repo and pipelines
- Communicate risks,
system health, and automation outcomes clearly to engineering leadership and cross-functional stakeholders - translating agent behaviour and reliability signals into business-readable insight.
- Define and continuously refine the observability strategy - what to monitor, how to act on it, and how to suppress noise programmatically. Drive adoption of agent-assisted monitoring across product and infrastructure layers.
- Analyse operational behaviour patterns across user personas and platform workloads to inform intelligent automation design and monitoring strategy evolution.
- Define and govern Service Level Indicators and Objectives - using error budget data to drive engineering prioritisation and calibrate automation intervention thresholds rather than to reactively defend the committed SLA.
- Define, instrument, and own platform reliability metrics - QoS, Uptime, MTTR, MTBF, and agent automation coverage - as leading indicators of system health and team maturity.
- Synthesise and publish key metrics to stakeholders
- Leverage deep AWS expertise and DevOps toolchain knowledge to design infrastructure automation that operates with the reliability and predictability of a software system.
- Continuously analyse infrastructure and tooling spend - identifying waste, right-sizing opportunities, and cost anomalies through automated FinOps signal processing.
- Build and operationalise cost governance frameworks where agent-driven policy enforcement, not periodic human review, is the primary control mechanism.
- Design and implement AI-augmented observability solutions across the stack - integrating SumoLogic, CloudWatch, Grafana, Prometheus, and PagerDuty with agentic reasoning layers that interpret signal and act, not just alert.
- Lead outage management with agentic support - automated problem detection, structured stakeholder communication, and agent-assisted resolution with clear human escalation protocols for high-severity events.
- Own incident management and disaster recovery strategy - with automated runbook execution, agent-supervised recovery playbooks, and validated DR testing cadences.
- Convert institutional knowledge into machine-executable runbooks and agent-accessible knowledge bases - ensuring operational intelligence is codified, versioned, and continuously improved
- Lead operational improvement through continuous automation - replacing recurring manual processes with agent-driven workflows and measuring success by the reduction in human intervention per unit of platform activity.
- Conduct rigorous incident postmortems and RCAs - using AI-assisted log correlation and timeline reconstruction to surface systemic gaps faster and drive durable corrective action.
- Partner with development, QA, and architecture teams to embed reliability and automation requirements early in the SDLC - shifting reliability left rather than absorbing complexity at the production boundary.
- Contribute to system design reviews with a reliability and automation lens - ensuring that agentic operability, observability hooks, and failure mode handling are first-class design considerations, not afterthoughts.
- Provide senior technical leadership - setting the bar for agentic automation design, code quality in SRE tooling, and engineering rigour across the team.
- Demonstrate Ownership and accountability.
Qualifications & Skills required
This is a critical service delivery role requiring experience with complex datacenter and cloud hosting environments. Pitney Bowes product hosting solutions leverage multiple technologies in complex data center and cloud environments that support multi-tiered high-availability applications. The role requires a talented self-directed and self-motivated individual with a strong work ethic and the following skills:
- Graduate or Post-Graduate in Computer Science, Engineering, or a related discipline - or equivalent demonstrated depth through qualified experience.
- 10+ years of SRE or platform engineering experience, with a demonstrable shift in recent years toward automation-first and AI-augmented operations.
- Excellent written and verbal communication skills - able to translate agent behaviour, reliability signals, and automation outcomes into clear engineering and executive narratives.
- Strong background in SaaS product operations - with hands-on experience running multi-region, high-availability platforms at enterprise scale.
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Sr Advisory Software Engineer (Pune)
🏢 Pitney Bowes (pbi)
📍 Pune