
Senior Advisory Software Engineer
Pitney Bowes15 days ago
Pune, IndiaSenior
Responsibilities
- Architect and operate agentic systems for autonomous monitoring, anomaly detection, diagnosis, remediation, incident response, and escalation.
- Build platform infrastructure automation for drift detection, capacity adjustment, provisioning, patching, configuration management, and cost optimization.
- Define and govern observability strategies, SLOs, SLIs, error budgets, reliability metrics, intelligent alerting, and automation intervention thresholds.
- Own CI/CD reliability through intelligent deployment gates, automated validation, agent-triggered rollbacks, and pipeline automation.
- Design AI-augmented observability integrations across SumoLogic, CloudWatch, Grafana, Prometheus, and PagerDuty.
- Lead outage management, incident management, disaster recovery, postmortems, root-cause analysis, and validated recovery playbooks.
- Convert institutional knowledge into versioned, executable runbooks and agent-accessible knowledge bases.
- Build and maintain Git repositories, automation pipelines, agent workflows, operational tooling, and human-in-the-loop safety controls.
- Analyze infrastructure spend and implement agent-driven FinOps policy enforcement and cost governance.
- Provide senior technical leadership and communicate system health, automation outcomes, risks, and reliability metrics to engineering leadership and cross-functional stakeholders.
- Partner with development, QA, product, architecture, and other teams to embed reliability, observability, and automation requirements throughout the SDLC.
Requirements
- 10+ years of SRE or platform engineering experience, including recent automation-first and AI-augmented operations experience.
- Graduate or postgraduate education in Computer Science, Engineering, or a related discipline, or equivalent demonstrated professional experience.
- Experience operating multi-region, high-availability SaaS platforms at enterprise scale across cloud and data center environments.
- Strong experience with Git, Argo, Docker, Kubernetes, and container lifecycle management.
- Deep expertise in AWS, cloud-native architectures, event-driven automation, Lambda-based remediation, and cloud control-plane integration.
- Strong Infrastructure as Code experience with Terraform, Ansible, and CloudFormation.
- Strong Python proficiency for agentic tool integrations, LLM API wrappers, and operational automation, plus Shell scripting and PowerShell familiarity.
- Hands-on experience with Prometheus, Grafana, SumoLogic, CloudWatch, PagerDuty, and OpsGenie.
- Working knowledge of Linux, Windows, networking fundamentals, firewall concepts, and distributed systems troubleshooting.
- Familiarity with Palo Alto firewalls and regulated or GovCloud environments is advantageous.
- Proficiency with Jira, Confluence, and SharePoint and experience integrating them into automated workflows.
- Experience with agent frameworks such as LangGraph, AutoGen, or CrewAI and LLM APIs including OpenAI or Anthropic.
- Experience with prompt engineering, tool or function calling, structured outputs, agent observability, guardrails, audit trails, and human-in-the-loop escalation protocols.
- Strong analytical, communication, collaboration, accountability, and senior technical leadership skills.
- Experience working within Agile delivery models and collaborating with Product Engineering, Product Management, Client Success, and senior leadership.
Benefits
- Opportunities to grow and develop a career at Pitney Bowes.
- Inclusive workplace environment that encourages diverse perspectives and ideas.
- Challenging and unique opportunities to contribute to a transforming organization.
- Comprehensive global benefits and wellbeing programs.
- Job location is Pune.