3 days ago
Morristown, NJ, USAStaff+
Base Salary
$169k - $214k/yr
Responsibilities
- Lead AIOps strategy and architecture during AWS migration and AI-native operations expansion.
- Define intelligent operational patterns for incident detection, triage, remediation, root cause analysis, anomaly detection, alert correlation, and decision-making.
- Architect and scale observability across cloud and application environments, including metrics, logs, traces, dashboards, alerting, and service health visibility.
- Drive AWS operational excellence through scalable monitoring, resilience, reliability, automation, and governance patterns.
- Design and implement AI-agent, agentic, and multi-agent operational workflows for troubleshooting, remediation, operator assistance, and workflow automation.
- Build and mature ChatOps capabilities for collaboration, incident response, visibility, and operational execution.
- Partner with SRE, Cloud Engineering, Infrastructure, Security, and Application teams to embed reliability practices into operations and platform design.
- Establish standards for telemetry, service-level objectives, alert quality, escalation workflows, incident readiness, and post-incident learning.
- Create scalable playbooks, standards, reusable patterns, and operating models for AIOps adoption.
- Mentor engineers and operators in observability, automation, SRE, ChatOps, and AI-assisted operations.
Requirements
- Typically requires a BS plus 12 years of experience or an MS plus 10 years of experience, or equivalent.
- Strong hands-on AWS experience across compute, networking, storage, security, automation, and cloud operations.
- Deep experience with observability and monitoring platforms such as New Relic and OpenSearch.
- Proven experience applying AI to event correlation, anomaly detection, alert reduction, root cause analysis, remediation support, and operational workflow automation.
- Strong experience designing or implementing AI agents, agentic workflows, or multi-agent systems.
- Strong grounding in SRE principles, including SLOs, SLIs, error budgets, automation, incident management, resilience, and continuous improvement.
- Demonstrated success building or scaling ChatOps practices.
- Strong knowledge of scripting, infrastructure automation, operational tooling, APIs, event-driven systems, and platform integration patterns.
- Ability to translate operational problems into scalable technical solutions and influence technical teams and senior leaders.
- Experience implementing operational AI with controls for accuracy, security, compliance, explainability, and human oversight.
- Preferred qualifications include leading enterprise AIOps initiatives, supporting AWS migrations, integrating AI agents with operational systems, working in healthcare or regulated environments, and exposure to platform engineering, DevOps, and developer experience practices.
Benefits
- Flexible hybrid work arrangement with in-office presence guided by team and business needs; candidates within approximately 50 miles of a U.S. office are generally preferred.
- U.S.-based benefits include a 401(k) plan with employer match, flexible paid time off, holidays, parental leave, life and disability insurance, and medical, dental, vision, and prescription drug coverage.
- May include discretionary bonus plans, commissions, or other incentives depending on the role.
