2 months ago
Chicago, IL, USAMid Level / Senior
Base Salary
$103k - $159k/yr
Responsibilities
- Develop and maintain deep technical knowledge of iManage platform services while troubleshooting logs, queries, infrastructure, and distributed cloud services.
- Use telemetry, anomaly detection, and trend analysis to identify systemic issues and partner with technical stakeholders on proactive fixes.
- Serve as incident commander for P1/P2 incidents, coordinating communication and engineering response to restore services.
- Build and tune dashboards, alerts, and synthetic monitoring for real-time visibility into system and end-user health.
- Serve as the Platform Support escalation point and subject matter expert for service consumption and observability.
- Partner with Engineering and SRE teams on internal tooling, observability, supportability, and complex troubleshooting.
- Lead customer advisories, technical escalations, problem isolation, data gathering, and resolution validation.
- Own end-to-end root cause analyses and blameless postmortems while driving systemic fixes that prevent recurrence.
- Advocate for users and stakeholders by identifying product friction and reliability concerns.
- Drive automation, self-service, knowledge management improvements, and AI-powered operational improvements.
- Engage iManage partners to establish shared reliability practices and improve ecosystem knowledge.
Requirements
- 3–5 years of experience in a technical escalation role in Support, Customer Reliability Engineering, Development, or Site Reliability Engineering.
- Ability to understand and communicate highly complex technical issues to non-technical and executive audiences.
- Hands-on experience leading P1/P2 incident response, running postmortems, and driving systemic fixes.
- Experience troubleshooting and supporting distributed cloud services.
- Strong technical proficiency with SQL, Python, Bash/Shell, PowerShell, and REST APIs.
- Solid understanding of Azure Kubernetes Service and related Azure services.
- Deep knowledge of observability and support platforms including Splunk, Grafana, Kibana, and Prometheus.
- Strong technical troubleshooting, problem-solving, automation, and self-service improvement skills.
Benefits
- Hybrid work policy with required in-office presence Tuesdays and Thursdays and flexible working hours.
- Market-competitive annual base salary and annual performance-based bonus.
- Health, vision, dental, and life insurance plus a 401(k) plan with company match up to 4%.
- Enhanced paid parental leave of 20 weeks for primary leave and 10 weeks for secondary leave.
- Flexible time off and multiple company wellness days.
- Access to RethinkCare behavioral health and well-being resources.
- Unlimited access to LinkedIn Learning and interactive Microsoft courses and training.
- Internal career development framework, certifications, open-plan workspace, free snacks and drinks, gaming area, and social events.
Tech Stack
Categories
Site Reliability
