
Staff Software Engineer - Resiliency and Platform Engineering
Choice Hotels International, Inc.2 months ago
Scottsdale, AZ, USAStaff+
Responsibilities
- Design and implement shared platform libraries, frameworks, tooling, automation, and guardrails that improve resiliency, runtime safety, and developer experience.
- Identify and eliminate systemic failure modes, including JVM memory leaks, unsafe defaults, brittle error handling, poor failure propagation, and resource exhaustion.
- Define and roll out developer-facing standards and paved roads for resiliency, observability, error handling, and operational readiness.
- Standardize logging, monitoring, alerting, and observability practices to improve signal quality and speed diagnosis and recovery.
- Partner with Principal Software Engineers, Solution Architects, Engineering Managers, and application teams to identify systemic risks and develop platform initiatives.
- Collaborate across software engineering resiliency, data engineering resiliency, and platform engineering teams on cross-cutting solutions.
- Engage with application codebases to understand system behavior, identify failure patterns, and validate resiliency improvements before shifting focus to systemic solutions.
- Participate in incident postmortems and operational reviews and convert recurring lessons into durable platform improvements.
- Evaluate and introduce tools and technologies that improve developer productivity, platform safety, and application resiliency.
- Apply AI-assisted development, diagnostic, and operational tools when they measurably improve engineering effectiveness or reliability outcomes.
- Influence technical direction through design reviews, reference implementations, mentorship, and technical leadership without formal delivery ownership.
Requirements
- Bachelor’s degree in computer science or a related technical field, or equivalent practical experience building and operating production systems.
- Typically 8–10+ years of hands-on experience designing, building, and supporting large-scale production software systems.
- Hands-on experience designing, building, and operating Java-based services, including Spring Boot applications in virtualized and containerized environments.
- Experience developing and supporting cloud-native and serverless workloads, including Python services and event-driven architectures.
- Strong practical experience with AWS public cloud environments and the reliability, scalability, and operational effects of cloud-managed services.
- Working knowledge of relational and non-relational data stores and their persistence, availability, and failure characteristics.
- Experience with application monitoring and observability platforms such as AppDynamics, OpenSearch, Amazon CloudWatch, or similar tools.
- Ability to diagnose complex production issues using metrics, logs, traces, and runtime signals.
- Understanding of Site Reliability Engineering principles and judgment to apply them selectively to platform and resiliency improvements.
- Experience delivering platform-level capabilities such as shared libraries, frameworks, internal tooling, or enablement platforms used by multiple application teams.
- Experience creating and rolling out paved roads, guardrails, or standardized patterns that balance safety, usability, and developer autonomy.
- Experience using AI-assisted tools for engineering effectiveness, log or trace analysis, incident analysis, or system reliability.
- Ability to influence technical direction and engineering practices across teams without direct ownership of delivery backlogs.
- Cloud or technology certifications, such as AWS certifications or equivalent, are preferred.
Benefits
- Competitive compensation and benefits, including medical, dental, and vision coverage.
- Paid leave and holidays, vacation, personal, family, volunteer, sick, jury duty, bereavement, military, and religious-observance leave.
- Retirement and health savings benefits.
- Employee recognition programs.
- Discounts at Choice hotels worldwide.
- Four days onsite in a hybrid arrangement at the N. Scottsdale office.
- Individual contributor role reporting to the Manager, Resiliency and Platform Engineering.