3 months ago
Richmond, VA, USA or Jersey City, NJ, USASenior
Responsibilities
- Establish Exiger's SRE function, reliability standards, SLIs, SLOs, and error budgets.
- Build and own observability for availability, latency, and system health.
- Use metrics and measured outcomes to set reliability priorities and evaluate changes.
- Own production-service reliability from design consultation and launch reviews through steady-state operation.
- Automate repetitive manual operations and build infrastructure-as-code and self-service tooling.
- Plan for scale through capacity planning and performance analysis.
- Improve resilience through chaos engineering, fault-injection testing, and game days.
- Lead sustainable, blameless incident response and postmortems and participate in an on-call rotation.
- Use AI-assisted development tools such as Codex and Claude to accelerate automation, tooling, and investigations.
- Set practices and standards that can be adopted by other engineering teams.
Requirements
- A bachelor's or master's degree in Computer Science or a related field, or equivalent practical experience.
- At least 6 years of software or systems engineering experience, including at least 4 years in a dedicated Site Reliability Engineering, production engineering, or platform reliability role.
- At least 4 years of experience designing, analyzing, and troubleshooting large-scale distributed systems.
- Strong knowledge of Unix/Linux internals and networking fundamentals including TCP/IP, DNS, routing, and load balancing.
- Hands-on experience establishing SLIs, SLOs, error budgets, monitoring, observability, capacity planning, and automation.
- Experience with chaos engineering or fault-injection testing, such as game days, Chaos Monkey, Gremlin, or LitmusChaos.
- Proven incident-management experience, including on-call ownership, leading incident response, and conducting blameless postmortems.
- Experience supporting web services, data storage, databases, and data pipelines on Linux/Unix or other operating systems.
- Familiarity with AWS and secure system integration.
- Experience using AI coding assistants such as Claude and Codex in engineering workflows.
- At least 4 years of programming experience in Go or C is preferred; Java is also welcome.
- Experience supporting ML or data platforms in production is preferred.
- Familiarity with Snowflake, Redshift, and/or Apache Iceberg is preferred.
- Experience with FedRAMP or other regulated or government environments is preferred.
- The applicant must be a U.S. citizen and eligible for a U.S. security clearance.
- Willingness to travel as needed for customer engagements.
- Ability to translate ambiguous mission problems into structured technical solutions and operate independently in high-stakes environments.
Benefits
- Discretionary Time Off with no maximum limits.
- Health, vision, and dental benefits.
- 16 weeks of fully paid parental leave.
- Flexible hybrid work-from-home and office arrangement.
- Wellness stipends and dedicated wellness programming.
- Career development programs with reimbursement for educational certifications.
- The role is based in the U.S. and may require travel for customer engagements.
About Exiger
Exiger builds AI- and data-driven software for supply chain, third‑party risk, and compliance management, used by corporations, banks, and government agencies. Its 1Exiger platform provides supplier visibility, risk detection, and due diligence, delivered as SaaS with supporting analytics and advisory services. Headquartered in McLean, Virginia, Exiger is FedRAMP authorized for U.S. public‑sector use and is majority‑owned by The Carlyle Group.
