9 months ago
Base Salary
$150k - $350k/yr
Responsibilities
- Identify architectural changes that improve reliability, performance, and availability.
- Foster a culture of reliability across the engineering organization.
- Design and implement operational processes for deployments, upgrades, rollbacks, and postmortem reviews.
- Join a core engineering team and participate in the on-call rotation.
- Build monitoring systems that ensure high-quality service for customers.
- Debug production issues across all services and levels of the stack.
Requirements
- At least five years of experience writing high-quality production code.
- At least two years of on-call experience for critical production services.
- Strong cloud skills and deep familiarity with at least one hyperscaler cloud, with AWS preferred.
- Familiarity with autoscaling, fleet management, and capacity planning at scale.
- Experience owning and scaling Kubernetes clusters to thousands of nodes is a plus.
- Experience with systems safety research such as STAMP and control theory is a plus.
- Ability to work in person in the NYC, San Francisco, or Stockholm office.
- Ability to participate in an on-call rotation and respond to production incidents.
Benefits
- Work in person at the NYC, San Francisco, or Stockholm office.
- Participate in an on-call rotation for production incidents.
- Opportunity to define reliability systems and practices as the company's first reliability-focused hire.
Tech Stack
Categories
DevOpsSite Reliability
About Modal
Modal builds a serverless compute platform for AI and data workloads, offering instant GPU access, sub-second container starts, and native storage to run inference, fine-tuning, and batch jobs. It sells a usage-based cloud service to developers and ML teams to deploy generative models and pipelines. Privately held and headquartered in New York City, its customers include companies like DoorDash and Ramp.
