6 months ago
Responsibilities
- Design and implement scalable Kubernetes-based solutions for secure enterprise workloads.
- Define and contribute to reliability initiatives including SLOs, capacity planning, and resilience improvements.
- Identify and resolve system misconfigurations and architectural risks while developing clear failure modes and mitigation strategies.
- Own and continuously improve internal platform systems, deployments, environments, and developer tooling.
- Improve reliability, observability, and operational workflows.
- Build guardrails and self-service tooling that help product engineers deliver more effectively.
- Contribute to internal engineering standards and design principles and collaborate with engineering, product, and leadership to accelerate delivery.
Requirements
- Extensive production-grade experience scaling Kubernetes infrastructure from early-stage systems to production-critical workloads.
- Deep understanding of distributed systems tradeoffs and a strong systems mindset.
- Experience in big tech followed by a startup, or at a startup that successfully scaled its infrastructure.
- Ability to operate in ambiguity and act as a technical owner with broad company impact.
- Experience with single-tenant architectures is a bonus.
- Background in financial services or regulated environments is a bonus.
- Familiarity with banking-level security and data governance is a bonus.
Benefits
- Opportunity to work on a rapidly scaling enterprise AI platform serving investment banks, hedge funds, and private equity firms.
- Hands-on ownership of platform architecture, reliability, scalability, and developer experience.
- Fast-paced startup environment with exposure to frontier AI technology, reinforcement learning, and published research.
Tech Stack
Categories
About Rogo
Rogo is the purpose-built AI platform for finance.
