1 day ago
Base Salary
$153k - $205k/yr
Responsibilities
- Design, build, secure, and operate scalable Kubernetes platforms for critical production services across hybrid and public-cloud environments.
- Build reusable Terraform modules and reviewable, automated infrastructure delivery workflows.
- Develop backend services, internal tools, and operational automation in Go, Python, or JavaScript/TypeScript.
- Partner with engineering and product teams to design reliable, performant, secure, cost-effective solutions for workload requirements.
- Improve the production lifecycle through deployment automation, progressive delivery, observability, and clear operational ownership.
- Define and operate SLIs, SLOs, error budgets, capacity plans, disaster-recovery testing, and resilience improvements.
- Participate in on-call, lead incident response, perform root-cause analysis, and drive blameless postmortems and durable corrective actions.
- Embed security and compliance into platform operations and partner with Security teams in regulated environments.
- Apply AI-assisted and data-driven operational techniques to improve signal detection, reduce alert noise, and accelerate root-cause analysis.
- Contribute through code reviews, documentation, knowledge sharing, mentoring, and team development.
Requirements
- 5+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or a closely related software engineering role supporting production systems.
- Deep hands-on experience designing, operating, securing, and troubleshooting production Kubernetes clusters and containerized workloads at scale.
- Strong Terraform experience, including reusable modules, state and environment management, and automated infrastructure delivery.
- Production software-development experience in Go, Python, or JavaScript/TypeScript for maintainable backend services, tooling, and automation.
- Demonstrated success improving the reliability, performance, scalability, or cost efficiency of distributed production systems.
- Experience with cloud infrastructure and networking concepts including IAM, DNS, load balancing, routing, service networking, and secure connectivity.
- Strong observability and troubleshooting skills across metrics, logs, traces, alerting, and incident data.
- Experience with SLIs, SLOs, error budgets, incident management, postmortems, and disaster-recovery practices.
- Familiarity with CI/CD, GitOps or deployment automation, and canary or blue-green deployment strategies.
- Security-minded approach and experience partnering with Security and engineering teams in regulated or high-availability environments.
- Clear written and verbal communication, strong ownership, and sound judgment balancing speed, risk, and operational excellence.
- Experience applying AI-assisted tooling to engineering or operations workflows is preferred.
Benefits
- Remote work arrangement is indicated by the #LI-Remote designation.
- Circle offers an inclusive work environment centered on transparency, collaboration, and flexible work practices.
Categories
DevOpsSite Reliability
About Circle
Circle builds a platform for businesses to use digital dollars and blockchain for payments, treasury, and programmable finance. It issues USDC, a widely used dollar stablecoin, and offers APIs, the Circle Payments Network, and Arc, an enterprise-focused blockchain. Founded in 2013 and publicly traded on the NYSE (CRCL), Circle operates remote-first and serves enterprises, financial institutions, and developers worldwide.
