Principal, Platform Engineering
Office of the Comptroller of the Currency1 day ago
Base Salary
$175k - $292k/yr
Responsibilities
- Elevate SRE, observability, monitoring, reliability, incident response, SLO/SLI, error-budget, and postmortem practices across the platform organization.
- Act as product owner for observability by defining its roadmap, requirements, and priorities in partnership with the dedicated SRE and monitoring team.
- Drive DORA metrics, cycle time, throughput, cloud utilization, rightsizing, FinOps, waste reduction, and performance optimization.
- Build and improve reliable, scalable, and secure cloud platform technology in partnership with engineers, architects, and leaders.
- Champion infrastructure-as-code, automated provisioning, configuration management, version control, and IaC governance across teams.
- Plan cloud capacity, lead technology roadmaps and end-of-life plans, and translate reliability signals into platform investments.
- Communicate operational issues and recommendations to senior management and participate in production changes, maintenance windows, and on-call rotation.
- Serve as an escalation point for platform reliability and observability issues while maintaining system documentation and regulatory compliance.
Requirements
- Bachelor’s degree, preferably in a technical discipline, or an equivalent combination of education and experience; a master’s degree is also considered.
- 10+ years of progressive hands-on software engineering experience with large-scale computing solutions and 10+ years of hands-on IT systems installation, operations, administration, and cloud or virtualized-server maintenance.
- Demonstrated expertise building and maturing SRE and observability practices at scale, including metrics, distributed tracing, centralized logging, dashboards, alerting, SLOs, SLIs, error budgets, incident management, and postmortems.
- Experience driving data-based reliability improvements, toil reduction, cloud efficiency, performance optimization, and software delivery metrics such as DORA metrics.
- Hands-on experience with Terraform, Kubernetes, Jenkins or other CI/CD tooling, Kafka, GitHub, and configuration management tools such as Puppet, Chef, or Ansible.
- Deep infrastructure-as-code experience with tools such as Terraform, CloudFormation, CDK, or Pulumi, including adoption, governance, module libraries, policy-as-code, or drift detection.
- Experience with AWS, Azure, infrastructure design, servers, operating systems, networks, storage, highly available mission-critical environments, and 24x7 availability.
- Strong consultative, analytical, communication, technical influence, and informal mentorship skills, including the ability to lead without formal authority.
- Experience acting as a product owner for a platform capability and defining roadmaps, requirements, and priorities.
- Preferred experience includes engineering metrics platforms such as LinearB, Jellyfish, Sleuth, or Haystack; scripting or development in Python, Ruby, Go, or Java; regulated or financial-services environments; production change control; and centralized SRE organizations.
- AWS Solutions Architect Associate certification or higher is strongly desired; Azure, Google Cloud, and other relevant industry certifications are preferred.
Benefits
- Hybrid work environment with up to 2 days per week of remote work.
- Tuition reimbursement and student loan repayment assistance.
- Technology stipend for remote network access and device use.
- Generous paid time off and parental leave.
- 401(k) employer match.
- Competitive medical, dental, and vision benefits.
- Potential discretionary annual incentive compensation is available in addition to base salary.
Tech Stack
AnsibleApache KafkaAWSAzureChefDatadogGoGoogle CloudGrafanaJavaJenkinsKubernetesPrometheusPuppetPythonRubyTerraform
Categories
DevOpsSite Reliability