
Site Reliability Engineer - Observability Platform
Ford Motor10 hours ago
Remote, United StatesMid Level
Base Salary
$85k - $193k/yr
Responsibilities
- Design and implement observability pipelines covering metrics, logging, tracing, and alerting.
- Define and operationalize SLIs, SLOs, and error budgets to improve availability and uptime.
- Build reusable infrastructure-as-code templates and frameworks for observability instrumentation and onboarding.
- Develop automation that improves system resilience, recoverability, availability, scalability, reliability, and time-to-market.
- Reduce operational toil through automation and improve supported software solutions.
- Collaborate with development teams to design, build, and operate scalable, resilient, cloud-native systems.
- Identify stability risks, develop mitigation plans, and review technical performance and capacity metrics.
- Conduct performance analysis and optimization for new and production systems.
- Troubleshoot distributed production systems, lead root-cause analysis, and participate in incident response, recovery, and postmortems.
- Evaluate emerging technologies and integrate AI/ML capabilities for anomaly detection, alerting precision, and performance insights.
- Mentor team members and promote observability best practices across engineering teams.
Requirements
- Bachelor’s degree in Computer Science or equivalent experience.
- At least 3 years of experience in an SRE role.
- At least 5 years of programming experience with Python, Go, Java/Scala, C, or C++.
- At least 3 years of experience building reusable infrastructure-as-code templates and frameworks with Terraform or ToFu.
- At least 3 years of experience with APM and monitoring tools such as Dynatrace, New Relic, ELK, Splunk, Prometheus, Sensu, Nagios, Kafka, or DataDog.
- At least 3 years of experience with J2EE, NoSQL/SQL datastores, Spring Boot, GCP/AWS/Azure, and Docker/Kubernetes for multi-tier applications.
- Experience with RESTful APIs and microservices platforms.
- Working knowledge of TCP/IP, internet routing, and load balancing.
- Strong proficiency with Google Cloud Platform and its services.
- Experience with automated, test-driven development in CI/CD pipelines.
- Thorough understanding of software development and agile methodologies.
- Ability to implement observability strategies that improve mean time to detect and mean time to resolve.
Benefits
- Immediate medical, dental, vision, and prescription drug coverage.
- Flexible family care days, paid parental leave, new-parent ramp-up programs, and subsidized backup childcare.
- Family-building benefits including adoption and surrogacy reimbursement and fertility treatments.
- Vehicle discount programs for employees and family members and management leases.
- Tuition assistance and active employee resource groups.
- Paid time off for individual and team community service.
- Paid holidays, including the week between Christmas and New Year’s Day, plus paid time off and the option to purchase additional vacation time.
- Remote position, indicated by the posting’s remote location tag.
Tech Stack
Apache KafkaAWSAzureCC++DatadogDockerGoGoogle Cloud PlatformJavaKubernetesNagiosPrometheusPythonScalaSplunkSpring BootSQLTerraform
Categories
DevOpsSite Reliability
About Ford Motor
Ford Motor Company designs, manufactures, and sells cars, trucks, SUVs, and commercial vehicles for consumers and businesses, and provides financing through Ford Credit. A public company headquartered in Dearborn, Michigan and listed on the NYSE, it operates globally under the Ford and Lincoln brands, including electric and connected vehicles such as the F‑150 Lightning and Mustang Mach‑E. Founded in 1903, it also offers software-driven fleet, mobility, and aftersales services.