
Sr. Platform Engineer
Alkami Technology1 month ago
Remote, United StatesSenior
Base Salary
$145k - $165k/yr
Responsibilities
- Investigate, troubleshoot, and resolve reliability issues in MANTL application code, microservice communications, and system integrations.
- Design and maintain monitoring, dashboards, alerting, and distributed tracing to improve platform visibility and incident diagnosis.
- Diagnose and remediate application performance issues involving caching, inefficient code paths, and slow queries.
- Harden services through fault injection, resilience testing, defensive design patterns, idempotency, retries, and circuit breakers.
- Build, maintain, and troubleshoot GitHub Actions deployment pipelines and Kubernetes workloads.
- Maintain a roadmap of reliability risks and create documentation and runbooks for root causes and remediations.
- Act as an escalation point and partner with cloud infrastructure and application engineering teams on cross-domain reliability issues.
- Contribute to defining SLOs and SLIs for key platform services.
Requirements
- 4–7 years of experience in software engineering, platform engineering, or a hybrid development/reliability engineering role.
- Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent work experience.
- Strong production proficiency with TypeScript in a Node.js/TypeScript service environment.
- Hands-on experience deploying and troubleshooting containerized workloads in Kubernetes.
- Experience building and maintaining CI/CD pipelines, with GitHub Actions preferred.
- Experience configuring monitoring, dashboards, and alerting in an APM or observability tool, with Datadog preferred.
- Experience troubleshooting distributed systems, microservice communication failures, integrations, and distributed tracing.
- Working familiarity with relational databases and experience diagnosing application performance and query issues.
- Experience designing for and testing failure modes, including fault injection and resilience or chaos-style testing.
- Strong analytical, troubleshooting, communication, and independent problem-solving skills.
- Preferred experience with Kafka or other message brokers/event-streaming platforms, OpenTelemetry or comparable tracing frameworks, regulated environments, Terraform or similar infrastructure-as-code tooling, and infrastructure/SRE partnerships.
Benefits
- Remote-first work arrangement in the United States.
- Unlimited paid time off.
- 401(k) with employer match.
- Diverse and inclusive company culture and additional benefits.
Tech Stack
Categories
Site Reliability