Alkami Technology

Sr. Platform Engineer

Alkami Technology
Apply
1 month ago
Remote, United StatesSenior

Base Salary

$145k - $165k/yr

Responsibilities

  • Investigate, troubleshoot, and resolve reliability issues in MANTL application code, microservice communications, and system integrations.
  • Design and maintain monitoring, dashboards, alerting, and distributed tracing to improve platform visibility and incident diagnosis.
  • Diagnose and remediate application performance issues involving caching, inefficient code paths, and slow queries.
  • Harden services through fault injection, resilience testing, defensive design patterns, idempotency, retries, and circuit breakers.
  • Build, maintain, and troubleshoot GitHub Actions deployment pipelines and Kubernetes workloads.
  • Maintain a roadmap of reliability risks and create documentation and runbooks for root causes and remediations.
  • Act as an escalation point and partner with cloud infrastructure and application engineering teams on cross-domain reliability issues.
  • Contribute to defining SLOs and SLIs for key platform services.

Requirements

  • 4–7 years of experience in software engineering, platform engineering, or a hybrid development/reliability engineering role.
  • Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent work experience.
  • Strong production proficiency with TypeScript in a Node.js/TypeScript service environment.
  • Hands-on experience deploying and troubleshooting containerized workloads in Kubernetes.
  • Experience building and maintaining CI/CD pipelines, with GitHub Actions preferred.
  • Experience configuring monitoring, dashboards, and alerting in an APM or observability tool, with Datadog preferred.
  • Experience troubleshooting distributed systems, microservice communication failures, integrations, and distributed tracing.
  • Working familiarity with relational databases and experience diagnosing application performance and query issues.
  • Experience designing for and testing failure modes, including fault injection and resilience or chaos-style testing.
  • Strong analytical, troubleshooting, communication, and independent problem-solving skills.
  • Preferred experience with Kafka or other message brokers/event-streaming platforms, OpenTelemetry or comparable tracing frameworks, regulated environments, Terraform or similar infrastructure-as-code tooling, and infrastructure/SRE partnerships.

Benefits

  • Remote-first work arrangement in the United States.
  • Unlimited paid time off.
  • 401(k) with employer match.
  • Diverse and inclusive company culture and additional benefits.

Categories

Site Reliability
Alkami Technology

About Alkami Technology

1,001-5,000 employees
Contact me