Alkami Technology

Staff Platform Engineer (MANTL)

Alkami Technology
Apply
1 month ago
Remote, United StatesStaff+

Base Salary

$140k - $175k/yr

Responsibilities

  • Investigate, troubleshoot, and resolve complex reliability issues in MANTL application code, including microservice communication failures and failure-condition correctness problems.
  • Define resilience strategies and failure-mode practices for third-party and internal integrations.
  • Own monitoring, dashboards, alerting, and distributed tracing strategy across the platform.
  • Diagnose and remediate application performance issues involving caching, inefficient code paths, and query performance.
  • Lead fault-injection, resilience, and defensive-design practices, including idempotency, retries, and circuit breakers.
  • Own and improve GitHub Actions CI/CD pipeline architecture and troubleshoot Kubernetes workloads.
  • Define and execute a roadmap for known reliability risks independently of feature-delivery timelines.
  • Establish reliability documentation, runbook, SLO, and SLI standards.
  • Serve as the senior escalation point for complex reliability issues and advise leadership on reliability risks and tradeoffs.
  • Provide technical mentorship and guidance to other Platform Engineers.

Requirements

  • 7 to 10 years of experience in software engineering, platform engineering, or a hybrid development/reliability engineering role.
  • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent work experience.
  • Deep production experience with TypeScript and Node.js service environments.
  • Strong experience deploying, operating, and troubleshooting containerized workloads in Kubernetes.
  • Experience designing, building, and maintaining CI/CD pipelines, with GitHub Actions preferred.
  • Experience designing monitoring, dashboards, and alerting strategies using APM or observability tools, with Datadog preferred.
  • Deep experience troubleshooting distributed systems and microservice communication failures.
  • Strong experience with distributed tracing tools and practices at scale.
  • Working familiarity with relational databases and diagnosing complex query and schema-level performance issues.
  • Experience with fault injection, resilience or chaos-style testing, and defensive software patterns.
  • Ability to independently lead ambiguous, high-impact reliability work and set technical direction.
  • Excellent communication skills for presenting root causes, risks, and remediation plans to technical and non-technical leadership.
  • Experience mentoring engineers.
  • Preferred experience with Kafka or other message brokers and event-streaming platforms.
  • Preferred experience with OpenTelemetry or comparable distributed tracing frameworks.
  • Experience in a regulated or compliance-driven environment is preferred.
  • Familiarity with Terraform or similar infrastructure-as-code tooling is preferred.
  • Experience partnering with infrastructure or SRE teams and influencing engineering roadmap prioritization is preferred.

Benefits

  • Remote-first work environment with this position eligible for remote work in the United States.
  • Unlimited paid time off.
  • 401(k) with employer match.
  • Diverse and inclusive company culture.
  • Employment sponsorship is not available; candidates must be eligible to work full-time in the United States.

Categories

DevOpsSite Reliability
Alkami Technology

About Alkami Technology

1,001-5,000 employees
Contact me