
Staff Platform Engineer (MANTL)
Alkami Technology1 month ago
Remote, United StatesStaff+
Base Salary
$140k - $175k/yr
Responsibilities
- Investigate, troubleshoot, and resolve complex reliability issues in MANTL application code, including microservice communication failures and failure-condition correctness problems.
- Define resilience strategies and failure-mode practices for third-party and internal integrations.
- Own monitoring, dashboards, alerting, and distributed tracing strategy across the platform.
- Diagnose and remediate application performance issues involving caching, inefficient code paths, and query performance.
- Lead fault-injection, resilience, and defensive-design practices, including idempotency, retries, and circuit breakers.
- Own and improve GitHub Actions CI/CD pipeline architecture and troubleshoot Kubernetes workloads.
- Define and execute a roadmap for known reliability risks independently of feature-delivery timelines.
- Establish reliability documentation, runbook, SLO, and SLI standards.
- Serve as the senior escalation point for complex reliability issues and advise leadership on reliability risks and tradeoffs.
- Provide technical mentorship and guidance to other Platform Engineers.
Requirements
- 7 to 10 years of experience in software engineering, platform engineering, or a hybrid development/reliability engineering role.
- Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent work experience.
- Deep production experience with TypeScript and Node.js service environments.
- Strong experience deploying, operating, and troubleshooting containerized workloads in Kubernetes.
- Experience designing, building, and maintaining CI/CD pipelines, with GitHub Actions preferred.
- Experience designing monitoring, dashboards, and alerting strategies using APM or observability tools, with Datadog preferred.
- Deep experience troubleshooting distributed systems and microservice communication failures.
- Strong experience with distributed tracing tools and practices at scale.
- Working familiarity with relational databases and diagnosing complex query and schema-level performance issues.
- Experience with fault injection, resilience or chaos-style testing, and defensive software patterns.
- Ability to independently lead ambiguous, high-impact reliability work and set technical direction.
- Excellent communication skills for presenting root causes, risks, and remediation plans to technical and non-technical leadership.
- Experience mentoring engineers.
- Preferred experience with Kafka or other message brokers and event-streaming platforms.
- Preferred experience with OpenTelemetry or comparable distributed tracing frameworks.
- Experience in a regulated or compliance-driven environment is preferred.
- Familiarity with Terraform or similar infrastructure-as-code tooling is preferred.
- Experience partnering with infrastructure or SRE teams and influencing engineering roadmap prioritization is preferred.
Benefits
- Remote-first work environment with this position eligible for remote work in the United States.
- Unlimited paid time off.
- 401(k) with employer match.
- Diverse and inclusive company culture.
- Employment sponsorship is not available; candidates must be eligible to work full-time in the United States.
Tech Stack
Categories
DevOpsSite Reliability