3 months ago
Tel Aviv-Yafo, IsraelSenior
Responsibilities
- Support the development and production operation of AI-based application solutions and AI-powered travel experiences.
- Partner with teams integrating AI capabilities, providers, and APIs while addressing reliability, authentication, quotas, rate limits, latency, and provider-specific operational constraints.
- Diagnose failures across AI-powered workflows, provider APIs, configurations, permissions, degraded responses, and related systems.
- Implement and operate reliable cloud infrastructure while helping product teams move quickly without compromising reliability.
- Build dashboards, alerts, traces, logs, and runbooks tied to SLOs and customer impact.
- Prototype and productionize AI-assisted systems that improve SRE workflows.
- Create tools, workflows, and automation to reduce repetitive operational work and improve access to operational knowledge.
- Collaborate with product, platform, data, security, support, and incident-response teams to improve production resilience.
Requirements
- At least 5 years of experience as a Senior SRE, Infrastructure Software Engineer, Production Engineer, or DevOps Engineer.
- At least 3 years of experience operating production, 24x7 customer-facing systems.
- Hands-on experience delivering production infrastructure, platform tooling, and automation used by engineering teams.
- Strong software engineering skills in Python, Go, Java, or a similar language, with an emphasis on production-quality code, tests, monitoring, and documentation.
- Experience with cloud infrastructure, container orchestration, Linux systems, networking, CI/CD, and infrastructure as code such as Terraform or CloudFormation.
- Experience building, tuning, and automating observability systems such as Grafana, Prometheus, New Relic, Datadog, Splunk, or similar tools.
- Familiarity with SLOs, incident response, on-call practices, root cause analysis, and blameless postmortems.
- Practical experience or strong interest in AI solutions, AI providers, agents, AI APIs, provider integrations, or AI-assisted internal tools.
- Ability to troubleshoot AI tools and provider/API issues involving rate limits, quotas, authentication, permissions, latency, SDK or API contract changes, content quality, and service degradations.
- Excellent communication skills and ability to work with stakeholders and domain experts across the company.
Benefits
- Position based out of Navan's Tel Aviv office.
About Navan
Navan builds an AI-powered travel and expense management platform for businesses, combining online booking, spend controls, automated expense reporting, and corporate cards with 24/7 support. The company sells its T&E software and payments tools to enterprises and high-growth firms, generating revenue from subscriptions, travel transactions, and card interchange. Founded in 2015 and headquartered in Palo Alto, it rebranded from TripActions and serves customers globally.