
Site Reliability Engineer | AI Infrastructure
Jones Lang LaSalle Incorporated2 months ago
Tel Aviv-Yafo, IsraelSenior
Responsibilities
- Design and maintain deployment and release infrastructure for cloud-native AI agents on Azure and AWS.
- Build monitoring and observability for AI services, including error tracking, cost monitoring, response-quality monitoring, and upstream API change detection.
- Own the implementation of infrastructure security and compliance standards for sensitive enterprise workflows.
- Create developer tooling and internal platform patterns for building, testing, and deploying services.
- Participate in on-call support, production incident response, and post-mortems.
- Write production code in TypeScript and Python and contribute to data pipelines and product features.
Requirements
- Require 5+ years of experience in SRE, platform engineering, DevOps, or infrastructure roles with end-to-end infrastructure ownership.
- Require strong experience with Azure or AWS, Docker, Kubernetes, and CI/CD pipelines.
- Require infrastructure-as-code experience with Terraform, CDK, or CloudFormation.
- Require monitoring and observability experience with Datadog, Splunk, CloudWatch, or similar tools.
- Require infrastructure fundamentals including Linux, networking, and security.
- Require production incident and on-call experience, including post-mortems.
- Require independent work, broad ownership, high accountability, and strong written and verbal English.
- Prefer experience with model serving, LLM API integration, vector databases, evaluation pipelines, self-service developer tooling, internal platforms, cloud and API cost optimization, and enterprise security engineering.
- Prefer the ability to write production TypeScript or Python rather than only scripts.
Benefits
- Hybrid work arrangement with three days in the Tel Aviv office.
- Compensation package includes base salary, RSUs with standard four-year vesting, keren hishtalmut, and an annual bonus; range details are shared early in the process.
- Distributed team collaboration across Israel, Czechia, the US, India, and Singapore.
- English is the team’s working language; Hebrew is not required.
- Technical interview process includes a hiring-manager conversation, system-design and live-troubleshooting deep dive without LeetCode, and a team conversation, typically completed in 2–3 weeks.
Tech Stack
Categories
DevOpsSite Reliability