1 month ago
Chicago, IL, USA +2 moreSenior
Base Salary
$102k - $179k/yr
Responsibilities
- Design, develop, and maintain backend services, APIs, platform components, and distributed systems for AI/ML applications.
- Build scalable cloud infrastructure with Pulumi or other infrastructure-as-code tools and develop deployment workflows and containerized environments.
- Support production deployment of ML models, LLM applications, RAG pipelines, agentic AI systems, model-serving systems, inference pipelines, and data workflows.
- Improve observability through monitoring, logging, tracing, alerting, incident management, and production support practices.
- Establish standards for testing, code quality, security, maintainability, scalability, latency optimization, and cost efficiency.
- Define reusable platform patterns, developer tooling, and engineering workflows that improve developer productivity and operational consistency.
- Partner with product, security, data, platform, AI/ML, and cross-functional teams on production-ready solutions and long-term platform strategy.
- Troubleshoot complex production issues, conduct root cause analysis, and drive remediation to improve system reliability.
- Mentor engineers through technical collaboration, code reviews, knowledge sharing, and best-practice guidance.
- Evaluate emerging AI engineering trends, AI-assisted development tools, and modern software practices.
Requirements
- A relevant degree is preferred but not required.
- At least 5 years of relevant experience is required.
- Strong Python development experience building production backend services, APIs, distributed systems, and cloud-based applications is required.
- Hands-on experience with Pulumi, Terraform, or other infrastructure-as-code tools and cloud platforms such as AWS, Azure, or GCP is required.
- Experience with Docker, CI/CD pipelines, infrastructure automation, container orchestration, Kubernetes, serverless architectures, and cloud deployment practices is preferred.
- Knowledge of observability tools, monitoring, logging, tracing, alerting frameworks, and production support practices is required.
- Familiarity with ML/AI systems, model serving, LLM applications, inference pipelines, RAG workflows, data pipelines, or related AI platform technologies is required.
- Experience with Databricks, Azure AI Foundry, or similar AI/ML platform technologies is preferred.
- Strong analytical, troubleshooting, problem-solving, verbal communication, and written communication skills are required.
- Ability to collaborate across technical and business teams and operate with ownership, accountability, and adaptability in evolving environments is required.
Benefits
- Comprehensive benefits plan.
- Estimated base salary range of $102,400.00 to $179,000.00 annually, with incentive eligibility.
- Geographic factors may adjust the salary range, and hires typically fall below the top of the range.
