1 day ago
Remote, CroatiaSenior
Responsibilities
- Lead reliability automation projects end to end, from problem definition and design through delivery, adoption, and measurable operational impact.
- Build and maintain shared automation, tooling, and platform services that improve production reliability, scalability, and operability.
- Automate SLOs, SLIs, error budgets, alerting, deployment safety, service recovery, self-healing, and auto-remediation.
- Automate incident-response workflows, diagnostics, runbook execution, monitoring, alerting, logging, and tracing.
- Own the reliability and availability of the team’s automation and platform services, including on-call response.
- Use incident and post-incident findings to create durable engineering fixes and eliminate operational toil.
- Design and scale chaos engineering, Gameday, disaster recovery, and resilience-testing tooling.
- Partner with product engineering, platform, infrastructure, operations, security, and other teams to drive adoption of reliability capabilities.
- Participate in a 24x7 on-call rotation for owned automation, tooling, and platform services.
Requirements
- B.S. or higher in Computer Science, Engineering, or a related technical discipline, or equivalent practical experience.
- At least 7 years of experience in Software Engineering, Site Reliability Engineering, Platform Engineering, or a closely related discipline.
- Strong software engineering fundamentals, including API and service design, testing, code review, and production software delivery.
- Practical experience applying SLOs, SLIs, observability, incident response, and toil reduction.
- Experience building on AWS, Azure, or another public cloud platform.
- Strong programming skills in Python, Go, Java, or similar languages.
- Experience with Infrastructure as Code, CI/CD pipelines, and deployment automation.
- Ability to work independently, handle ambiguous problems, and drive cross-team projects to completion.
- Experience building production automation, self-healing or event-driven systems, internal platforms, developer tooling, or services adopted by engineering teams.
- Experience with containers, Kubernetes, cloud-native architectures, distributed systems, observability platforms, incident-management platforms, workflow orchestration, job scheduling, chaos engineering, disaster recovery, or resilience testing is preferred.
- Experience with AI-assisted operational tooling, including LLM-based incident triage, diagnostics, or agentic remediation workflows, is preferred.
- Strong written, verbal, and cross-functional collaboration skills.
Benefits
- For Croatia-based roles, the stated starting base salary is €38,000 to €55,000 per year.
- The compensation package may include annual cash bonuses, commissions for sales roles, stock grants, and comprehensive benefits.
- The role includes a 24x7 on-call rotation.
- The role may require in-person onboarding and/or in-person identity verification.
About Autodesk
Autodesk builds design and engineering software used in architecture, construction, manufacturing, and media and entertainment. Its subscription-based portfolio includes AutoCAD, Revit, Fusion 360, Maya, and 3ds Max, along with cloud collaboration and simulation services. Founded in 1982 and headquartered in San Francisco, the public company trades on NASDAQ as ADSK and serves customers ranging from AEC firms to product designers and film and game studios.
