ToolsGroup

Senior Infrastructure Engineer

ToolsGroup
Apply
15 days ago
Turin, Italy or Milan, ItalySenior

Base Salary

$55k - $68k/yr

Responsibilities

  • Lead major incidents from impact assessment and containment through recovery, communication, root-cause analysis, and corrective actions.
  • Troubleshoot distributed production systems across Windows, Linux, Kubernetes, containers, networking, DNS, cloud infrastructure, identity, databases, storage, APIs, and service dependencies.
  • Operate and improve Azure, OCI, or comparable cloud environments, including monitoring, access controls, backup and recovery, reliability, and cost-aware scaling.
  • Define and improve service-level indicators and objectives, observability, alert quality, capacity, resilience, dependency mapping, and recovery readiness.
  • Automate operational tasks and controls using PowerShell, Python, infrastructure as code, and deployment pipelines with validation, secure credential handling, and rollback.
  • Apply incident, change, and problem management while protecting service availability.
  • Lead and mentor technical professionals and align IT/Ops, Engineering, Product, Security, Support, and business stakeholders on service ownership, risks, priorities, and commitments.

Requirements

  • 5+ years of experience operating business-critical, customer-facing, or high-availability services in SRE or production engineering roles.
  • Recent direct technical ownership of high-severity incidents, including triage, recovery decisions, communications, and measurable follow-through.
  • Strong troubleshooting fundamentals across Windows, Linux, TCP/IP, DNS, routing, firewalls, proxies, and load balancers.
  • Hands-on cloud operations experience with Azure, OCI, or a similar platform, including compute, storage, networking, IAM, observability, backup, and recovery.
  • Production experience with containers and Kubernetes, including workload health, scheduling, networking, persistent storage, secrets, deployment, rollback, scaling, and recovery.
  • Deep experience with metrics, logs, distributed tracing, actionable alerting, SLI/SLO design, capacity analysis, failure-mode thinking, and post-incident engineering.
  • Ability to diagnose database-backed and API-driven services across application, query, connection-pool, storage, certificate, network, and downstream dependency layers.
  • Practical identity and access management experience with Active Directory and Microsoft Entra ID or equivalent, including hybrid identity, privileged access, MFA, service identities or gMSAs, lifecycle controls, and recovery.
  • Experience with safe automation and infrastructure-as-code using PowerShell, Python, Terraform, or comparable tools and CI/CD, including testing, idempotency, error handling, credential security, rollout, and rollback.
  • Experience leading, mentoring, or serving as a senior escalation point for technical professionals, with strong business-facing and cross-functional communication skills.
  • Additional relevant experience with Microsoft 365, endpoint management, EDR, device compliance, hybrid workplace operations, or formal ITIL, cloud, security, or infrastructure certifications.

Tech Stack

Categories

DevOpsSite Reliability
ToolsGroup

About ToolsGroup

201-500 employees

ToolsGroup builds AI-powered supply chain planning software for retailers, consumer goods brands, manufacturers, and distributors. Its platform covers demand forecasting, inventory optimization, S&OP, and replenishment using probabilistic models, and integrates with ERP systems and major clouds. Headquartered in Boston and privately held with private equity investment in 2018, the company has deployments in more than 44 countries.

Contact me