2 months ago
Buenos Aires, Argentina +3 moreSenior
Responsibilities
- Design, implement, and maintain secure, reliable, scalable cloud infrastructure.
- Build and improve CI/CD pipelines and automate infrastructure provisioning, configuration, deployment, and environment management.
- Maintain infrastructure as code and consistent development, testing, staging, and production environments.
- Build and maintain containerized application environments and orchestration capabilities.
- Implement monitoring, logging, tracing, alerting, dashboards, production-health reporting, SLIs, SLOs, and operational metrics.
- Improve resiliency, scalability, availability, disaster recovery readiness, deployment strategies, and automated rollback capabilities.
- Manage secrets, certificates, identity, permissions, network controls, cloud security configurations, vulnerability remediation, patching, and infrastructure hardening.
- Diagnose and resolve production incidents involving infrastructure, networking, deployments, performance, capacity, and cloud services.
- Participate in incident response, post-incident reviews, root-cause analysis, corrective-action planning, production support, and an appropriate on-call rotation.
- Reduce cloud waste, improve cost visibility, and build self-service tools and reusable deployment patterns for application engineering teams.
- Document infrastructure architecture, operational procedures, recovery processes, and troubleshooting guidance.
- Use AI-assisted engineering tools for scripting, troubleshooting, documentation, and infrastructure analysis.
Requirements
- Approximately 5+ years of experience in DevOps, SRE, Cloud Engineering, Platform Engineering, or a closely related role.
- Strong experience operating production workloads in AWS and with infrastructure as code using Terraform, CloudFormation, or Pulumi.
- Experience building and maintaining CI/CD pipelines with tools such as GitHub Actions, GitLab CI, Jenkins, CircleCI, or AWS CodePipeline.
- Strong knowledge of Docker, containerized application delivery, and orchestration platforms such as Kubernetes or Amazon ECS.
- Strong understanding of cloud networking, including VPCs, subnets, routing, load balancers, DNS, firewalls, gateways, and private connectivity.
- Experience with cloud IAM, role-based access, least-privilege design, secrets management, logging, metrics, dashboards, alerts, and distributed tracing.
- Familiarity with Datadog, New Relic, Grafana, Prometheus, CloudWatch, or OpenTelemetry.
- Strong Linux and command-line skills, plus scripting experience with Python, Bash, JavaScript, or a comparable language.
- Understanding of deployment patterns for modern web applications, APIs, background workers, and event-driven systems.
- Experience supporting relational databases, backups, restore processes, and high availability.
- Knowledge of incident management, root-cause analysis, capacity planning, production-readiness reviews, cloud security, vulnerability remediation, encryption, certificates, key management, and compliance-oriented controls.
- Preferred experience includes multi-tenant SaaS, serverless infrastructure, event-driven cloud services, PostgreSQL, managed databases, CDNs, media or video workloads, enterprise security and compliance frameworks, internal developer platforms, resilience or disaster-recovery testing, cloud cost optimization, and globally distributed engineering teams.
Benefits
- Work on real-world AI-driven projects across industries and enterprise transformation initiatives.
- Collaborate with a global team across continents and cultures.
- Inclusive environment emphasizing continuous learning, innovation, and ethical AI standards.
Tech Stack
AWSBashCircleCIDatadogDockerGitHub ActionsGitLab CI/CDGrafanaJavaScriptJenkinsKubernetesLinuxPostgreSQLPrometheusPythonTerraform
Categories
DevOpsSite Reliability
About Solvd
Solvd is a privately held software engineering and consulting firm that builds AI, data, cloud, application development, and quality engineering solutions for enterprises across technology, ecommerce, retail, fintech, hospitality, and banking. Headquartered in Walnut Creek, California, it is a portfolio company of Siguler Guff & Company. Following its acquisition of Tooploox, Solvd delivers end-to-end engagements from strategic advisory to custom AI development and enterprise-scale implementation for Fortune 500 clients.
