
Future openings - SRE - Compute (Linux, Kubernetes, Containers)
Virtasant Inc.12 months ago
Remote, United StatesMid Level
Responsibilities
- Support internal developers through onboarding, community forums, Slack, ticketing systems, and deep technical troubleshooting.
- Investigate and resolve Kubernetes workload, scheduling, service, networking, and RBAC issues.
- Troubleshoot Docker containers, images, builds, runtime environments, and Linux server behavior.
- Diagnose TCP/IP, DNS, routing, firewall, filesystem, systemd, and resource-constraint problems using Linux tools.
- Write technical tickets, incident summaries, runbooks, and knowledge-base articles.
- Participate in postmortems and drive follow-up actions.
- Collaborate with SRE, Platform, and Engineering teams to escalate issues, identify platform bugs, and propose improvements.
- Build tools, scripts, and processes that improve support efficiency.
Requirements
- Strong Linux fundamentals, including processes, services, filesystems, and kernel basics.
- Strong Linux networking knowledge, including TCP/IP, routing, DNS, load balancers, and firewalls.
- Hands-on Kubernetes experience, preferably beyond managed cloud platforms.
- Strong Docker and containerization expertise.
- Experience providing production support for large-scale systems.
- Ability to write and understand Python scripts or similar languages.
- Comfort reviewing pull requests, reading developer code, and communicating with development teams.
- Strong Git proficiency.
- Exceptional written and spoken English.
- Strong organization, attention to detail, follow-through, problem-solving, analytical thinking, and structured debugging skills.
- Ability to work independently with distributed teams.
- Bachelor's degree in Computer Science or a similar engineering discipline is preferred.
- Experience with private cloud or internal platform engineering teams and familiarity with Spark, Kafka, or distributed systems are preferred.
Benefits
- Remote-first role open across Brazil, Chile, Colombia, and Mexico, aligned to 8 AM–5 PM Pacific time.
- Work from anywhere with autonomy and flexibility.
- Global collaboration, continuous learning, and exposure to cutting-edge systems across clients and sectors.
Categories
Site Reliability