
Future openings INDIA - SRE - Compute (Linux, Kubernetes, Containers)
Virtasant Inc.5 months ago
Remote, United StatesMid Level
Responsibilities
- Provide high-touch onboarding and technical support to internal customers using the compute platform.
- Respond through Slack, community forums, and ticketing systems, serving as the first engineering support line for Kubernetes and container issues.
- Troubleshoot Kubernetes workloads, scheduling, services, networking, and RBAC.
- Investigate Docker containers, images, builds, runtime environments, and Linux server issues.
- Diagnose Linux processes, services, filesystems, resource constraints, and networking problems.
- Use tools including tcpdump, ss, strace, traceroute, dig, journalctl, top, and htop for system investigation.
- Write technical tickets, incident summaries, runbooks, knowledge-base articles, and documentation improvements.
- Participate in postmortems and drive follow-up actions.
- Collaborate with SRE, Platform, and Engineering teams to escalate issues, identify platform bugs, and propose improvements.
- Provide structured feedback based on support trends and contribute tools, scripts, and processes that improve support efficiency.
Requirements
- Strong Linux fundamentals covering processes, services, filesystems, and kernel basics.
- Strong Linux networking knowledge, including TCP/IP, routing, DNS, load balancers, and firewalls.
- Hands-on Kubernetes experience, preferably beyond managed cloud platforms.
- Strong Docker and containerization expertise.
- Experience providing production support for large-scale systems.
- Ability to write and understand Python scripts or similar languages.
- Ability to review pull requests, read developer code, and communicate with development teams.
- Strong Git proficiency.
- Exceptional written and spoken English with clear technical communication.
- Strong organization, attention to detail, follow-through, problem-solving, analytical thinking, and structured debugging skills.
- Ability to work independently with distributed teams.
- Bachelor’s degree in Computer Science or a similar engineering discipline is preferred.
- Experience supporting private cloud or internal platform engineering teams is preferred.
- Familiarity with Spark, Kafka, or distributed systems is preferred.
Benefits
- Remote-first work from anywhere with autonomy and respect for working time.
- Collaboration with a global team across more than 130 countries.
- Continuous learning and exposure to systems across clients and sectors.
- Trust-based, diverse, and globally distributed work environment.
- Opportunity to solve technically complex problems affecting systems used by hundreds of millions of people.
Categories
Site Reliability