
Senior Platform Engineer, Cloud Infrastructure
Virtasant Inc.1 month ago
Remote, United StatesSenior
Responsibilities
- Design, build, and operate production Kubernetes clusters across networking, workload isolation, scheduling, resource management, and multi-region topologies.
- Develop and maintain Go, Python, or Java services, controllers, operators, middleware, and internal HTTP, REST, and gRPC interfaces.
- Implement service mesh capabilities including mTLS, workload identity, service-account authentication, authorization, and traffic management.
- Lead platform incident response, root-cause analysis, postmortems, performance troubleshooting, SLO definition, and actionable alerting.
- Own infrastructure as code, CI/CD, and GitOps delivery workflows and support cloud migrations between providers or environments.
- Build observability systems with metrics, dashboards, alerting, distributed tracing, and service instrumentation.
- Partner with product, security, and infrastructure teams on architecture and requirements, contribute to design reviews, mentor engineers, and help set technical direction.
Requirements
- 6+ years of professional experience in software, platform, infrastructure, or site reliability engineering, including significant experience operating production distributed systems.
- Demonstrated experience building and operating production Kubernetes platforms rather than only deploying workloads onto Kubernetes.
- Production experience writing Go, Python, or Java and designing systems from ambiguous starting points through production delivery.
- Experience planning and executing production cloud migrations between providers or environments.
- Degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
- Strong knowledge of Kubernetes internals, networking, NetworkPolicy, resource management, cluster behavior under load, and service mesh technologies such as Istio, Envoy, or Linkerd.
- Hands-on experience with mTLS, workload identity, Linux, cgroups, infrastructure as code, Terraform or an equivalent, and at least one major cloud platform such as GCP, AWS, or Azure.
- Experience with Prometheus, Grafana, OpenTelemetry, PromQL, Docker, relational databases, PostgreSQL or managed PostgreSQL-compatible services, replication, and failover.
- Strong debugging and performance-profiling skills, with preferred experience in Ginkgo, Gomega, Kubernetes controllers and operators, Python, IAM, SSO, Keycloak, OIDC, SAML, Vault, high availability, disaster recovery, monorepos, Bazel, Java, SOC 2, GDPR, Alibaba Cloud, Apache Spark, or Apache Flink.
- Ability to communicate clearly through design documents, postmortems, and cross-team requirements gathering, work independently in a distributed team, and operate effectively in a fast-moving technical environment.
Benefits
- Remote role with Pacific-time coverage from 8:00 AM to 5:00 PM PST.
Tech Stack
Alibaba CloudAmbassadorApache FlinkApache SparkAWSAzureBazelDockerGoGoogle Cloud PlatformGrafanagRPCIstioJavaKubernetesLinuxPostgreSQLPrometheusPythonTerraformVault