1 day ago
Base Salary
$187k - $334k/yr
Responsibilities
- Design, build, enhance, deploy, maintain, and operate critical distributed messaging, streaming, caching, and NoSQL services.
- Develop core software modules, APIs, client libraries, SDKs, and cloud-native controllers for distributed systems.
- Create observability modules, alerts, automation, and dashboard lifecycle management capabilities.
- Deploy and operate distributed infrastructure across Kubernetes, OpenStack, bare metal, AWS, GCP, and private cloud environments.
- Champion resiliency, high availability, fault tolerance, replication, sharding, and operational excellence for streaming, messaging, and caching platforms.
- Evaluate and implement open-source and cloud-native tools and technologies.
- Participate in the on-call rotation and optimize distributed services in production.
- Lead architectural direction and collaboration across distributed software development teams, including technical writing and presentations to senior leaders and architects.
Requirements
- 12+ years of experience in software development engineering.
- 6+ years focused on designing, building, and operating distributed systems such as Redis, Kafka, RabbitMQ, or NoSQL solutions.
- 5+ years designing and implementing complex distributed-system architectures with high availability and fault tolerance.
- 8+ years of experience with at least two of Java, Python, Go, or C/C++, including production-level distributed-systems code.
- Expertise with Chef configuration management and Kubernetes service deployment using Helm and ArgoCD.
- Bachelor’s degree in Computer Science, Engineering, or a related discipline, or equivalent practical experience; a master’s degree is strongly preferred.
- Expertise in algorithmic thinking, including CAP theorem, queuing theory, and consensus protocols.
- Deep experience with API and client-library development, including RESP and Kafka wire protocols.
- Experience with cloud-native controllers, with familiarity with operators such as Strimzi preferred.
- Experience with consistency and linearizability testing, chaos testing, and fault-injection strategies.
- Deep knowledge of distributed-systems principles, replication, sharding, high availability, and global replication with 99.99% SLOs.
- Experience with large-scale data-processing technologies such as Kafka, Redis, RabbitMQ, Spark, and Flink.
- Strong understanding of object-oriented design, architectural patterns, system security, authentication, and authorization.
- Extensive experience with source-control and CI/CD tools such as Git, Jenkins, and Harness.
- Ability to inspect complex open-source service internals, lead cross-team collaboration, drive architecture, and communicate through technical documentation and presentations.
Benefits
- Flexible work arrangement combining in-person and remote work, with at least 50% of each quarter spent in the office or in the field.
- Role may be eligible for the Workday Bonus Plan or a role-specific commission/bonus and annual refresh stock grants.
- Workday offers comprehensive benefits and reasonable accommodations during the application process.
Tech Stack
Apache FlinkApache KafkaApache SparkAWSCC++ChefGitGoGoogle Cloud PlatformHarnessHelmJavaJenkinsKubernetesOpenStackPythonRabbitMQRedis
About Workday
Workday builds cloud-based enterprise applications for human capital management and financial management, including payroll, time tracking, expenses, planning, and procurement, sold on a subscription basis with professional services. Founded in 2005 and headquartered in Pleasanton, California, it is a public company traded on NASDAQ as WDAY. Organizations worldwide use Workday to unify HR and finance data, automate processes, and apply AI to workforce and financial operations.
