2 months ago
Houston, TX, USA or San Francisco, CA, USASenior
Responsibilities
- Architect, deploy, and operate workflow orchestration platforms supporting simulation, machine learning and model training, data pipelines, CI/CD, and general-purpose workloads.
- Build internal platforms, abstractions, SDKs, and self-service tooling on top of orchestration engines.
- Operate workflow platforms at scale on Kubernetes across AWS and on-premises data centers, including scheduling, autoscaling, GPU and heterogeneous resources, and cross-cluster orchestration.
- Improve workload reliability, performance, and cost efficiency through observability, queuing, prioritization, retries, and resource optimization.
- Partner with ML, simulation, data, and infrastructure teams to deliver fit-for-purpose pipelines.
- Integrate workflow platforms with storage, data streaming and event systems, artifact and model registries, and CI/CD tooling.
- Establish workflow authoring and operations best practices, templates, and documentation, and mentor engineers.
- Respond to user-impacting issues, provide short-term mitigation, and implement durable long-term solutions.
Requirements
- Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent experience.
- 5+ years of hands-on experience in platform engineering, infrastructure, DevOps, or SRE roles.
- Significant experience with workflow orchestration platforms such as Argo Workflows, Airflow, or comparable systems.
- Strong software development skills in one or more of Python, Go, Java, or JavaScript/TypeScript.
- Solid understanding of Kubernetes and distributed systems.
- Preferred: expert-level experience with Argo Workflows, Airflow, Prefect, Dagster, Temporal, Kubeflow Pipelines, or Flyte.
- Preferred: experience orchestrating ML training, simulation, or large-scale data and batch workloads, including GPU scheduling.
- Preferred: in-depth Kubernetes experience with EKS, GKE, AKS, RKE2, or Rancher, including cross-cluster orchestration.
- Preferred: proficiency with Terraform, Pulumi, OpenTofu, or Ansible.
- Preferred: experience with NATS JetStream, Kafka, Pulsar, or RabbitMQ.
- Preferred: familiarity with Prometheus, Grafana, Loki, OpenTelemetry, or comparable observability stacks.
- Preferred: demonstrated ability to optimize workload cost and performance without compromising reliability.
Tech Stack
AnsibleApache AirflowApache KafkaAWSGoGrafanaJavaJavaScriptKubernetesPrometheusPythonRabbitMQRancherTerraformTypeScript
Categories
About Bot Auto
Bot Auto develops and operates Level 4 autonomous trucks, selling freight capacity as Transportation-as-a-Service to shippers rather than licensing technology. The Houston-based, privately held company (founded 2023) owns the trucks and runs end-to-end operations, pairing fleet operators with autonomy engineers to deploy long-haul routes. Its platform spans autonomy software, vehicle controls, sensors, compute, and fleet operations with a documented safety case for commercial freight service.
