
Staff Platform Engineer - Developer Infrastructure
Persona AI Inc6 hours ago
Responsibilities
- Own the build graph and CI capacity for a mixed C++/Python/Rust/CUDA monorepo, including incremental correctness, remote caching, cross-compilation, and self-hosted or hardware-attached runners.
- Ensure reproducible builds that produce bit-identical artifacts from tagged commits over time.
- Manage versioning, artifact promotion, container registries, package mirrors, signing, and provenance.
- Design safe, resumable, bandwidth-aware robot rollouts using staged channels, canary robots, and reliable rollback.
- Build reproducible development environments and self-service tooling for engineers to test and deploy software to robots without requiring Kubernetes expertise.
- Improve onboarding so new engineers can build, simulate, and deploy to a bench robot on their first day.
- Operate cloud and on-prem compute, multi-terabyte robot-log storage, and the operational layer of training and simulation clusters.
- Provide fleet observability through metrics, logs, and traces, and maintain VPN, overlay networking, local mirrors, and offline-capable registry authorization for deployment sites.
Requirements
- At least 8 years of experience operating production infrastructure for a software engineering organization, with direct ownership of CI/CD or developer-platform work.
- Deep Linux systems fluency, including networking, storage, systemd, kernel debugging, and driver debugging.
- Hands-on ownership of a major CI system such as GitHub Actions, GitLab CI, Buildkite, or Jenkins, with understanding of underlying build and cache layers.
- Experience with a C/C++ build system at scale, such as Bazel or CMake, including cross-compilation and dependency pinning.
- Experience with infrastructure as code using Ansible or Terraform and reviewable, reproducible, version-controlled changes.
- Fluency in Python and Bash, with an emphasis on automating manual procedures.
- Strong written communication skills.
- Preferred experience shipping software to embedded or edge Linux targets, including ARM64, Jetson, Yocto or custom images, A/B partitions, and OTA update systems.
- Preferred experience with hybrid on-premises and cloud environments, open-source observability using Prometheus, Grafana, or OpenTelemetry, and self-hosted registries, artifact stores, or object storage.
- Robotics, autonomous vehicle, aerospace, or other hardware-heavy experience is a bonus, as are Nix and Bazel remote execution.
Benefits
- Competitive compensation, performance-based bonus, early-stage equity, competitive PTO, and a company-wide paid winter break from December 24 through January 2.
- Medical benefits are 99% employer covered.
- Full-time role located in Houston, Texas, with 10% travel.