about 3 hours ago
Base Salary
$30k - $50k/yr
Responsibilities
- Own the infrastructure and runtime environment for OpenClaw agents.
- Build, configure, deploy, and troubleshoot agents on Linux-based virtual machines.
- Develop Python services, automation, and operational tooling.
- Design and operate agent harnesses for memory and context management.
- Create a standardized platform for safe agent development and deployment.
- Deploy and manage agent workloads on Kubernetes and cloud infrastructure.
- Investigate system behavior using Bash and Linux tools.
- Build reliable systems for agent scheduling and task management.
- Develop integrations for agents to interact with internal services and tools.
- Establish secure approaches for managing credentials and access control.
- Implement observability for agent performance and resource consumption.
- Create dashboards and incident-response processes for platform health.
- Diagnose failures across various system components.
- Improve platform efficiency and resilience as agent usage grows.
- Collaborate with product engineers to transition prototypes to production.
- Establish best practices for operating autonomous agents.
Requirements
- Professional experience in building and operating production software or infrastructure.
- Strong proficiency in Python for agent development and automation.
- Excellent familiarity with Linux environments and debugging software.
- Strong Bash and command-line skills for process management and diagnostics.
- Experience managing software across fleets of virtual machines.
- Hands-on experience with Google Cloud Platform or Amazon Web Services.
- Experience deploying and managing Linux virtual machines and containerized workloads.
- Hands-on experience with Kubernetes and Docker.
- Experience with CI/CD and infrastructure as code.
- Experience with large language models and autonomous agents.
- Understanding of agent concepts like memory and context management.
- Experience monitoring and debugging distributed systems.
- Strong understanding of production reliability and incident response.
- Ability to balance experimentation with security and reliability.
- Strong ownership and communication skills across teams.
- A builder mindset for creating evolving software infrastructure.
