5 months ago
Remote, AmericasStaff+
Responsibilities
- Drive platform architecture decisions and align the engineering team on scalable, maintainable patterns.
- Review code, design documents, and architectural proposals for scalability, reliability, security, and operability.
- Mentor engineers, remove technical blockers, improve production readiness, and establish platform best practices.
- Own and evolve the Django, Django REST Framework, and ASGI backend platform for performance and correctness.
- Scale asynchronous execution across Celery, Dramatiq, and Temporal/Cortex using resilient workflow patterns such as retries, circuit breakers, and graceful degradation.
- Optimize PostgreSQL, pgvector, connection pooling, and caching strategies.
- Maintain and improve Kubernetes deployment infrastructure, including GKE, Helm, Terraform/OpenTofu, CI/CD, rollout strategies, KEDA autoscaling, and worker-pool resource allocation.
- Own the reliability of RabbitMQ, Redis, and PostgreSQL infrastructure, including incident response and post-mortems.
- Extend OpenTelemetry and Datadog instrumentation, dashboards, alerts, and SLOs while reducing latency and memory bottlenecks.
- Identify scaling bottlenecks proactively, improve platform primitives and documentation, and systematically reduce technical debt.
Requirements
- 10+ years building and operating production backend systems at scale.
- Deep expertise in Python and relational databases, preferably Django and PostgreSQL.
- Hands-on experience with Kubernetes, Helm, and cloud infrastructure, preferably GCP.
- Strong background in distributed systems, including message queues, event sourcing, and workflow orchestration.
- Production experience with asynchronous task systems such as Celery, Dramatiq, or similar tools.
- Track record of debugging complex production issues across multiple services.
- Ability to work autonomously and drive technical initiatives without close supervision.
- Clear technical communication skills, including explaining tradeoffs and building consensus.
- Experience with Temporal or similar workflow engines is preferred.
- Background in LLM infrastructure, RAG systems, or AI/ML platforms is preferred.
- Familiarity with OpenTelemetry, Datadog, or similar observability stacks is preferred.
- Experience with KEDA or other Kubernetes autoscaling solutions is preferred.
- Experience with multi-tenant SaaS platform architecture is preferred.
- History of improving developer experience and platform abstractions is preferred.
Benefits
- Remote full-time position
- Requires overlap with Americas time zones
- Reliable high-speed internet required
