Senior Backend / Distributed Systems Engineer
Aivar Innovations Private Limited15 days ago
Bengaluru, IndiaSenior
Responsibilities
- Own the architecture and implementation of core control-plane backend services.
- Design APIs and domain services for platform entities including organizations, projects, clusters, resources, workloads, models, experiments, webhooks, credentials, events, and alerts.
- Model long-running infrastructure operations as durable, recoverable state machines with explicit intermediate, retry, failure, cancellation, and terminal states.
- Build asynchronous jobs, queues, retries, idempotency, deduplication, leases, backoff, cancellation, timeout handling, and failure recovery.
- Design versionable APIs with resource semantics, filtering, pagination, validation, error models, concurrency behavior, and compatibility guarantees.
- Define domain ownership and consistency boundaries across PostgreSQL, Kubernetes, caches, derived state, and authoritative systems.
- Implement multi-tenant authorization with organization-, tenant-, project-, and resource-scoped permissions and auditable enforcement.
- Design event, notification, alert, audit, webhook, and cluster-agent communication infrastructure.
- Design for high availability, failover, duplicate delivery, leader election, database failure, stale cache state, and process restarts.
- Own schema design, indexing, migrations, transactions, optimistic concurrency, query performance, retention, archival, and operational safety.
- Build observability with structured logs, metrics, traces, request correlation, job execution traces, and auditability.
- Review architecture, mentor engineers, and establish backend standards for correctness, testing, operability, and maintainability.
- Partner with frontend and Kubernetes engineers to evolve product UX, APIs, controllers, and persistence models together.
Requirements
- At least 5 years building production backend or distributed systems, preferably in infrastructure platforms, developer platforms, control planes, enterprise SaaS, or other stateful systems.
- Alternatively, 4+ years building product services or enterprise APIs.
- Strong Go experience, including concurrency, interfaces, context propagation, testing, profiling, networking, HTTP services, and production-quality error handling.
- Strong distributed-systems fundamentals, including idempotency, retries, delivery semantics, eventual consistency, leases, leader election, backpressure, ordering, and partial failure.
- Strong PostgreSQL and relational database fundamentals, including schema design, transactions, locking, indexing, migrations, query planning, and data lifecycle.
- Experience designing APIs used by other engineers, including resource modeling, compatibility, versioning, pagination, filtering, authorization, error contracts, and SDK or client usability.
- Experience with asynchronous processing such as queues, workers, job systems, event buses, workflow engines, or durable execution patterns.
- Knowledge of authentication, authorization, service identities, secrets, API keys, signing, tenant isolation, audit logs, and least privilege.
- Ability to model long-running infrastructure actions as explicit state machines.
- Strong operational instincts and the ability to make production failures diagnosable.
- Working knowledge of Kubernetes controllers, resources, reconciliation, watches, namespaces, RBAC, and control-plane versus cluster state.
- Fluency with agentic coding tools for implementation, testing, investigation, and code generation.
- Preferred experience includes control-plane or infrastructure products, multi-cluster systems, PostgreSQL at scale, event or workflow infrastructure, Kubernetes client libraries, distributed tracing, webhook or audit architectures, and enterprise IAM integrations.