
Staff MaaS Backend Engineer
BitDeer Technologies Group14 days ago
Remote, United States or San Jose, CA, USAStaff+
Responsibilities
- Co-own the end-to-end MaaS system design with the Principal Architect, author technical decision records, and lead the maturity roadmap.
- Implement and maintain Go backend services, APIs, migrations, and independently shippable platform changes.
- Own inference gateway compatibility for OpenAI and Anthropic formats, including streaming, tool calling, structured output, versioning, and deprecation contracts.
- Develop load-aware, model-aware, and prefix-cache-aware routing with circuit breaking, fallback mechanisms, and regional failover.
- Optimize token throughput, KV-cache tiering, time-to-first-token latency, cost per million tokens, and model-serving performance.
- Integrate deployment tooling, LoRA multiplexing, and cold-start-aware autoscaling with Kubernetes GPU infrastructure.
- Design and operate globally distributed regional inference pools and an active-active control plane with defined SLOs.
- Run peak-concurrency load tests, maintain on-call runbooks, and implement predictable load shedding and incident response.
- Enforce fail-closed authorization, multi-tenant isolation, zero-retention data paths, API-key and OAuth-credential lifecycle management, quotas, and abuse controls.
- Build exactly-once token metering, high-volume usage ledgers, asynchronous prepaid spend caps, and accurate invoice reconciliation.
- Execute incremental zero-downtime migrations using strangler-style replacements, shadow traffic, and stateful dual writes.
- Deliver end-to-end request tracing and cost telemetry while preventing unauthorized retention of prompts, completions, and PII.
- Set Go and API engineering standards, mentor engineers, and align cross-functional technical execution.
Requirements
- 8+ years of backend engineering experience.
- 3+ years owning a high-traffic, multi-tenant API platform serving paying customers.
- Proven experience scaling distributed systems through multi-region active-active deployments, caching, backpressure, and performance engineering.
- Demonstrated success executing zero-downtime brownfield migrations for stateful subsystems such as metering or ledgers without regressions or accounting gaps.
- Deep hands-on proficiency with Go-based services and production Kubernetes, including Envoy and GPU-aware scheduling.
- Strong architectural experience with PostgreSQL, Redis, and Kafka.
- Experience owning end-to-end observability strategies using OpenTelemetry and high-cardinality analytics stores.
- Systems-level understanding of LLM serving, server-sent-event streaming, KV and prefix caching, and time-to-first-token versus throughput trade-offs.
- Experience building reliable exactly-once metering and billing systems that reconcile large event volumes under partial failure conditions.
- Expertise designing fail-closed authorization, verified identities, and strict cross-tenant isolation.
- Ability to author rigorous design documents, make architectural trade-offs, collaborate with architects, and execute decisions effectively.
- Willingness to carry on-call responsibilities, lead blameless incident reviews, and convert outages into structural improvements.