
SDE II - AI Platform
Safe Security2 days ago
Bengaluru, IndiaMid Level
Responsibilities
- Build APIs, SDKs, and self-service workflows for deploying, versioning, evaluating, and monitoring AI/ML workloads.
- Design, provision, schedule, isolate, and optimize shared GPU clusters across multiple teams.
- Operate production inference using vLLM, Triton, KServe, Ray, or equivalent systems, including batching, autoscaling, model routing, and multi-model serving.
- Optimize tokens per second, p99 latency, GPU utilization, capacity, and cost per token through quantization and serving improvements.
- Engineer GPU-aware Kubernetes workloads using GPU Operator, device plugins, GPU-aware scheduling, MIG, and NCCL.
- Troubleshoot GPU out-of-memory issues, driver and CUDA mismatches, interconnect bottlenecks, and throughput regressions.
- Build model registries, versioning systems, CI/CD workflows, evaluation harnesses, regression gates, and safe rollout and rollback processes.
- Implement observability for model quality, accuracy drift, latency, throughput, GPU utilization, tenant cost, SLOs, alerting, and incident response.
- Build secure multi-tenant services with authentication, authorization, secrets management, data protection, tenant isolation, and governed access to models and GPUs.
- Use Terraform to codify infrastructure and design for high availability and disaster recovery.
- Contribute to secure multi-tenant microservices and APIs on AWS and own feature delivery, code reviews, mentoring, and architecture direction.
Requirements
- 2–4 years of experience building and operating production software and infrastructure with end-to-end ownership.
- Bachelor’s or Master’s degree in Computer Science, Engineering, or equivalent practical experience.
- Strong Python and/or Go experience developing and reviewing production services.
- Deep Kubernetes and Docker experience, including scheduling, resource management, operators, networking, and debugging under load.
- Strong Linux, networking, storage, and distributed-systems fundamentals.
- Production experience with AWS, Azure, or GCP and Terraform, including cloud services such as Lambda, API Gateway, EC2, S3, and RDS.
- Experience building and operating backend services and APIs for a multi-tenant SaaS product.
- Experience with high availability, scalability, disaster recovery, observability, CI/CD, metrics, traces, SLOs, and automated delivery.
- Hands-on experience with managed AI platforms such as AWS SageMaker, Amazon Bedrock, Google Cloud Vertex AI, or Azure Machine Learning, alongside self-managed infrastructure.
- Experience building or operating AI evaluation systems, including quality measurement, LLM-as-a-judge pipelines, evaluation datasets, and model or prompt regression testing.
- Background in platform engineering, infrastructure, SRE, distributed systems, AI infrastructure, or MLOps/LLMOps.
- Hands-on GPU cluster operations, including provisioning, capacity planning, scheduling, allocation, isolation, utilization optimization, and day-two ownership.
- Experience with NVIDIA technologies and CUDA, including drivers, container runtime, toolkit compatibility, and troubleshooting.
- Production experience with LLM inference and model serving using vLLM, Triton, KServe, Ray, or equivalent, with demonstrated cost or performance improvements.
- Preferred experience includes SageMaker HyperPod, Bedrock, Vertex AI at scale, DCGM, Prometheus, Grafana, MIG, NCCL, FP8, INT8, AWQ, GPTQ, speculative decoding, KV-cache tuning, Langfuse, OpenTelemetry, PyTorch profiling, vector databases, RAG, LLM gateways, agentic systems, TypeScript, Express, React, and human-in-the-loop labeling.
Benefits
- Meaningful equity is provided to employees.
- Unlimited leave is offered.
- Comprehensive medical insurance and wellness benefits are provided.
- The company offers career advancement opportunities in a rapidly growing organization.
Tech Stack
AWSAzureDockerExpressGoGoogle Cloud PlatformGrafanaKubernetesPrometheusPythonPyTorchReactSQLTerraformTypeScript
Categories
About Safe Security
Safe Security builds an AI-driven cyber risk quantification and management platform used by CISOs, GRC, and third-party risk teams to measure and prioritize enterprise, vendor, and AI-related risks. It sells its software to large enterprises as a subscription service, integrating with security and business systems to produce board-level, dollar-based risk insights. Founded in 2012 and headquartered in Palo Alto, it is privately held.