Safe Security

SDE II - AI Platform

Safe Security
Apply
2 days ago
Bengaluru, IndiaMid Level

Responsibilities

  • Build APIs, SDKs, and self-service workflows for deploying, versioning, evaluating, and monitoring AI/ML workloads.
  • Design, provision, schedule, isolate, and optimize shared GPU clusters across multiple teams.
  • Operate production inference using vLLM, Triton, KServe, Ray, or equivalent systems, including batching, autoscaling, model routing, and multi-model serving.
  • Optimize tokens per second, p99 latency, GPU utilization, capacity, and cost per token through quantization and serving improvements.
  • Engineer GPU-aware Kubernetes workloads using GPU Operator, device plugins, GPU-aware scheduling, MIG, and NCCL.
  • Troubleshoot GPU out-of-memory issues, driver and CUDA mismatches, interconnect bottlenecks, and throughput regressions.
  • Build model registries, versioning systems, CI/CD workflows, evaluation harnesses, regression gates, and safe rollout and rollback processes.
  • Implement observability for model quality, accuracy drift, latency, throughput, GPU utilization, tenant cost, SLOs, alerting, and incident response.
  • Build secure multi-tenant services with authentication, authorization, secrets management, data protection, tenant isolation, and governed access to models and GPUs.
  • Use Terraform to codify infrastructure and design for high availability and disaster recovery.
  • Contribute to secure multi-tenant microservices and APIs on AWS and own feature delivery, code reviews, mentoring, and architecture direction.

Requirements

  • 2–4 years of experience building and operating production software and infrastructure with end-to-end ownership.
  • Bachelor’s or Master’s degree in Computer Science, Engineering, or equivalent practical experience.
  • Strong Python and/or Go experience developing and reviewing production services.
  • Deep Kubernetes and Docker experience, including scheduling, resource management, operators, networking, and debugging under load.
  • Strong Linux, networking, storage, and distributed-systems fundamentals.
  • Production experience with AWS, Azure, or GCP and Terraform, including cloud services such as Lambda, API Gateway, EC2, S3, and RDS.
  • Experience building and operating backend services and APIs for a multi-tenant SaaS product.
  • Experience with high availability, scalability, disaster recovery, observability, CI/CD, metrics, traces, SLOs, and automated delivery.
  • Hands-on experience with managed AI platforms such as AWS SageMaker, Amazon Bedrock, Google Cloud Vertex AI, or Azure Machine Learning, alongside self-managed infrastructure.
  • Experience building or operating AI evaluation systems, including quality measurement, LLM-as-a-judge pipelines, evaluation datasets, and model or prompt regression testing.
  • Background in platform engineering, infrastructure, SRE, distributed systems, AI infrastructure, or MLOps/LLMOps.
  • Hands-on GPU cluster operations, including provisioning, capacity planning, scheduling, allocation, isolation, utilization optimization, and day-two ownership.
  • Experience with NVIDIA technologies and CUDA, including drivers, container runtime, toolkit compatibility, and troubleshooting.
  • Production experience with LLM inference and model serving using vLLM, Triton, KServe, Ray, or equivalent, with demonstrated cost or performance improvements.
  • Preferred experience includes SageMaker HyperPod, Bedrock, Vertex AI at scale, DCGM, Prometheus, Grafana, MIG, NCCL, FP8, INT8, AWQ, GPTQ, speculative decoding, KV-cache tuning, Langfuse, OpenTelemetry, PyTorch profiling, vector databases, RAG, LLM gateways, agentic systems, TypeScript, Express, React, and human-in-the-loop labeling.

Benefits

  • Meaningful equity is provided to employees.
  • Unlimited leave is offered.
  • Comprehensive medical insurance and wellness benefits are provided.
  • The company offers career advancement opportunities in a rapidly growing organization.
Safe Security

About Safe Security

1,001-5,000 employees

Safe Security builds an AI-driven cyber risk quantification and management platform used by CISOs, GRC, and third-party risk teams to measure and prioritize enterprise, vendor, and AI-related risks. It sells its software to large enterprises as a subscription service, integrating with security and business systems to produce board-level, dollar-based risk insights. Founded in 2012 and headquartered in Palo Alto, it is privately held.

Contact me