Responsibilities
- Manage and optimize AWS compute, networking, storage, IAM policies, and cloud spend using Terraform.
- Operate and troubleshoot Kubernetes clusters, including pod scheduling, resource limits, autoscaling with Karpenter, service mesh with Istio, and container lifecycle management.
- Build monitoring, alerting, dashboards, incident-response processes, root-cause analyses, and runbooks using Prometheus and Grafana.
- Architect and maintain CI/CD pipelines with SAST, dependency scanning, secrets detection, and secure deployments through TrueFoundry.
- Support Databricks environments, GPU workload scheduling, model deployment pipelines, and AI compute cost optimization.
- Automate repetitive operational work and build internal productivity tooling with Python and Bash.
- Drive SOC2 compliance and embed security guardrails across infrastructure and engineering workflows.
Requirements
- 3–5 years of hands-on DevOps or DevSecOps experience in a product company or startup.
- B.Tech, BE, or M.Tech in Computer Science or a related discipline.
- Production experience with AWS, including EC2, VPC, IAM, S3, and CloudWatch.
- Production Kubernetes experience covering deployments, troubleshooting, and scaling.
- Experience with CI/CD pipelines such as GitHub Actions, GitLab CI, or ArgoCD, including SAST and DAST security testing.
- Infrastructure-as-code experience with Terraform or CloudFormation.
- Python and/or Bash scripting experience for automation and tooling.
- Experience setting up monitoring, dashboards, alerts, and runbooks with Prometheus, Grafana, or similar tools.
- Working knowledge of IAM least privilege, secrets management, SOC2, and PCI-DSS concepts.
- Ability to own ambiguous infrastructure projects and drive them to completion.
- Clear communication and documentation skills for runbooks, pull requests, and architecture notes.
- Preferred experience with GPU instances, model serving, TorchServe, Triton, or TrueFoundry.
- Preferred Databricks or Spark experience involving cluster management, job scheduling, and cost controls.
- Preferred SOC2 or PCI-DSS compliance experience.
- Relevant certifications such as AWS Solutions Architect, CKA, AWS Security Specialty, or CompTIA Security+ are a bonus.
Benefits
- Competitive compensation and ESOPs.
- Direct access to senior engineering leadership and short feedback loops.
- Hands-on ownership in a small two-person DevSecOps team.
- Exposure to production-scale real-time voice AI, GPU workloads, model serving, and AI infrastructure cost optimization.
Tech Stack
About Prodigal
Prodigal maximizes payments for lenders and debt collectors by building dynamic strategies and motivating consumers with highly engaging, personalized treatments. Our advanced genAI has been trained on over 400 million consumer finance conversations, delivering unmatched industry expertise so you can drive record recovery rates. Experience the power of intelligent debt resolution with Prodigal’s AI that pays. Prodigal is headquartered in Mountain View, California, and our global team is on a mission to build the intelligence layer that powers consumer finance. With the backing of domain experts, technology leaders, and top investors, including Accel, Menlo Ventures, and Y-Combinator, Prodigal is poised to become the next iconic vertical SaaS company.