Responsibilities
- Manage and optimize AWS compute, networking, storage, cost monitoring, rightsizing, and cloud spend.
- Operate and troubleshoot Kubernetes clusters, including pod scheduling, resource limits, autoscaling with Karpenter, Istio service mesh, and container lifecycle management.
- Build monitoring, alerting, dashboards, runbooks, and incident-response processes using Prometheus and Grafana.
- Architect and maintain CI/CD pipelines with SAST, DAST, dependency scanning, secrets detection, and secure deployments through TrueFoundry.
- Manage AI/ML infrastructure, including Databricks environments, GPU workload scheduling, model deployment pipelines, and AI compute-cost optimization.
- Automate repetitive operational work and build internal tooling with Python and Bash.
- Own IAM policies, AWS hardening, security guardrails, and SOC2 compliance initiatives.
Requirements
- 3–5 years of hands-on DevOps or DevSecOps experience in a product company or startup.
- B.Tech, BE, or M.Tech in Computer Science or a related discipline.
- Production AWS experience with EC2, VPC, IAM, S3, and CloudWatch, including daily console and CLI use.
- Production Kubernetes experience covering deployment management, troubleshooting, and scaling.
- Experience with CI/CD systems such as GitHub Actions, GitLab CI, or ArgoCD, including integrating automated security testing.
- Infrastructure-as-code experience with Terraform or CloudFormation.
- Python and/or Bash scripting experience for automation and tooling.
- Experience setting up monitoring, dashboards, alerts, and runbooks with Prometheus, Grafana, or similar tools.
- Working knowledge of IAM least privilege, secrets management, SOC2, and PCI-DSS concepts.
- Ability to independently own ambiguous infrastructure projects and communicate through readable runbooks, pull requests, and architecture documentation.
- Preferred experience includes GPU instances, model serving with TorchServe, Triton, or TrueFoundry, inference optimization, Databricks, Spark, SOC2 or PCI-DSS audits, and certifications such as AWS Solutions Architect, CKA, AWS Security Specialty, or CompTIA Security+.
Benefits
- Competitive compensation and ESOPs are offered.
- Work on real-time voice AI infrastructure with sub-one-second latency pipelines and production-scale GPU workloads.
- Gain exposure to AI/ML infrastructure, model serving, and inference-cost optimization.
- Join a two-person DevSecOps team with direct ownership and access to senior engineering leadership.
- Work in a fast-paced, entrepreneurial environment at a YC-, Accel-, and Menlo Ventures-backed company.
Tech Stack
About Prodigal
Prodigal maximizes payments for lenders and debt collectors by building dynamic strategies and motivating consumers with highly engaging, personalized treatments. Our advanced genAI has been trained on over 400 million consumer finance conversations, delivering unmatched industry expertise so you can drive record recovery rates. Experience the power of intelligent debt resolution with Prodigal’s AI that pays. Prodigal is headquartered in Mountain View, California, and our global team is on a mission to build the intelligence layer that powers consumer finance. With the backing of domain experts, technology leaders, and top investors, including Accel, Menlo Ventures, and Y-Combinator, Prodigal is poised to become the next iconic vertical SaaS company.
