2 months ago
Pune, IndiaSenior
Responsibilities
- Design, build, and maintain scalable, cost-efficient AWS infrastructure with Terraform, including PCI-scoped network segmentation.
- Own and evolve production observability using SigNoz, Grafana, structured logging, and related monitoring practices.
- Build and maintain Bitbucket Pipelines CI/CD workflows with deployment guardrails, rollback strategies, and environment promotion.
- Design and run performance and load tests with k6 or comparable tools and address bottlenecks before production.
- Participate in on-call rotations, lead infrastructure incident response, and drive postmortems and durable fixes.
- Define and track SLOs, error budgets, and operational KPIs for payment-critical services.
- Improve deployment velocity and safety through faster pipelines, preview environments, and safer migration tooling.
- Partner with security and compliance teams to maintain PCI DSS posture and support audits.
- Manage AWS cost and capacity through right-sizing, reserved capacity strategy, and optimization efforts.
- Use AI-assisted tools to accelerate infrastructure work while maintaining review quality and security.
Requirements
- 5+ years of professional DevOps, SRE, or Platform Engineering experience.
- Computer Science or Engineering degree or equivalent experience.
- Strong hands-on AWS experience across ECS, Lambda, VPC, security groups, RDS/Aurora, SQS, and EventBridge.
- Production Terraform experience including module design, state management, drift handling, and multi-environment patterns.
- Experience operating production observability stacks and familiarity with SigNoz, Grafana, Datadog, New Relic, or Prometheus.
- Experience designing and running load tests with k6, JMeter, Gatling, Locust, or comparable tools.
- Experience building and maintaining CI/CD pipelines with Bitbucket Pipelines, GitHub Actions, CircleCI, or similar tools.
- Strong shell scripting ability or comparable scripting experience with Python or Go.
- Ability to understand backend and frontend code to implement infrastructure effectively.
- Understanding of networking fundamentals, access control, identity management, IAM, OIDC, Cognito, and OAuth.
- Comfort leading incident response and writing postmortems that drive systemic improvement.
- Fluency with AI-assisted development tools such as Claude Code, Cursor, or Copilot, with a clear perspective on their appropriate use.
- Preferred experience includes PCI-compliant or regulated environments, payments or high-availability transactional systems, Vanta, PostgreSQL tuning, KMS-based cryptography, secrets management, key rotation, chaos engineering, or game-day exercises.
- Must be legally authorized to work in the country of application.
Benefits
- 100% remote work with flexible working arrangements.
- Paid parental leave benefit programs.
- Three extra #GiveBackDays for volunteering and community impact.
- Diversity and inclusion initiatives including a D&I Council and Global Mentorship Program.
- Access to free mental health support.
- The company supports a global network of colleagues and ongoing professional growth.
