14 days ago
Bengaluru, IndiaStaff+
Responsibilities
- Design, develop, and maintain scalable, highly available AI-SRE platform services.
- Own critical system components from design through production operation.
- Define architecture, author functional specifications, and lead technical reviews.
- Diagnose complex distributed-system and production issues involving multi-source event processing.
- Build AI-assisted capabilities for incident detection, diagnosis, remediation, and automation.
- Design REST, gRPC, GraphQL, and event-driven APIs and integrations.
- Define observability strategies using metrics, logs, traces, and actionable alerts.
- Partner with SRE, platform, product, infrastructure, and cross-functional teams during incident investigations.
- Set engineering standards for quality, scalability, security, performance, and reliability.
- Identify technical debt and scaling risks and drive platform improvements.
- Mentor Software Engineers and Senior Software Engineers through design reviews and architectural guidance.
- Influence technical direction across teams and evaluate emerging AI and platform technologies.
Requirements
- 7–10 years of professional software development experience building scalable, distributed applications or platforms.
- Strong experience with Java and deep understanding of distributed systems, concurrency, resiliency, failure handling, data structures, and algorithms.
- Experience designing REST, gRPC, GraphQL, and asynchronous service integrations.
- Hands-on experience with Kubernetes, containers, and cloud-native architectures.
- Experience with observability, incident management, on-call rotations, and resolving complex production issues.
- Proven ability to lead architecture and delivery of production-critical systems without direct authority.
- Experience with AIOps, intelligent observability, generative AI, LLMs, agents, RAG, production AI pipelines, or automated operational runbooks is highly preferred.
- Experience with internal developer platforms, reliability tooling, SLO/SLI frameworks, AI-assisted root-cause analysis, or complete incident lifecycles is preferred.
- Experience with AWS, Azure, or Google Cloud Platform is preferred.
- Customer-facing experience representing technical products to enterprise stakeholders is preferred.
- Bachelor’s degree in Computer Science or a related discipline is preferred; an advanced degree is preferred.
- Equivalent professional experience will also be considered.
Tech Stack
Categories
About Harness
Harness builds a SaaS platform for software delivery, combining CI/CD, feature flags, cloud cost management, reliability, and security tooling for engineering and DevOps teams. The privately held company, founded in 2016 and headquartered in San Francisco, sells to enterprises adopting cloud and Kubernetes. Notable customers include United Airlines, Morningstar, and Choice Hotels, and its offerings span testing, deployments, governance, and automation across the software delivery lifecycle.
