
Observability Engineer - Site Reliability Engineering (SRE)
Oversea-Chinese Banking Corporation Limited2 hours ago
Singapore, SingaporeMid Level / Senior
Responsibilities
- Design, implement, and maintain observability solutions covering metrics, logs, and distributed traces across Kubernetes and OpenShift.
- Instrument applications and infrastructure with OpenTelemetry and drive observability adoption across engineering teams.
- Build dashboards, alerting rules, SLI/SLO frameworks, observability data pipelines, and self-healing or auto-remediation workflows.
- Develop Python-based internal tooling, automation utilities, REST API integrations, and AWS EventBridge event-driven workflows.
- Create and maintain Ansible playbooks for configuration management, agent deployment, and infrastructure provisioning.
- Monitor and manage Kubernetes and OpenShift workloads, namespaces, resource usage, pod health, and container tracing across SIT, UAT, PROD, and DR environments.
- Collaborate with DevOps and application engineering teams to embed observability checks and quality gates into CI/CD pipelines.
- Maintain automation code in Bitbucket, manage delivery through Jira, and integrate checks into Jenkins pipelines.
- Participate in incident response, post-incident reviews, operational runbook improvement, and cluster onboarding.
Requirements
- Bachelor’s degree in Computer Science, Information Technology, or a related engineering discipline.
- 3–6 years of experience in an SRE, Platform Engineering, or DevOps role with a strong observability focus.
- Hands-on experience operating Kubernetes and/or Red Hat OpenShift clusters, namespaces, workloads, and resources.
- Production-quality Python scripting, automation, and API integration experience.
- Experience with OpenTelemetry instrumentation, collectors, exporters, and logs, metrics, and traces.
- Experience building or consuming RESTful APIs, using AWS services including AWS EventBridge, and working with Ansible.
- Working knowledge of Bitbucket, Git, Jira, and Jenkins.
- Relevant certifications such as CKA/CKAD, Red Hat OpenShift, or AWS Solutions Architect are advantageous.
- Experience with Elastic Stack, Moogsoft, Dynatrace, Prometheus, Grafana, Istio, SLI/SLO frameworks, cloud-native CI/CD, GitOps, distributed tracing, or regulated financial-services environments is advantageous.
- Comfort working with compliance, audit, and governance frameworks.
Benefits
- Competitive base salary.
- Holistic and flexible benefits.
- Community initiatives.
- Industry-leading learning and professional development opportunities.
- High-impact work supporting critical financial-services infrastructure, with opportunities to grow into senior SRE, platform architecture, or engineering tracks.
Tech Stack
Categories
DevOpsSite Reliability