Site Reliability Engineer
Tookitaki Holding PTE LTD1 month ago
Remote, PhilippinesMid Level / Senior
Responsibilities
- Build and maintain monitoring, alerting, and logging systems.
- Respond to incidents and outages, conduct post-mortems, and implement corrective actions.
- Automate infrastructure provisioning and application deployment using Terraform, Ansible, or Helm.
- Contribute to CI/CD pipelines and improve software delivery reliability and speed.
- Manage and troubleshoot Docker containers and Kubernetes clusters, including scaling, resource management, and workload health.
- Operate AWS or GCP environments and monitor system availability and resource usage.
- Implement and monitor TLS/SSL, RBAC, SSO, and secure API practices.
- Support compliance and security audits through logging, access controls, and operational hygiene.
- Collaborate with developers, infrastructure engineers, client success, and support teams on production readiness.
- Maintain playbooks, runbooks, and system documentation.
Requirements
- Bachelor’s degree in Computer Science, Engineering, or a related technical field.
- 3–6 years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or a related role.
- Experience with production environments and live system debugging.
- Experience deploying and scaling services with Kubernetes, Docker, and Helm.
- Linux administration and command-line debugging experience.
- Hands-on experience with AWS or GCP cloud platforms.
- Bash and Python scripting experience for automation and monitoring.
- Experience with Prometheus, Grafana, ELK, or Datadog.
- Familiarity with databases such as MariaDB and ScyllaDB and with SQL or CQL querying.
- Ability to participate in on-call rotations and work in high-pressure production environments.
- Strong problem-solving, debugging, communication, and documentation skills.
Benefits
- Competitive compensation.
- Work on a globally recognized RegTech platform focused on financial crime prevention.
- Exposure to AI and big data infrastructure, including Spark, Kafka, ScyllaDB, and Flink.
- Location: Manila.
Tech Stack
AnsibleApache FlinkApache KafkaApache SparkAWSBashDatadogDockerGitLab CI/CDGoogle Cloud PlatformGrafanaHelmJenkinsKubernetesMariaDBPrometheusPythonSQLTerraform
Categories
Site Reliability