
Platform Observability Engineer
StoneX Group Inc.2 hours ago
Bogotá, ColombiaSenior
Responsibilities
- Design, implement, and operate scalable observability platforms and services centered on OpenSearch, Kubernetes, and Datadog or similar platforms.
- Operate and improve production OpenSearch clusters, including scaling, upgrades, index lifecycle management, and performance and reliability troubleshooting.
- Adopt OpenTelemetry instrumentation standards and drive telemetry by default across services and platforms.
- Build and manage observability data pipelines using tools such as Cribl, Vector, Fluent Bit, and OpenTelemetry Collector.
- Implement telemetry filtering, enrichment, redaction, normalization, retention, tiered storage, and archiving strategies.
- Create actionable alerting, dashboards, runbooks, and analytics to reduce noise and improve mean time to recovery.
- Implement observability infrastructure as code with Terraform and integrate observability into CI and CD workflows.
- Ensure platform stability, performance, security, governance, and cost efficiency through monitoring, tuning, access controls, and incident response.
- Partner with security, compliance, architecture, development, product, and operations teams.
- Create documentation, deliver training, participate in on-call rotations, and contribute to incident response and post-incident reviews.
Requirements
- At least 5 years of experience in observability, SRE, platform, or infrastructure roles.
- Hands-on experience with Datadog or a similar platform covering metrics, logs, APM, tracing, RUM, synthetics, dashboards, alerting, and service catalogs.
- Experience with Prometheus, Grafana, InfluxDB, and production operation of Elasticsearch or OpenSearch clusters, ideally on Kubernetes.
- Experience with index design, query tuning, and index lifecycle policies in Elasticsearch or OpenSearch.
- Practical knowledge of OpenTelemetry concepts and instrumentation patterns.
- Experience with ETL or telemetry pipeline tools such as Cribl Stream, Fluent Bit, Logstash, Kafka, or OpenTelemetry Collector.
- Understanding of data schemas, normalization, enrichment, log routing, sampling, and observability cost optimization.
- Working experience with Terraform or similar infrastructure-as-code tooling and reusable modules.
- Experience with Docker, Kubernetes, Git, Helm, Ansible, and CI and CD pipeline integration.
- Proficiency in Python or Go for APIs, automation, and integrations.
- Strong understanding of Linux fundamentals, networking basics, and security best practices.
- Bachelor’s degree in computer science, engineering, or a related field, or equivalent practical experience.
- Preferred certifications include Datadog, Terraform Associate, CKA, or CKAD.
- Ability to solve complex problems independently and collaborate effectively in a complex platform environment.
Benefits
- Medical and life insurance.
- Public transportation support.
- Meal and food allowances.
- Full-time employment contract.
- Hybrid schedule with four days per week in the office and one day per week remote.
- Role can be based in São Paulo, Brazil, or Bogotá, Colombia.
Tech Stack
AnsibleApache KafkaDatadogDockerElasticsearchGitGoGrafanaHelmInfluxDBKubernetesLinuxLogstashPrometheusPythonTerraform