22 hours ago
Responsibilities
- Design, develop, test, and deploy scalable software applications and platform services primarily using Go.
- Build components that collect, process, aggregate, analyze, and export observability and telemetry data.
- Develop services integrating with Google Cloud Platform observability services, cloud-native monitoring solutions, and internal SRE tooling.
- Create and enhance monitoring and observability capabilities for application health, performance, reliability, and user experience.
- Implement application instrumentation that generates metrics, traces, and logs.
- Design and develop APIs, plugins, connectors, and integrations supporting telemetry collection and observability workflows.
- Collaborate with SRE, Product, Operations, and Engineering teams on observability requirements and scalable solutions.
- Troubleshoot performance, scalability, and reliability issues across cloud-based environments.
- Participate in architecture discussions involving distributed systems, cloud infrastructure, and monitoring technologies.
- Mentor junior engineers through technical guidance, code reviews, and engineering best practices.
- Implement automated testing strategies and deployment pipelines to improve software quality and efficiency.
- Stay current with observability, cloud-native architecture, infrastructure monitoring, and software reliability technologies.
Requirements
- Bachelor's degree in Computer Science, Software Engineering, or a related technical field.
- 5+ years of professional software development experience building scalable applications and services.
- Strong proficiency in Go with experience designing, developing, and maintaining production-grade applications.
- Deep knowledge of Go concurrency, performance optimization, testing, and code quality practices.
- Experience developing cloud-native applications in Google Cloud Platform, AWS, or Azure.
- Working knowledge of cloud networking concepts, VPCs, networking fundamentals, and distributed system architecture.
- Experience building and consuming RESTful APIs and microservices.
- Strong understanding of metrics, traces, logs, telemetry pipelines, monitoring architectures, and observability concepts.
- Experience instrumenting applications with OpenTelemetry or similar observability technologies.
- Experience with telemetry data collection, aggregation, and analysis.
- Familiarity with Grafana, Prometheus, Datadog, New Relic, Dynatrace, Google Cloud Operations Suite, or related monitoring platforms.
- Experience working alongside Site Reliability Engineering teams and supporting production monitoring strategies.
- Strong understanding of SQL and NoSQL database technologies.
- Experience with Agile development methodologies, including Scrum and Kanban.
- Excellent problem-solving, analytical, and communication skills.
- Preferred: experience building internal developer platforms, monitoring systems, or observability products.
- Preferred: familiarity with Google Cloud observability tools, telemetry ecosystems, distributed tracing, service meshes, cloud-native observability architectures, high-scale production SaaS environments, and Terraform.
About NCR
NCR builds enterprise technology for retail, restaurants, and banks, including point-of-sale and self-checkout systems, payments software, ATM networks, and managed services. The company separated in 2023 into two public entities: NCR Voyix (unified commerce and payments for retail and hospitality) and NCR Atleos (banking and ATM-as-a-service). Founded in 1884 and headquartered in Atlanta, it sells hardware, software subscriptions, and support to customers in many countries.
