
Observability Lead-Sr. Infrastructure Engineer-ON-SITE REQUIRED IN LISTED TRUIST HUBS
Truist Financial Corporation2 hours ago
Charlotte, NC, USAStaff+
Responsibilities
- Define and implement enterprise observability strategies covering metrics, logs, traces, telemetry pipelines, event streaming, and developer-enabled observability.
- Design, build, manage, and implement infrastructure technology platforms across cloud, network, database, storage, platform, computing, and middleware domains.
- Develop automation, monitoring, optimization techniques, APIs, integrations, collectors, dashboards, and reusable operational tooling.
- Instrument and troubleshoot application code, identify performance bottlenecks, and guide improvements to reliability and performance.
- Design and support Kafka and event-driven telemetry flows, including producers, consumers, topics, partitions, consumer groups, and consumer-lag analysis.
- Collaborate directly with engineering, SRE, platform teams, software developers, business stakeholders, and external partners as a forward-deployed technical partner.
- Provide technical guidance, training, and direction to infrastructure teams and lower-level technical professionals.
- Lead infrastructure projects and programs, review designs and configurations, and ensure adherence to technology standards, governance, and regulatory requirements.
- Translate customer implementations and field learnings into reusable code assets, platform capabilities, enterprise standards, and strategic recommendations.
Requirements
- Bachelor’s degree in Computer Science, Engineering, Information Systems, or a related field.
- At least 7 years of professional infrastructure engineering experience.
- Advanced knowledge of enterprise cloud, network, database, storage, platform, computing, and middleware technologies.
- Hands-on OpenTelemetry experience, including application instrumentation, collector configuration, semantic conventions, and telemetry pipeline design.
- Experience with observability platforms such as Prometheus, Grafana, Jaeger, Elastic, Splunk, Dynatrace, or Datadog.
- Strong software development experience in one or more of Python, Go, Java, JavaScript/TypeScript, .NET/C#, Bash, or PowerShell.
- Experience designing APIs, automation frameworks, command-line utilities, integrations, collectors, dashboards, or reusable internal tools.
- Strong Kafka or event-streaming experience, including Kafka Connect, event-driven integration, consumer-lag analysis, throughput troubleshooting, and monitoring Kafka-based services.
- Experience with Git-based source control, code review, CI/CD pipelines, automated testing, release practices, secure coding, and developer enablement.
- Strong background in Kubernetes, cloud platforms, containers, microservices, service-oriented architectures, event-driven architectures, and cloud-native application patterns.
- Experience with infrastructure as code or configuration automation tools such as Terraform, Ansible, or Helm.
- Experience defining observability strategies, technical standards, developer enablement patterns, event-streaming observability patterns, and engineering roadmaps.
- Experience in forward-deployed engineering, solutions engineering, developer advocacy, or customer-facing technical leadership.
- Preferred qualifications include a bachelor’s degree and ten or more years of experience, or an equivalent combination of education and work experience.
- Ability to convert customer-specific implementations, code assets, event-streaming patterns, and field learnings into reusable platform capabilities and enterprise standards.
Benefits
- Medical, dental, vision, life insurance, disability, accidental death and dismemberment, tax-preferred savings accounts, and a 401k plan are available to eligible regular teammates working 20 or more hours per week.
- At least 10 days of vacation, 10 sick days, and paid holidays are provided during the first year, prorated as applicable.
- Depending on position and division, benefits may include a defined benefit pension plan, restricted stock units, and/or deferred compensation plan.
- Regular, non-temporary employment; office-centric schedule requiring five days per week at one of the listed Truist hub locations.
Tech Stack
AnsibleApache KafkaBashC#DatadogGitGoGrafanaHelmJavaJavaScriptKubernetes.NETPowerShellPrometheusPythonSplunkTerraform
Categories
DevOpsForward Deployed