4 months ago
Raleigh, NC, USA or Dallas, TX, USASenior
Responsibilities
- Design, build, secure, maintain, optimize, and document Elastic Cloud Stack and AWS Managed Prometheus observability solutions.
- Design and configure ETL pipelines using Elastic Common Schema for application logs and metrics.
- Configure index templates and manage Elastic Index Lifecycle Management data retention policies.
- Develop Ansible playbooks to deploy Beat agents across on-premises and AWS systems.
- Use Terraform and Infrastructure as Code practices to manage production infrastructure in AWS.
- Create Elastic alerts using Watcher and Kibana Alerts and integrate them with ticketing tools and Microsoft Teams.
- Develop machine learning jobs and observability AI solutions for dynamic monitoring, incident detection, and proactive resolution.
- Collaborate with application owners, engineers, and development teams to gather requirements and architect monitoring solutions.
- Lead IT infrastructure monitoring projects, manage vendors, and provide daily SME escalation support.
- Tune Elastic clusters, indexing, search performance, security, configurations, and administration.
- Support solution transitions from development through QA and production and participate in agile team meetings.
Requirements
- A technical degree in Information Technology is required.
- Experience with Elastic Cloud and AWS Managed Prometheus is required.
- Knowledge of installation, system tasks, data collection, network troubleshooting, data pipelines, and cluster administration is required.
- Proficiency in Python, Bash, PowerShell, Painless, and other scripting languages is required.
- Extensive ELK Stack experience with Elasticsearch, Logstash, Kibana, Beats, Machine Learning, APM, X-Pack, and REST API integration is required.
- Experience with Prometheus, Grafana, AWS observability tools, and their performance, security, and management is required.
- Experience with Elasticsearch security integrations including Windows SAML, LDAP, and Kerberos is required.
- Experience with AWS CloudWatch, CloudTrail, Kubernetes, Docker, and Lambda is required.
- Experience integrating Elastic alerting with third-party ticketing tools is required.
- Experience implementing and integrating observability AI agents and frameworks is required.
Benefits
- Remote and hybrid working arrangements are available.
- Fully paid comprehensive health and well-being benefits are provided.
- The company offers career development frameworks, structured learning paths, training, and progression plans.
