5 hours ago
Remote, United StatesStaff+
Responsibilities
- Assess the customer’s observability estate, including vendors, agents, collectors, query surfaces, volumes, and operating processes.
- Analyze metric cardinality, active series, scrape targets, log volumes, trace sampling, retention, and telemetry fidelity.
- Build TCO comparisons covering vendor costs, self-managed infrastructure, data transfer, egress, and AWS estimates.
- Design the AWS-native target architecture across collection, pipelines, storage, querying, and alerting.
- Define retention, data classification, ownership tagging, and compliance archive policies with Security, Legal, and Engineering.
- Test performance, alert latency, and data fidelity against agreed success criteria.
- Lead migration execution, including pipeline cutover, OpenTelemetry log transform porting, dashboard rebuilding, and alert deduplication.
- Define OpenTelemetry conventions, collector topology, resource attributes, and sampling strategies.
- Lead the embedded TechPod, coordinate with customer observability leadership, and present technical and commercial recommendations.
Requirements
- 8+ years of experience in SRE, DevOps, observability, or platform engineering, including 4+ years owning production observability platforms and experience at technical lead, staff, or principal level.
- Deep experience operating metrics, logging, and tracing platforms at very large scale and improving their cost and reliability.
- Advanced production experience with Prometheus and a long-term storage system such as Thanos, Cortex, or Mimir, including PromQL, recording rules, remote write, and cardinality management.
- Hands-on experience with Datadog or a comparable commercial observability platform and its pricing models.
- Production experience with AWS observability services, OpenTelemetry Collector or ADOT, log pipelines, Kubernetes, and Terraform.
- Experience with SLO-based alerting, incident management tooling, cost modeling, telemetry assessment, and automation using Python, Go, or Bash.
- Ability to explain technical, cost, privacy, and compliance tradeoffs to engineers, Security, Legal, and executive stakeholders.
- Preferred experience includes vendor migrations, dashboard and alert automation, Grafana ecosystem tools, continuous profiling, data governance, analytics platforms, consumer-scale systems, consulting engagements, and relevant certifications.
Benefits
- 100% remote workplace.
- Unlimited paid time off.
- Equity participation.
- 401(k) with company contribution and sponsored healthcare.
- Training and certification programs for professional growth.
Tech Stack
Categories
Forward Deployed
About EverOps
EverOps is an IT services firm that helps engineering organizations modernize cloud and IT operations, improve security, and control cloud spend. It delivers managed services and project-based consulting through embedded TechPod teams that plan, build, and run DevOps, SRE, NOC, and automation workflows. Founded in 2012 and headquartered in San Francisco, it is privately held and serves technology-driven companies migrating to or scaling on public cloud platforms.
