Lead Site Reliability Engineer (Performance & Scalability) | Contract | Remote MX
Tech Holding2 hours ago
Responsibilities
- Establish performance, throughput, latency, capacity, SLO, error-budget, dashboard, alert, and reliability baselines for critical workflows.
- Instrument and analyze request paths across application services, infrastructure, databases, networking, caches, queues, DNS, registries, and third-party services.
- Identify bottlenecks, lead cross-functional remediation, and drive architecture hardening, graceful degradation, dependency-failure planning, and resilience improvements.
- Build capacity and cost models that forecast platform limits, future constraints, and required infrastructure investments.
- Lead load, stress, soak, spike, failure, recovery, and scalability testing using representative demand scenarios.
- Partner with Test Automation and Scalability Engineering on automated performance testing, regression coverage, and production release gates.
- Own technical readiness assessments, production-readiness criteria, and go/no-go evidence for pilots, partnerships, and launches.
- Create runbooks for scale-up events, incidents, rollback, recovery, and dependency failures.
- Lead performance and reliability investigations during incidents and communicate risks, tradeoffs, and investment needs to leadership.
Requirements
- Significant experience in Site Reliability Engineering, performance engineering, platform engineering, distributed systems, or a closely related engineering discipline.
- Experience supporting production systems with meaningful scale, traffic, latency, or availability requirements.
- Deep understanding of observability, performance analysis, capacity planning, and reliability engineering.
- Strong hands-on experience with cloud infrastructure and production distributed systems.
- Deep knowledge of databases, networking, caching, queueing, compute, storage, and distributed-system failure modes.
- Experience defining and operating against SLOs, SLIs, error budgets, and production reliability metrics.
- Hands-on experience with load, stress, soak, scalability, and resilience testing.
- Ability to profile systems, diagnose bottlenecks, tune architecture, and implement improvements with engineering teams.
- Experience designing for graceful degradation, dependency failures, recovery, and high-demand scenarios.
- Strong incident management and root-cause analysis experience.
- Ability to communicate technical performance and reliability risks in clear business terms to senior leadership.
- Experience with high-scale SaaS, identity, DNS, registry, infrastructure, or other highly distributed platforms is preferred.
- Experience creating capacity-cost models, forecasting infrastructure requirements, building performance and reliability gates into CI/CD pipelines, or preparing platforms for enterprise-scale traffic is preferred.
Benefits
- Contract employment
- Remote position in Mexico
Categories
Site Reliability
About Tech Holding
Tech Holding is a full-service consulting firm that was founded on the premise of delivering predictable outcomes for our clients. Our founders and team members have industry experience and held senior positions in a wide variety of companies – from emerging startups to large Fortune 50 firms – and we have taken our combined experiences and developed a unique approach that is supported by the principles of deep expertise, integrity, transparency, and dependability.