Tech Holding

Lead Site Reliability Engineer (Performance & Scalability) | Contract | Remote MX

Tech Holding
Apply
2 hours ago
Remote, MexicoStaff+
H1B Sponsor

Responsibilities

  • Establish performance, throughput, latency, capacity, SLO, error-budget, dashboard, alert, and reliability baselines for critical workflows.
  • Instrument and analyze request paths across application services, infrastructure, databases, networking, caches, queues, DNS, registries, and third-party services.
  • Identify bottlenecks, lead cross-functional remediation, and drive architecture hardening, graceful degradation, dependency-failure planning, and resilience improvements.
  • Build capacity and cost models that forecast platform limits, future constraints, and required infrastructure investments.
  • Lead load, stress, soak, spike, failure, recovery, and scalability testing using representative demand scenarios.
  • Partner with Test Automation and Scalability Engineering on automated performance testing, regression coverage, and production release gates.
  • Own technical readiness assessments, production-readiness criteria, and go/no-go evidence for pilots, partnerships, and launches.
  • Create runbooks for scale-up events, incidents, rollback, recovery, and dependency failures.
  • Lead performance and reliability investigations during incidents and communicate risks, tradeoffs, and investment needs to leadership.

Requirements

  • Significant experience in Site Reliability Engineering, performance engineering, platform engineering, distributed systems, or a closely related engineering discipline.
  • Experience supporting production systems with meaningful scale, traffic, latency, or availability requirements.
  • Deep understanding of observability, performance analysis, capacity planning, and reliability engineering.
  • Strong hands-on experience with cloud infrastructure and production distributed systems.
  • Deep knowledge of databases, networking, caching, queueing, compute, storage, and distributed-system failure modes.
  • Experience defining and operating against SLOs, SLIs, error budgets, and production reliability metrics.
  • Hands-on experience with load, stress, soak, scalability, and resilience testing.
  • Ability to profile systems, diagnose bottlenecks, tune architecture, and implement improvements with engineering teams.
  • Experience designing for graceful degradation, dependency failures, recovery, and high-demand scenarios.
  • Strong incident management and root-cause analysis experience.
  • Ability to communicate technical performance and reliability risks in clear business terms to senior leadership.
  • Experience with high-scale SaaS, identity, DNS, registry, infrastructure, or other highly distributed platforms is preferred.
  • Experience creating capacity-cost models, forecasting infrastructure requirements, building performance and reliability gates into CI/CD pipelines, or preparing platforms for enterprise-scale traffic is preferred.

Benefits

  • Contract employment
  • Remote position in Mexico

Categories

Site Reliability
Tech Holding

About Tech Holding

201-500 employees

Tech Holding is a full-service consulting firm that was founded on the premise of delivering predictable outcomes for our clients. Our founders and team members have industry experience and held senior positions in a wide variety of companies – from emerging startups to large Fortune 50 firms – and we have taken our combined experiences and developed a unique approach that is supported by the principles of deep expertise, integrity, transparency, and dependability.