1 day ago
London, United KingdomStaff+
Responsibilities
- Own the multi-year technical direction, architecture, consolidation, and deprecation strategy for a significant fleet-automation problem area.
- Design and evolve large-scale workflow orchestration engines involving workflow modeling, scheduling, conflict resolution, execution state machines, and cross-system contracts.
- Design safety guardrails including unavailability budgets, rate limiting, graceful cancellation, handbrakes, and blast-radius controls for production fleet operations.
- Lead through others by decomposing ambiguous work, reviewing designs, mentoring engineers, and owning consequential technical decisions.
- Drive alignment and formalize contracts and SLAs across capacity management, provisioning, remediation, hardware and data center operations, and external infrastructure providers.
- Replace one-off scripts, escalation paths, and per-case integrations with scalable platforms and abstractions that reduce operational load.
- Operate critical systems through on-call ownership, incident leadership, observability, SLO design, and structural reliability improvements.
- Raise engineering standards and contribute technical framing to roadmap and investment discussions.
Requirements
- BS/MS in Computer Science or equivalent practical experience.
- 8+ years of professional software engineering experience, including significant experience as the technical owner of a large production system.
- Staff-level scope leading the design and delivery of systems spanning multiple teams and planning cycles with measurable impact.
- Deep expertise in distributed systems such as workflow or orchestration engines, control planes, scheduling, state machines, or resource management systems.
- Professional production experience with one or more of Python, C++, Java, Rust, or Go.
- Experience operating critical systems, including on-call ownership, incident leadership, observability, SLO design, and reliability improvement.
- Ability to lead ambiguous cross-organizational work and communicate clearly with engineers and senior leadership.
- Experience developing engineers through mentoring, design reviews, and technical direction-setting.
- Preferred experience with constraint-based scheduling, planning, solver-backed systems, infrastructure automation, fleet or capacity management, hardware lifecycle, data center operations, or cloud infrastructure.
- Preferred experience applying AI or agentic automation to operational workflows, integrating AI tools, practicing responsible AI, and developing prompt/context engineering or agent orchestration skills.
- Preferred experience building workflow, orchestration, or job-execution engines; consolidating legacy systems; and integrating heterogeneous external infrastructure APIs and public cloud lifecycle models.
- Comfort working in large mature codebases with heavy cross-team dependencies.
About Meta
Meta builds social platforms and communication apps—including Facebook, Instagram, WhatsApp, and Messenger—and develops AR/VR hardware and software such as Quest to power immersive computing. It monetizes primarily through advertising tools for businesses, with additional revenue from devices and services, and operates a massive global infrastructure. Founded in 2004 and headquartered in Menlo Park, California, Meta Platforms, Inc. is a public company traded on Nasdaq under the ticker META.
