5 days ago
Base Salary
$155k - $190k/yr
Responsibilities
- Collaborate with technology and production teams to shape Platform as a Service implementations for stateless applications and AI infrastructure.
- Contribute to Agentic Engineering and MCP Server Development efforts.
- Establish governance, security, observability, reliability, and responsible-development practices for AI and agentic systems.
- Automate application deployment and administration across on-premises and cloud environments using Kubernetes orchestration and operators.
- Develop, deploy, and operate AI workloads on enterprise Kubernetes platforms across development and production environments.
- Monitor platform infrastructure and manage performance, alerts, and capacity using tools such as Datadog.
- Support cybersecurity vulnerability mitigation and compliance efforts.
- Create runbooks, operating procedures, system documentation, and architecture diagrams.
- Develop and maintain microservices, tools, and automation, and provide production troubleshooting and break-fix support.
Requirements
- Bachelor’s degree in Computer Science, Engineering, or a related technical field.
- 8+ years of experience in software, DevOps, or platform engineering, including 1+ year building AI systems for real workflows.
- Hands-on experience with cloud-native technologies, Docker, and Kubernetes.
- Experience developing and deploying AI agents, tool-calling systems, and multi-agent systems.
- Proficiency with Git and DevOps best practices.
- Expert Linux experience, including Red Hat or CentOS, and proficiency with Unix shell scripting such as Bash or C-Shell.
- Proficiency in Python, Java, or JavaScript development.
- Operational support experience monitoring and troubleshooting platform infrastructure with Datadog, Prometheus, or similar tools.
- Strong experience managing and troubleshooting production Kubernetes workloads.
- Experience developing or deploying custom application servers, RAG models, and MCP servers.
- Understanding of LLM integration and tuning, token-related operational costs, and GPU compute utilization.
- Deep expertise in at least one major cloud ecosystem, including AWS, Azure, OCI, or GCP, and infrastructure provisioning.
- Experience designing automation for complex, multi-stakeholder environments.
- Experience with AMD Inference Microservice Platform (AIMS) or NVIDIA NIM is a plus.
- Strong communication skills and the ability to work across artists, production management, and engineering teams.
Benefits
- Fully remote position, with work performed from a non-DreamWorks worksite, most commonly the employee’s residence.
- Company-sponsored medical, dental, and vision insurance.
- 401(k), paid leave, tuition reimbursement, discounts, and other perks.
Tech Stack
Categories
About NBC
NBCUniversal is a global media and entertainment company that produces and distributes film, television, news, sports, and streaming content for consumer audiences and advertisers. Its portfolio includes NBC, Telemundo, Universal Pictures, DreamWorks Animation, and Peacock, and it also operates Universal theme parks and consumer products licensing. Formed in 2004 and headquartered in New York City, it is a subsidiary of Comcast.
