Responsibilities
- Implement OpenTelemetry instrumentation, custom exporters, span enrichers, semantic conventions, context propagation hooks, and telemetry pipeline components.
- Build and maintain dashboards, alerting rules, SLO/SLA definitions, quality gates, automated quality reports, and observability data connectors.
- Instrument safety, security, responsible AI, governance, memory, MCP, agent fleet, marketplace, protocol, physical AI, and multimodal signals.
- Support red-team exercises, secure trace handling, PII redaction, audit-log retention, anomaly monitoring, and security observability processes.
- Build continuous quality monitoring and standardized evaluation harnesses for agentic solutions, including regression and behavioral-drift detection.
- Write production-grade Python observability tooling, SDKs, signal aggregators, anomaly detectors, and data transformation pipelines.
- Participate in on-call rotations, incident response, root-cause-analysis documentation, code reviews, testing, and observability platform documentation.
- Collaborate with AI platform engineers, data scientists, SRE, product teams, and security stakeholders.
Requirements
- Bachelor's or master's degree in Computer Science, Software Engineering, AI/ML, Data Science, or a related technical field.
- 11+ years of software engineering, platform engineering, or data engineering experience, including at least two years of hands-on observability, monitoring, or distributed systems work.
- Strong production Python skills, including async patterns, type hints, Poetry, Ruff, pytest, and maintainable testable code.
- Working knowledge of metrics, logs, traces, OpenTelemetry SDKs, trace context propagation, semantic conventions, distributed systems, event streaming, REST/gRPC APIs, and containerized deployment.
- Hands-on experience with Azure, AWS, or GCP, including managed services, IAM fundamentals, and cost awareness.
- Experience with CI/CD pipelines, GitOps, infrastructure-as-code concepts, automated testing, and data analysis or visualization using tools such as Grafana, Datadog, Splunk, or Prometheus.
- Hands-on experience with agentic AI frameworks such as LangChain, LangGraph, AutoGen, Semantic Kernel, or CrewAI.
- Preferred experience includes open-source observability or OTEL contributions, reinforcement learning or self-supervised learning, model fine-tuning, AI security tooling, responsible AI frameworks, fairness evaluation, explainability tools, and production AI platform, MLOps, or LLMOps deployments.
Tech Stack
Categories
About PepsiCo
PepsiCo is a playground for curious people. We invite thinkers, doers, and changemakers to champion innovation, take calculated risks, and challenge the status quo. From executives to team members on the front lines, we’re excited about the future. We take chances. Together, we dare to make the world a better place. Our associates are the magic ingredient. Each of them plays an integral role in helping create deep connections between people and our products. Think about your last group celebration: Chances are, one of our iconic brands was by your side. At PepsiCo, you’re invited to be a part of a global team of innovators who make, move, and sell these products—which are enjoyed by more than 1 billion people a day. A career at PepsiCo means working in a culture where everyone’s welcome. Here, you can dare to be yourself. No matter who you are or where you’re from, you can influence the people around you and the world at large. By showing up, you’ll have the opportunity to learn, develop and grow your skills for the future. Our supportive teams can fuel your professional goals to make a global impact on people and the planet. Join us. Dare for Better.
