11 hours ago
Hyderābād, IndiaStaff+
Responsibilities
- Design and implement OpenTelemetry instrumentation, custom exporters, span enrichment, semantic conventions, distributed tracing, and telemetry pipelines for assigned AI platforms.
- Deploy and operate collectors, processors, exporters, dashboards, alerts, SLOs, and production on-call capabilities for metrics, logs, traces, and governance signals.
- Instrument AI safety, red-teaming, Responsible AI, policy, guardrail, audit, and security signals, including PII redaction and secure trace handling.
- Build quality gates, evaluation monitoring, regression detection, performance alerts, automated quality reports, and post-go-live monitoring for agentic solutions.
- Instrument agent memory, MCP interactions, harnesses, reinforcement-learning components, fleet coordination, marketplaces, registries, and agent communication protocols.
- Write production-quality Python telemetry components and apply time-series and statistical analysis to tune thresholds and reduce alert noise.
- Collaborate with peer architects, contribute to engineering standards and documentation, participate in agile ceremonies, and review code.
Requirements
- Bachelor's or master's degree in Computer Science, Software Engineering, AI/ML, Data Science, or a related technical field.
- 10–12 years of overall experience, including 5–10 years of professional software and AI/ML engineering or platform engineering experience and at least 2 years of observability, distributed-systems monitoring, or telemetry-pipeline development.
- Demonstrated ability to deliver production-grade software from design through deployment and on-call operation.
- Hands-on OpenTelemetry SDK experience covering custom exporters, collector configuration, semantic conventions, and distributed trace propagation.
- Strong Python skills, including async patterns, type hints, pytest, CI/CD integration, and production telemetry tooling.
- Working knowledge of distributed systems, Kafka or equivalent event streaming, containerized deployment with Docker and Kubernetes, and cloud platforms such as Azure, AWS, or GCP.
- Experience analyzing time-series and log data with Grafana, Datadog, Prometheus, Splunk, or equivalent tools.
- Familiarity with CI/CD pipelines, GitOps, automated testing, infrastructure-as-code, LLM workflows, agentic AI, RAG, memory, tool calling, and multi-step planning.
- Understanding of AI safety, guardrails, policy enforcement, prompt injection, PII handling, access control, audit logging, quality engineering, evaluation frameworks, and Responsible AI principles.
- Preferred experience with LangChain, LangGraph, AutoGen, Semantic Kernel, CrewAI, Bedrock Agents, MCP, A2A, UCP, AP2, reinforcement learning, vector databases, semantic search, AI safety frameworks, adversarial ML, red-team tooling, or open-source observability and AI projects.
Tech Stack
Categories
About PepsiCo
PepsiCo is a global food and beverage company that manufactures, markets, and distributes snacks and drinks under brands such as Pepsi, Mountain Dew, Gatorade, Lay’s, Doritos, and Quaker Oats. It sells to consumers through retail, foodservice, and e-commerce channels, with products available in more than 200 countries and territories. Founded in 1965 and headquartered in Purchase, New York, PepsiCo is a public company listed on NASDAQ under the symbol PEP.
