9 hours ago
Mexico City, MexicoSenior
Responsibilities
- Own complex and high-impact customer investigations involving AI applications and agentic workflows.
- Lead structured investigations, root-cause analysis, recovery efforts, and technical escalations across application, platform, and data layers.
- Investigate agent workflows, tool calls, model-provider integrations, memory, checkpoints, streaming, retries, and timeouts.
- Read and debug Go, TypeScript, and Python services, SDKs, and agent implementations using logs, metrics, traces, and diagnostic tooling.
- Troubleshoot Kubernetes and containerized workloads, including deployments, images, secrets, resource failures, and networking.
- Create troubleshooting methodologies, knowledge articles, diagnostic scripts, and reusable runbooks.
- Mentor and coach engineers in AI fundamentals, cloud-native operations, customer communication, and troubleshooting methodology.
- Partner with Engineering, Product, Security, and Technical Services teams to communicate customer impact and influence product improvements.
- Review recurring issues and improve support processes, tooling, and customer self-service.
- Contribute to MongoDB database and cloud support cases when needed.
Requirements
- Typically 8+ years of relevant technical experience, including significant experience troubleshooting production systems, supporting complex customers, or leading technical escalations; equivalent demonstrated expertise is considered.
- Hands-on experience building or troubleshooting agentic AI applications and related workflows, including LangGraph, A2A, MCP, tool calling, model-provider integrations, memory, checkpoints, streaming, and human-in-the-loop workflows.
- Ability to read and debug Go, TypeScript, and Python services, SDKs, and agent workflows, including asynchronous execution, HTTP clients, retries, and timeouts.
- Experience with Docker and Kubernetes, including deployments, containers, networking, secrets, resource management, and deployment troubleshooting.
- Familiarity with REST/JSON APIs, streaming protocols, API keys, OAuth/JWT, RBAC, and service accounts.
- Experience using metrics, events, logs, traces, error monitoring, and alerting to diagnose production issues.
- Experience investigating distributed-system issues across application, platform, and data layers and driving customer problems through resolution.
- Bonus experience with MongoDB, MongoDB Atlas, Atlas Vector Search, Voyage AI, another distributed database, AWS, Azure, GCP, Terraform, GitHub Actions, CodeBuild, Helm, operators, container registries, or advanced Kubernetes networking.
Benefits
- Hybrid working model for candidates based in Mexico City.
- Supportive culture with employee affinity groups, fertility assistance, and generous parental leave.
- Disability accommodations are available during the application and interview process.
Tech Stack
Categories
Solutions Engineering
About MongoDB
MongoDB builds the MongoDB document database and the MongoDB Atlas cloud database service for application developers and enterprises. Its business model centers on cloud consumption (Atlas), commercial subscriptions, and support, with tooling for migration, synchronization, and disaster recovery across on‑prem and multi‑cloud environments. Founded in 2007 and headquartered in New York, it is a public company listed on NASDAQ and widely used across industries worldwide.
