5 months ago
Responsibilities
- Build and operate the LLM control plane and gateway with routing, quotas, failover, and token and cost tracking.
- Deliver unified APIs and SDKs with normalized schemas, structured outputs, caching, and traces, logs, and metrics.
- Enforce content filtering, prompt and response validation, privacy controls, and PII redaction.
- Enable multi-model and multi-vendor LLM usage with automated canarying and versioning.
- Own the agent runtime, including tool registration, permissions, function calling, grounding, retrieval, orchestration, state, and long-running workflows.
- Build platform components for classical ML and deep model training, scoring, experiment tracking, and packaging.
- Monitor model and data drift and support retraining and tuning to maintain accuracy and relevance.
- Implement human-in-the-loop review and safe actioning before agents interact with dealer systems.
- Evolve the domain graph and entity resolution and build reliable data ingestion pipelines.
- Serve real-time contextual data with access controls and lineage and support hybrid graph, vector, and keyword retrieval.
- Run offline and online evaluations for quality, factuality, bias, and safety.
- Define latency, uptime, and cost SLOs and enable autoscaling and spend controls.
- Maintain model and agent registries, versioning, approvals, audit trails, compliance support, and reproducibility.
- Provide templates, CLIs, sandboxes, and documentation while mentoring engineers and promoting MLOps and AI safety practices.
Requirements
- 5+ years building large-scale data, machine learning, or platform systems.
- Strong software engineering fundamentals in API design, concurrency, and distributed systems.
- Production experience with Python plus one of Java, Scala, or Go.
- Experience with microservices, API design, MLOps pipelines, model tracking and registries, model CI/CD, A/B testing, shadow and canary deployments, and online feature computation.
- Experience with AWS, Docker, Kubernetes, performance, reliability, and cost engineering in multi-tenant SaaS.
- Practical ML knowledge covering feature engineering, training, evaluation, drift detection, and deployment of models in user-facing workflows.
- Experience building or operating an LLM gateway or control plane with provider adapters, routing, policies, caching, quotas, rate limits, cost accounting, and token accounting.
- Experience with agentic systems, including tool use, function calling, orchestration, human-in-the-loop workflows, safety guardrails, and online evaluation telemetry.
- Experience with knowledge graphs, GraphQL, vector search, and hybrid retrieval patterns.
- Platform-as-a-product mindset with attention to developer experience, observability, fallback behavior, access control, cost awareness, vendor portability, documentation, and teaching.
Tech Stack
Apache AirflowApache FlinkApache KafkaApache SparkAWSDockerGoGraphQLgRPCJavaKubernetesLightGBMMLflowNeo4jPythonScalaXGBoost