about 1 month ago
Responsibilities
- Build and operate the LLM control plane and gateway with routing, quotas, failover, and token and cost tracking.
- Ship unified APIs and SDKs with normalized schemas, structured outputs, caching, and observability across traces, logs, and metrics.
- Enforce content filtering, prompt and response validation, privacy controls, and PII redaction by default.
- Enable multi-model and multi-vendor LLM usage with automated canarying and versioning.
- Own the agent runtime, including tool registries, permissions, function calling, grounding, retrieval, state, and long-running workflows.
- Build training, scoring, experiment tracking, packaging, drift monitoring, retraining, and tuning components for classical ML and deep models.
- Develop human-in-the-loop review and safe-actioning mechanisms before agents interact with dealer systems.
- Evolve the domain graph and entity resolution while building reliable data ingestion and real-time context services with access controls and lineage.
- Implement hybrid graph, vector, and keyword retrieval plus caching and TTL strategies.
- Run offline and online evaluations for quality, factuality, bias, and safety.
- Define latency, uptime, and cost SLOs while enabling autoscaling and spend controls.
- Maintain model and agent registries, versioning, approvals, audit trails, reproducibility, and compliance support.
- Provide templates, CLIs, sandboxes, and documentation, mentor engineers, and champion MLOps and AI safety practices.
Requirements
- 5+ years building large-scale data, ML, or platform systems.
- Strong software engineering fundamentals in API design, concurrency, and distributed systems.
- Production experience with Python plus one of Java, Scala, or Go.
- Experience with microservices, API design, MLOps pipelines, model tracking and registries, model CI/CD, A/B testing, shadowing, canarying, and online feature computation.
- Experience with AWS, Docker, Kubernetes, and performance, reliability, and cost engineering for multi-tenant SaaS.
- Practical ML knowledge covering feature engineering, training, evaluation, drift detection, and deployment of models supporting user-facing workflows.
- Experience building or operating an LLM gateway or control plane with provider adapters, routing, policies, caching, quotas, rate limits, and cost and token accounting.
- Experience with agentic systems, including tool use, function calling, orchestration, human-in-the-loop workflows, safety guardrails, and online evaluation and telemetry.
- Experience with knowledge graphs, GraphQL, vector search, and hybrid retrieval patterns.
- Ability to support developer experience, observability, access control, documentation, teaching, vendor portability, and cost-aware system design.
Tech Stack
Apache AirflowApache FlinkApache KafkaApache SparkAWSDockerGoGraphQLgRPCJavaKubernetesLightGBMMLflowNeo4jPythonScalaXGBoost