about 1 month ago
Responsibilities
- Build and operate the LLM control plane and gateway with routing, quotas, failover, and token and cost tracking.
- Ship unified APIs and SDKs with normalized schemas, structured outputs, caching, and observability through traces, logs, and metrics.
- Enforce content filtering, prompt and response validation, PII redaction, access controls, and safe actioning.
- Support multi-model and multi-vendor LLM usage with automated canarying, versioning, approvals, audit trails, and reproducibility.
- Own the agent runtime, including tool registration, permissions, function calling, grounding, retrieval, orchestration, state, and long-running workflows.
- Enable training and scoring pipelines for classical ML and deep learning models, including experiment tracking, packaging, monitoring, retraining, and tuning.
- Evolve the domain graph, entity resolution, data ingestion pipelines, and real-time context services for dealership data.
- Implement hybrid graph, vector, and keyword retrieval plus caching and TTL strategies to balance accuracy, latency, and cost.
- Run offline and online evaluations for quality, factuality, bias, and safety, and define SLOs for latency, uptime, and cost.
- Provide templates, CLIs, sandboxes, documentation, and mentorship to help product teams build and deploy AI systems safely and efficiently.
Requirements
- At least 8 years of experience building large-scale data, machine learning, or platform systems.
- Strong software engineering fundamentals, including API design, concurrency, and distributed systems.
- Production experience with Python and at least one of Java, Scala, or Go.
- Experience with microservices, API design, MLOps pipelines, experiment tracking, model registries, model CI/CD, A/B testing, shadow and canary deployments, and online feature computation.
- Experience with AWS, Docker, Kubernetes, and performance, reliability, and cost engineering for multi-tenant SaaS.
- Practical knowledge of feature engineering, model training and evaluation, drift detection, and deploying models for user-facing workflows.
- Experience building or operating an LLM gateway or control plane with provider adapters, routing, policies, caching, quotas, rate limits, and cost and token accounting.
- Experience with agentic systems, including tool use, function calling, orchestration frameworks, human-in-the-loop workflows, safety guardrails, and online evaluation and telemetry.
- Experience with knowledge graphs, GraphQL, vector search, and hybrid retrieval patterns.
- Ability to design for developer experience, observability, fallback behavior, access control, portability, resilience, cost awareness, documentation, and team enablement.
Benefits
- The role supports remote or location-flexible work only where stated in the posting; no specific work arrangement or other benefits are provided.
- Employees should apply through the Internal Job Board in Ashby effective 4 August 2026 if they are current Tekion employees.
Tech Stack
Apache AirflowApache FlinkApache KafkaApache SparkAWSDockerGoGraphQLgRPCJavaKubernetesLightGBMMLflowNeo4jPythonScalaXGBoost