7 months ago
Responsibilities
- Convert tabular, NLP/LLM, recommendation, and forecasting models into fast, reliable production services with defined API contracts.
- Integrate with the LLM Gateway/MCP and manage prompt and configuration versioning.
- Build and orchestrate CI/CD pipelines and develop reusable templates, SDKs, CLIs, sandbox datasets, and documentation.
- Review, refactor, optimize, containerize, deploy, version, and monitor data science models.
- Design batch and stream pipelines, online features, and domain-graph integrations.
- Build inference microservices with REST/gRPC, schema versioning, structured outputs, and stringent p95 latency targets.
- Manage feature stores, model and agent registries, versioning, lineage, and the overall model/feature lifecycle.
- Implement prompt versioning, A/B testing, guardrails, and feedback- and metrics-based orchestration.
- Instrument observability for traces, logs, metrics, drift, model performance, safety signals, and cost tracking.
- Ensure runtime reliability through autoscaling, caching, circuit breakers, retries, fallbacks, and graceful degradation.
- Collaborate with Applied Sciences, data engineering, product management, and architecture teams to deliver enterprise ML systems.
- Monitor and mitigate risks associated with LLMs and agentic systems.
Requirements
- 8–11+ years of experience in ML engineering, MLOps, or backend/platform engineering with production ML.
- Experience with LLMs, retrieval systems, vector stores, and graph or knowledge stores.
- Strong Python skills plus one of Java, Go, or Scala, along with API design, concurrency, and testing experience.
- Hands-on experience with orchestration frameworks and libraries such as LangChain, LlamaIndex, OpenAI Function Calling, or AgentKit.
- Knowledge of reactive, planning, and retrieval-augmented agent architectures and safe execution patterns.
- Experience with Airflow or Kubeflow, Spark or Flink, Kafka or Kinesis, and strong data quality practices.
- Experience with Docker, Kubernetes, service meshes, REST/gRPC, and performance and reliability engineering.
- Experience with experiment tracking, model registries such as MLflow, feature stores, A/B testing, shadow testing, and drift detection.
- Experience with OpenTelemetry, Prometheus, and Grafana, including diagnosing latency, tail behavior, and memory or CPU hotspots.
- Experience with AWS services including IAM, ECS/EKS, S3, RDS/DynamoDB, and Step Functions/Lambda, with cost optimization experience.
- Knowledge of secrets management, RBAC/ABAC, PII handling, and auditability.
- Preferred: product-oriented, reliability- and safety-first, systems-thinking, collaborative, and pragmatic approach.
Benefits
- Opportunity to work on a global, cloud-native, AI-driven automotive platform.
- Collaborative, innovative, and fast-paced work environment.
- Current Tekion employees should apply through the internal Ashby job board effective August 4, 2026.
Tech Stack
Amazon DynamoDBApache AirflowApache FlinkApache KafkaApache SparkAWSDockerGoGrafanagRPCJavaKubernetesMLflowPrometheusPythonScala