1 day ago

Base Salary

$100k - $125k/yr

Responsibilities

  • Design and build synchronous and asynchronous model-serving services using FastAPI, SQS, Kafka, Docker, and Kubernetes/EKS.
  • Implement batch and real-time inference patterns with authentication, validation, logging, tracing, metrics, scalability, and observability.
  • Develop MLflow-based model training, experiment tracking, registry, deployment, promotion, and inference workflows across environments.
  • Build reproducible Databricks and PySpark data, training, inference, and feature-engineering pipelines using Delta Lake, Databricks Workflows, and Databricks Asset Bundles.
  • Create reusable MLOps frameworks, shared libraries, starter templates, engineering standards, and production-ready implementation patterns.
  • Design target-state architectures, technical diagrams, solution blueprints, implementation plans, and operational runbooks with technical and business stakeholders.
  • Monitor model performance, prediction drift, data quality, service health, and pipeline reliability; diagnose incidents and implement durable fixes.
  • Apply software engineering practices including testing, code reviews, CI/CD, dependency management, versioning, and infrastructure automation.
  • Use GitHub Copilot and Claude Code in day-to-day engineering workflows and document solutions for internal team ownership.

Requirements

  • Deep hands-on Python expertise for data engineering and backend application development.
  • Strong experience developing production-grade backend services with FastAPI.
  • Hands-on Databricks experience, including MLflow, Delta Lake, Databricks Workflows, Databricks Asset Bundles, and Spark-based distributed processing.
  • Experience designing and implementing end-to-end MLflow training and inference architectures across multiple environments.
  • Experience with AWS services including SQS, EKS, and Aurora PostgreSQL.
  • Experience implementing event-driven and asynchronous architectures with Kafka and/or SQS.
  • Strong understanding of the full machine learning lifecycle, including feature engineering, training, deployment, monitoring, and retraining.
  • Experience creating architecture diagrams, technical design documentation, implementation plans, and operational runbooks.
  • Strong Docker and Kubernetes fundamentals.
  • Experience with CI/CD pipelines using GitHub Actions, Jenkins, or similar tools.
  • Hands-on experience using GitHub Copilot and Claude Code as part of regular software engineering workflows.
  • Excellent written and verbal communication skills for collaboration with technical and business stakeholders.
  • Prior experience in Group Insurance, Life Insurance, or Underwriting is preferred.
  • Experience operationalizing GenAI or LLM-based applications and services is preferred.

Benefits

  • Health, dental, vision, life insurance, and disability plans are available to eligible full-time employees and hourly employees working more than 30 hours per week.
  • Benefits begin on the first day of employment.
  • Eligibility for the company 401(k) plan begins after 30 days of employment.
  • The company provides 11 paid holidays and 12 weeks of parental leave.
  • A free-time PTO policy provides flexibility for sick time and vacation.
  • The role is offered on a consulting basis, with the posting also describing full-time and eligible hourly-employee benefits.

Tech Stack

Apache KafkaApache SparkAWSDatabricksDockerFastAPIGitHub ActionsJenkinsKubernetesMLflowPython
Fractal Analytics

About Fractal Analytics

5,001-10,000 employees

Fractal builds enterprise AI and analytics solutions—delivered as consulting services and software products—for large global enterprises, especially Fortune 500 firms. Its teams embed AI in decisions across marketing, pricing, supply chain, forecasting, and customer experience. Founded in 2000 and headquartered in New York, Fractal is publicly listed and operates across North America, EMEA, and APAC.

Contact me