18 hours ago
Durham, NC, USAStaff+
Responsibilities
- Architect and implement distributed AI/ML systems supporting model training, inference, observability, and continuous monitoring.
- Develop secure and trustworthy AI frameworks, including adversarial robustness pipelines, anomaly detection models, model governance, and compliance mechanisms.
- Build and optimize agentic AI workflows for data pipelines, model lifecycle operations, and system self-diagnostics.
- Research and prototype hybrid neural architectures, federated learning, adversarial learning, and multi-agent AI safety mechanisms.
- Evaluate AI systems for performance, reliability, robustness, and high-throughput operation under real-world and adversarial conditions.
- Integrate AI safety and model assurance into enterprise architectures with cybersecurity, cloud engineering, and data science teams.
- Mentor engineering teams on advanced ML algorithms, infrastructure patterns, and secure development practices.
Requirements
- Bachelor’s degree in Computer Science, Engineering, Information Technology, Information Systems, or a closely related field and five years of relevant experience, or a master’s degree in one of these fields and three years of relevant experience.
- Demonstrated expertise developing scalable, secure, distributed AI/ML applications and ML platform infrastructure across AWS, Azure, Google, or IBM cloud environments.
- Demonstrated expertise with IAM, encryption and KMS, infrastructure as code, Terraform, cloud-native platforms, Kubernetes, and Kubeflow.
- Experience with big data applications, ETL and batch processing, predictive analytics, autonomous multi-agent systems, and technologies including AWS Glue, EMR, Kinesis, Athena, DynamoDB, Hadoop, MongoDB, and PostgreSQL.
- Experience developing agentic workflows with Strands, CrewAI, LangGraph, OpenAI Swarm, MCP Server, and AGP.
- Experience building asynchronous and multithreaded Python and Java solutions, RAG pipelines, vector databases, OpenSearch, and Redis or Memcached.
- Experience automating deployment and MLOps with Airflow, AWS Step Functions, Jenkins, Git, MLflow, JFrog Artifactory, Go, and FastAPI.
- Experience implementing secure AI, adversarial robustness, federated learning, and governance using AIGAN, Google Federated Learning Framework, SageMaker Clarify, and MLflow.
- Experience with model serving and inference optimization using DJL, Triton, Flask, CPU/GPU acceleration, CloudWatch, Datadog, Splunk, Streamlit, and Gradio.
- Experience and/or expertise may be gained during a doctoral program.
Benefits
- Fidelity is transitioning to a full-time onsite working model through a phased rollout; onsite requirements vary by region and role and may evolve.
- The transition does not apply to fully remote roles.
Tech Stack
Amazon DynamoDBApache AirflowApache HadoopAWSAzureDatadogFastAPIFlaskGitGoGoogle CloudIBM CloudJavaJenkinsKubernetesMLflowMongoDBPostgreSQLPythonRedisSplunkTerraform
Categories
About Fidelity
Fidelity Investments provides brokerage, retirement plan recordkeeping, wealth management, and asset management services to individuals, employers, advisors, and institutions, plus online trading platforms and mutual funds and ETFs. It earns fees from managing and administering assets, advisory services, and brokerage transactions. Founded in 1946 and headquartered in Boston, it is privately held and administers trillions of dollars for U.S. and global customers.
